====== SAP Cloud ALM Health Monitoring ====== The **SAP Cloud ALM Health Monitoring** monitor retrieves health metrics from your SAP Cloud ALM tenant and publishes them into the monitoring system. It can also evaluate thresholds and raise alarms based on a selected measure per row. ===== Prerequisites ===== Network connectivity from the collector to your SAP Cloud ALM tenant API endpoint (e.g. ''https://your-tenant.api.xxx.hana.ondemand.com''). ==== Cloud ALM Connector (required) ==== The monitor requires a **Web Service** connector with authentication type **CLOUD_ALM**. [[..:sap_cloud_alm|SAP Cloud ALM Connector]] This connector is based on a service key of the **SAP Cloud ALM API** service instance in the BTP subaccount that contains your Cloud ALM entitlement. The following OAuth scopes **must** be included in the instance parameters (in the ''authorities'' list): * ''$XSMASTERAPPNAME.calm-api.hm.read'' * ''$XSMASTERAPPNAME.calm-api.metrics.read'' === Where to add the scopes === Even if you already created a service key for the ALM connector, you do **not** need to delete the instance. - In the BTP Cockpit go to your subaccount → **Services** → **Instances and Subscriptions**. - Locate your **SAP Cloud ALM API** instance. - Click the **⋯** (three dots) menu next to the instance. - Select **Update** to edit the instance. - Click **Next**. - Add the two scopes above to the ''authorities'' array in the parameters JSON and save. ===== API Endpoints used ===== ^ Endpoint ^ Purpose ^ | POST /oauth/token | Authentication (BTP UAA) | | GET /api/calm-metrics/v1/metrics?provider=hm | Retrieve health monitoring metrics | The metrics endpoint is paged. The monitor reads all pages on every run. If any page fails, the run fails with the HTTP code and the API message, so that no partial data is published. Each run reads the latest 5 minute bucket that has data. The monitor samples the state at run time, it does not replay the buckets between two runs. With a schedule longer than 5 minutes, a short peak between two runs is not seen. Keep the schedule at 5 or 10 minutes when the thresholds matter. ===== Key Features ===== ==== Metrics Collection ==== The monitor can collect and publish the following measures for Cloud ALM health metrics: * **usage** – Percentage utilization * **value** – Current measured value * **limit** – Configured or maximum limit * **okStatus** – Percentage of the time bucket in OK status * **warningStatus** – Percentage of the time bucket in warning status * **criticalStatus** – Percentage of the time bucket in critical status Only **usage** is published with the **PERCENT** unit. The other measures are published without a unit. ==== Alarm Capabilities ==== * Per-row thresholds using the standard ''G2W:80 W2M:90'' syntax. [[..:commonsettings#multi_thresholds_syntax|Multi Threshold Syntax]] * Choose what to alarm on: **NONE**, **USAGE**, **VALUE**, **OK_STATUS**, **WARNING_STATUS**, or **CRITICAL_STATUS** * Alarm evaluation is performed only on the selected measure for each row * **Attributes Filter** can further restrict which datapoints are evaluated * **Exclusive** controls whether a matching datapoint can be consumed by later rows * Wildcard support (''*'') for **Metric** and **Service name**, alone or inside a name, e.g. ''S4H*'' or ''*Workprocess*'' * Optional alarm tag for grouping or routing * Automatic clear is supported by the monitoring framework when the alarm condition no longer matches ==== Metric Publishing Behavior ==== The monitor supports both global and row-level metric publishing: * **Send all metrics = true** * publishes all supported datapoints for all returned metrics * metric names keep the original Cloud ALM measure in parentheses * the row-level **Send metric** setting is not used * **Send all metrics = false** * publishes only the selected measure for rows where **Send metric** is enabled * datapoints excluded by the **Attributes Filter** of the row are not published When row-level publishing is used, the monitor publishes the measure selected in **Alarm on** for matching rows. ==== Advanced Discovery ==== * **Load ALM Health Metrics** button: Automatically discovers metric names and service names reported by Cloud ALM during the last 24 hours and populates the configuration table * The loaded metrics also expose the available datapoint attributes, so you can use them to build an **Attributes Filter** for each row * The Value, Limit and Usage columns of the loaded list are the 24 hour aggregate, not the current value * Metrics reported once a day, such as ''ABAP.Certificates'', appear in the list but are present in the monitor run only during the 5 minutes after their collection. Alarms on them are not reliable at the default schedule ===== Configuration ===== ==== Monitor Configuration ==== === Method 1: Load ALM Health Metrics (Recommended) === - Open the monitor configuration. - Click the **Load ALM Health Metrics** button. - The table will be populated with **METRIC**, **SERVICE_NAME** and the available attribute names from your live Cloud ALM tenant. - Activate the rows you want, define thresholds, choose **Alarm on**, and save. === Method 2: Wildcard / Manual Mode === * Set **Metric** = ''*'' and/or **Service name** = ''*'' to monitor everything or a broad subset * A ''*'' can also be part of a name, for example **Service name** = ''S4H*'' or **Metric** = ''ABAP.*'' * If **Send all metrics** is enabled, all collected datapoints are published * If **Send all metrics** is disabled, only rows with **Send metric** enabled will publish performance data === Settings Reference === ^ Field ^ Description ^ Default ^ | Active | Enable/disable this configuration row | ''true'' | | Service name | Filter by service name, ''*'' matches any part of the name | ''*'' | | Metric | Filter by metric name, ''*'' matches any part of the name | ''*'' | | Attributes Filter | Optional filter applied to datapoint attributes before alarm evaluation / row matching, see below | (empty) | | Alarm on | Which measure to evaluate for alarms (NONE / USAGE / VALUE / OK_STATUS / WARNING_STATUS / CRITICAL_STATUS) | ''USAGE'' | | Thresholds | Multi-level thresholds (G2W / W2M / etc.) | ''G2W:80 W2M:90'' | | Alarm tag | Optional tag added to the alarm message | (empty) | | Exclusive | If enabled, matching datapoints of this row can be consumed and not reused by later rows | ''true'' | | Alarm | Enable/disable alarm creation for this row | ''true'' | | Send metric | Publish the selected measure for this row when **Send all metrics** is disabled | ''true'' | | Send all metrics | Publish all returned metrics and measures globally | ''true'' | ==== Attributes Filter ==== The **Attributes Filter** is an optional rule used to narrow down which datapoints a row can match. The filter is a comma separated list of clauses. All clauses must match for the datapoint to be used by the row. ^ Clause ^ Meaning ^ Example ^ | ''name'' | The datapoint has the attribute, whatever its value | ''Wp_type'' | | ''name:value'' | The datapoint has the attribute with this exact value | ''Wp_type:DIA'' | Attribute names and values are compared without regard to case. Resource attributes such as ''service.name'' can be used as well as datapoint attributes. When you use **Load ALM Health Metrics**, the **Attributes Filter** column of the loaded rows lists the attribute names available for each metric, for example ''Wp_type'' or ''Client''. Loaded as is, such a filter only requires the attributes to be present. Add '':value'' to a name to restrict the row to one value. It checks the datapoint’s attributes before the row is used for: * **Alarm evaluation** * **Metric publishing** when the row is processed in row-based mode A clause without a name or without a value after the colon, such as ''Client:'', is ignored and reported in the monitor log. If no valid clause remains, the row is not restricted by attributes. ===== Collected Metrics ===== All metrics are stored with the base key: ''promonitor.cloud_alm.hm.*'' Supported metric keys include: ^ Metric Key ^ Unit ^ Description ^ Tags ^ | promonitor.cloud_alm.hm.''''.usage | % | Current utilization percentage | service.name, sap.service.name, service.namespace, service.instance.id, sap.service.display_name, plus datapoint-specific tags | | promonitor.cloud_alm.hm.''''.value | - | Current measured value | same as above | | promonitor.cloud_alm.hm.''''.limit | - | Configured limit | same as above | | promonitor.cloud_alm.hm.''''.okstatus | - | OK status datapoint | same as above | | promonitor.cloud_alm.hm.''''.warningstatus | - | Warning status datapoint | same as above | | promonitor.cloud_alm.hm.''''.criticalstatus | - | Critical status datapoint | same as above | The displayed metric name in the UI is: '' (usage)'', '' (value)'', '' (limit)'', '' (okStatus)'', '' (warningStatus)'', or '' (criticalStatus)'' ===== Alarm Evaluation Notes ===== * Alarm evaluation is performed only on the measure selected in **Alarm on** * If **Alarm on = NONE**, no alarm is generated for that row * The monitor checks datapoints with the selected measure name and evaluates the configured thresholds * Thresholds fire when the value reaches the configured level. The status measures are the percentage of the time bucket spent in that status. To alarm on a status, use **WARNING_STATUS** or **CRITICAL_STATUS** with ''G2W:1'', which fires as soon as any time was spent in that status. **OK_STATUS** is 100 when the service is fine, so a threshold on it cannot detect a failure * If **Exclusive** is enabled, a matching datapoint can be consumed so that later rows do not evaluate it again. A row consumes only the datapoints of its selected measure, except a **NONE** row, which consumes all datapoints of the matched metrics. Use a **NONE** row with **Exclusive** to keep a metric away from the rows below * **Attributes Filter** can restrict evaluation to datapoints with matching attributes * A row raises **one alarm per matching datapoint**. Each datapoint has its own alarm, identified by the service, the metric, the measure, the datapoint attributes and the row. The alarm message lists the datapoint attributes, for example ''[Tags: Wp_type=DIA]'' * A row with an invalid threshold is skipped with a warning in the monitor log. The other rows are still evaluated ===== Examples ===== ==== 1. Publish all metrics for all services ==== No rows needed. Empty table = publish everything. Send all metrics = true by default. ==== 2. Alarm on usage above 80% for all metrics ==== ^ Active ^ Service name ^ Metric ^ Attributes Filter ^ Alarm on ^ Thresholds ^ Alarm tag ^ Exclusive ^ Alarm ^ Send metric ^ | true | * | * | (empty) | USAGE | G2W:80 W2M:90 | (empty) | true | true | true | Alarms when ''usage'' crosses 80% on any metric for any service. ==== 3. Alarm on usage for one service only ==== ^ Active ^ Service name ^ Metric ^ Attributes Filter ^ Alarm on ^ Thresholds ^ Alarm tag ^ Exclusive ^ Alarm ^ Send metric ^ | true | S4H.100 | * | (empty) | USAGE | G2W:80 W2M:90 | (empty) | true | true | true | Only datapoints for ''S4H.100'' are evaluated. Other services not matched. ==== 4. Alarm on ABAP work process usage with stricter threshold ==== ^ Active ^ Service name ^ Metric ^ Attributes Filter ^ Alarm on ^ Thresholds ^ Alarm tag ^ Exclusive ^ Alarm ^ Send metric ^ | true | S4H.100 | ABAP.Workprocess.Usage | (empty) | USAGE | G2W:70 W2M:85 | WP_OVERLOAD | true | true | true | Stricter threshold for work process usage only. Other metrics for ''S4H.100'' not matched by this row. ==== 5. Alarm on dialog work process usage only ==== The ''ABAP.Workprocess.Usage'' metric returns one datapoint per ''Wp_type'' (''DIA'' ''BTC'' ''UPD'' ''UP2'' ''SPO''). Use ''Attributes Filter'' to target dialog processes only. ^ Active ^ Service name ^ Metric ^ Attributes Filter ^ Alarm on ^ Thresholds ^ Alarm tag ^ Exclusive ^ Alarm ^ Send metric ^ | true | S4H.100 | ABAP.Workprocess.Usage | Wp_type:DIA | USAGE | G2W:70 W2M:85 | (empty) | true | true | true | Only datapoints where ''Wp_type = DIA'' are evaluated. Batch and update processes not alarmed. ==== 6. Alarm on short dumps per client ==== ''ABAP.ShortDumps.Today'' returns one datapoint per ''Client''. To alarm only on client 100: ^ Active ^ Service name ^ Metric ^ Attributes Filter ^ Alarm on ^ Thresholds ^ Alarm tag ^ Exclusive ^ Alarm ^ Send metric ^ | true | S4H.100 | ABAP.ShortDumps.Today | Client:100 | VALUE | G2W:50 W2M:200 | DUMPS_100 | true | true | true | ''Alarm on = VALUE'' because this metric has no usage percentage. Threshold is a raw count. ==== 7. Alarm on aborted batch jobs per client ==== ^ Active ^ Service name ^ Metric ^ Attributes Filter ^ Alarm on ^ Thresholds ^ Alarm tag ^ Exclusive ^ Alarm ^ Send metric ^ | true | S4H.100 | ABAP.Batchjobs.Aborted | Client:100 | VALUE | G2W:1 W2M:5 | BTC_ABORT | true | true | true | Alarms as soon as 1 aborted job appears in client 100. ''G2W:1'' means any non-zero value triggers warning. ==== 8. Alarm on HANA disk usage ==== ''Hana.Disk.Usage'' returns one datapoint per ''Usagetype'' (''DATA'' ''LOG'' ''TRACE'' ''DATA_BACKUP'' etc.). ^ Active ^ Service name ^ Metric ^ Attributes Filter ^ Alarm on ^ Thresholds ^ Alarm tag ^ Exclusive ^ Alarm ^ Send metric ^ | true | S4H.100 | Hana.Disk.Usage | Usagetype:DATA | VALUE | G2W:300 W2M:400 | HANA_DISK | true | true | true | Value is in GB. Threshold is a raw size. Adjust to your disk capacity. ==== 9. Alarm on batch response time ==== ''ABAP.ResponseTime'' returns one datapoint per ''Tasktype'' (''RFC'' ''Batch'' ''Update''). To alarm on slow batch jobs: ^ Active ^ Service name ^ Metric ^ Attributes Filter ^ Alarm on ^ Thresholds ^ Alarm tag ^ Exclusive ^ Alarm ^ Send metric ^ | true | S4H.100 | ABAP.ResponseTime | Tasktype:Batch | VALUE | G2W:2000 W2M:5000 | SLOW_BTC | true | true | true | Value is in ms. ''G2W:2000'' = warning above 2 seconds. ''W2M:5000'' = major above 5 seconds. ==== 10. Alarm without publishing metrics ==== ^ Active ^ Service name ^ Metric ^ Attributes Filter ^ Alarm on ^ Thresholds ^ Alarm tag ^ Exclusive ^ Alarm ^ Send metric ^ | true | * | * | (empty) | USAGE | G2W:80 W2M:90 | (empty) | true | true | false | ''Send metric = false'' and **Send all metrics** disabled: alarms fire but no datapoints written to time-series. ==== 11. Two services with different thresholds ==== ^ Active ^ Service name ^ Metric ^ Attributes Filter ^ Alarm on ^ Thresholds ^ Alarm tag ^ Exclusive ^ Alarm ^ Send metric ^ | true | S4H.100 | ABAP.Workprocess.Usage | (empty) | USAGE | G2W:70 W2M:85 | S4H_PROD | true | true | true | | true | * | ABAP.Workprocess.Usage | (empty) | USAGE | G2W:80 W2M:95 | (empty) | true | true | true | Row 1 consumes ''S4H.100'' datapoints (''Exclusive = true''). Row 2 applies a looser threshold to all other services. ==== 12. One threshold for all production systems ==== ^ Active ^ Service name ^ Metric ^ Attributes Filter ^ Alarm on ^ Thresholds ^ Alarm tag ^ Exclusive ^ Alarm ^ Send metric ^ | true | *PRD* | * | (empty) | USAGE | G2W:70 W2M:85 | PROD | true | true | true | | true | * | * | (empty) | USAGE | G2W:85 W2M:95 | (empty) | true | true | true | Row 1 matches every service whose name contains ''PRD''. Row 2 handles the other services with a looser threshold. ==== 13. Suppress a metric from alarming and publishing ==== ''DataCollector.Version'' and ''ABAP.Version'' are informational metrics that report 0 and should not alarm. ^ Active ^ Service name ^ Metric ^ Attributes Filter ^ Alarm on ^ Thresholds ^ Alarm tag ^ Exclusive ^ Alarm ^ Send metric ^ | true | * | DataCollector.Version | (empty) | NONE | G2W:80 W2M:90 | (empty) | true | false | false | | true | * | ABAP.Version | (empty) | NONE | G2W:80 W2M:90 | (empty) | true | false | false | | true | * | * | (empty) | USAGE | G2W:80 W2M:90 | (empty) | true | true | true | Rows 1 and 2 consume the version metrics (''Exclusive = true'', ''Alarm = false'', ''Send metric = false''). Row 3 handles everything else. ===== Troubleshooting ===== * **"No metrics returned"** or empty table after Load: Check that the connector uses **CLOUD_ALM** authentication and that the service key from the BTP **SAP Cloud ALM API** instance is valid. Also verify the tenant has Health Monitoring data. * **HTTP 401/403**: Token issue – regenerate the service key in the BTP instance. The monitor run fails and shows the HTTP code and the message returned by the API. * **Monitor run fails with "page fetch failed"**: One page of the metrics endpoint could not be read. The run is not published to avoid clearing alarms on missing data. Check the connectivity and the API message in the monitor log. * **Metrics show 0 but data exists in Cloud ALM**: Some metrics may legitimately report 0 for one or more measures. * **Alarms not triggering**: * Verify the row is **Active** * Verify **Alarm** is enabled * Verify **Alarm on** matches the intended measure * Verify **Metric** and **Service name** match the incoming Cloud ALM metric * Verify **Attributes Filter** is correct, if used. A filter with names only, as loaded, does not restrict the values * Verify threshold syntax is correct ([[products:promonitor:latest:monitorsguide:commonsettings|Multi thresholds syntax]]). A row with an invalid threshold is reported in the monitor log and skipped * **Expected status alarm not raised**: Some Cloud ALM metrics expose multiple datapoints for the same measure. The monitor evaluates matching datapoints for the selected alarm measure, but only raises alarms for matching data that passes the row filters. * **No metrics stored when data exists**: If **Send all metrics** is disabled, make sure **Send metric** is enabled on the relevant row. * **Stale data after adding new services in Cloud ALM**: Click **Load ALM Health Metrics** again or wait for the next scheduled run.