Start investigations from any bot's messages¶
For a team whose alerts reach a Slack channel through a bot that is not Alertmanager: at the end, a botMessage
trigger written in full reads that bot's posts, and each alert starts an agent investigation in its own thread.
Agent Kourier accepts the messages of one bot per trigger, named by its Slack bot ID, and reads them with the regexes
you give it. Only Alertmanager has a preset, and only Alertmanager is exercised in the sandbox and the end-to-end tests. For
another bot you write what the preset would fill in: match, extract, thread, limits and promptTemplate.
For Alertmanager, use Investigate Alertmanager alerts instead.
Before you start¶
- A Binding for the channel, and the Agent Kourier bot invited to it.
- The bot ID of the integration that posts the alerts (
B...). A dry run of the channel prints each message's sender. - A few of its real messages, to write the regexes against. Note which field holds what:
textis the message's own text;title,titleLink,fallback,footerandattachmentTextcome from its first legacy attachment.
Write the trigger¶
The example below is for a made-up message format, one line in text:
[FIRING] DiskAlmostFull on node-7: /var is 93% full
[RESOLVED] DiskAlmostFull on node-7: /var is 61% full
Replace the regexes with ones that fit your bot's messages.
chat:
connectionRef: {namespace: agent-kourier-system, name: slack}
channel: C0123456789
triggers:
- type: mention # (1)!
- type: botMessage
from:
botId: B0123456789 # (2)!
match:
text: '^\[(FIRING|RESOLVED)\] ' # (3)!
exclude:
text: '^\[[A-Z]+\] Heartbeat '
extract: # (4)!
state:
field: text
regex: '^\[(?P<state>FIRING|RESOLVED)\]'
alertname:
field: text
regex: '^\[[A-Z]+\] (?P<alertname>[A-Za-z0-9_.:-]+)'
key:
field: text
regex: '^\[[A-Z]+\] (?P<key>[A-Za-z0-9_.:-]+ on [A-Za-z0-9_.-]+)'
thread: # (5)!
by: key
cooldown: 4h
closeWhen:
state: RESOLVED
limits: # (6)!
rateLimit: {maxRuns: 5, per: 10m, maxConcurrent: 2}
maxRunsPerDay: 40
promptTemplate: | # (7)!
An alert is {{ .State }} in channel {{ .Channel }}: {{ .Fields.alertname }}.
Investigate it with your read-only tools. Report the likely cause, the evidence you
checked, and the next steps a person should take.
- Keep this if people should also be able to mention the bot. A list that names triggers is the whole list.
- The alert messages'
bot_id, never a display name. - Every
matchentry must match, and noexcludeentry may. Regexes are RE2, at most 512 bytes. - Each entry's regex has a named group with the entry's name, which supplies the value. A message in which an extract does not match does not match the trigger.
- Messages with the same
keyshare one thread. A repeat within the cooldown adds a note; a message whosestateisRESOLVEDcloses the thread with a note. Neither starts a turn. - Without a preset, a trigger with no
limitsis unlimited. These are the values the Alertmanager preset uses. - The first turn's prompt. Read extracted values through
.Fields: a value the template reads must be name-shaped (letters, digits,_.-:), and the key here holds spaces, so a template that reads.Keyis refused.
The Binding is validated when it loads: a regex that does not compile, an extract without its named group, a
thread.by: key without a key extract, or a template that reads a field the trigger does not extract rejects it.
Check it before going live¶
Replay the channel's history through the trigger with a dry run. Each of the bot's messages should
come out matched, with the state, alert name and key you expect, or unmatched with the regex that turned it away.
Tune the regexes until the lines read right, then load the Binding.
After it is live, agentkourier_trigger_matched_total and agentkourier_trigger_unmatched_total count the bot's
messages the trigger took and turned away, and agentkourier_trigger_decisions_total what each did to its thread.
Every field and decision is in Chat triggers.