Image: web-assets.dd-static.net · rights & removal
Teaching a 9B model to investigate production alerts
Reporting by Datadog Security LabsRead the original at datadoghq.com
Executive Summary
Facts Only
* A deployment, feature flag, or configuration change can trigger an alert requiring investigation.
* Rules-based approaches identify potentially relevant changes quickly but have limited accuracy on complex incidents.
* Agentic investigations using frontier models are too expensive at scale for every alert.
* Qwen3.5-9B was fine-tuned on traces from investigations generated by GLM-5.3.
* The fine-tuned model achieved 87% of GLM-5.3’s recall with only 5% of the investigation cost ($0.003 vs $0.06).
* A single 40 GB A100 GPU can support approximately 100,000 investigations per week at the fine-tuned rate.
* The change attribution process involved using Bits Investigation to derive proxy labels for changes, which were used during fine-tuning.
* The training data involved generating 100 teacher traces from internal incidents and selecting examples based on matching proxy labels.
* The process involved supervised fine-tuning using 16-bit LoRA on the assistant responses in selected traces.
* Customer incident results for the fine-tuned model reached 0.62 Recall@5, compared to 0.52 for the base GLM-5.3.
Full Take
From the original · Datadog Security Labs
Scott Kramer Staff Engineer Junaid Ahmed Vice President, Engineering A deployment, feature flag, or configuration change may trigger an alert, but identifying which recent change most likely contributed to the alert can require a time-consuming investigation across multiple services and systems.Read the full story at datadoghq.com
Sentinel — provisional
No strong signs of machine writing were found in the source article. Provisional estimate, not a finding that a person wrote it.
This text reads like a detailed engineering whitepaper or blog post detailing a machine learning methodology applied to observability data, showing high specificity and internal methodological depth.
This looks only at the wording of the original source article, not at this page's AI-written sections. A small local AI model made this estimate. It has not been checked against known human and machine texts, so treat it as provisional. It cannot show who wrote an article.
