What did AI cost you this quarter?
Reporting by Red Hat Developer BlogRead the original at developers.redhat.com
Executive Summary
A cost attribution challenge arises when AI usage is aggregated into a single line item, making it difficult for teams to understand their consumption and drive efficiency. The scenario involves multiple departments sharing GPU infrastructure for various AI tasks, leading to an opaque billing structure where the total cost is assigned to a central platform team without per-department breakdown. This lack of visibility prevents teams from optimizing usage, as there is no direct feedback loop connecting usage to cost generation.
The provided solution involves implementing a system built on Red Hat OpenShift AI that provides granular tracking through several integrated components: Models-as-a-Service (MaaS) for identity boundaries, vLLM/llm-d for serving models and reporting token counts to Prometheus, and the cluster observability operator (COO) for automatic metric collection. This is supplemented by Perses dashboards for macro-level visualization and MLflow for detailed per-inference tracing that links usage directly to specific departmental activities and costs.
The ultimate goal of this system is to shift cost attribution from a lump sum to actionable data, allowing organizations to identify inefficiencies—such as underutilized resources or inefficient model choices—and enable teams to proactively optimize their workloads rather than simply facing a bill.
Facts Only
* Three departments share a GPU cluster for AI workloads.
* Engineering runs code reviews using a 27-billion-parameter model every 5 seconds.
* Marketing generates campaign copy on the same model every 12 seconds.
* Support handles customer questions on a smaller, cheaper model every 20 seconds.
* A total cost of $14,200 was incurred for the quarter, but individual departmental spending is unknown.
* A demo showed Marketing utilized only 32% of GPU capacity, resulting in potential waste.
* The system uses Models-as-a-Service (MaaS) to create per-department API keys and identity boundaries.
* vLLM or llm-d serve models and report token counts to Prometheus.
* Prometheus collects token consumption metrics via the cluster observability operator without requiring custom instrumentation.
* Perses dashboards visualize usage in the OpenShift Console and AI console, allowing filtering by subscription (department) or model.
* MLflow traces every inference request, capturing department, prompt details, tokens, latency, and cost.
* A specific trace example showed an agent span for engineering with token and cost details for a chat model interaction.
* Cost is calculated based on defined per-model input/output token rates.
* The breakdown example showed Engineering spending $9,120 (60%), Marketing $3,550 (25%), and Support $1,530 (15%).
Full Take
The narrative pivots on the discovery that opaque cost aggregation silences operational problems. The shift from monitoring total spend to attributing usage based on workload context unlocks agency. The core pattern exposed is the asymmetry of information: platform teams possess granular performance data while business units operate under aggregate billing, creating a structural bottleneck where technical realities are obscured by financial abstraction.
The discussion around showback demonstrates a critical tension between control and trust. Framing cost monitoring as surveillance often triggers resistance; the success in addressing this came from reframing it as supportive discovery—helping teams resolve performance bottlenecks rather than police behavior. This suggests that visibility, when delivered through an invitation to solve a recognized problem (like inefficient prompting), is more effective than punitive measurement.
The system architecture itself reflects a layered approach to accountability: MaaS establishes the identity layer, Prometheus gathers raw telemetry, and MLflow contextualizes those metrics into application-level traces. The true insight lies in the cost model's granularity: distinguishing between input and output tokens across different models (e.g., high token-heavy code reviews versus short customer queries) provides the necessary context to shift decision-making from simple expenditure awareness to intelligent resource allocation. The implication is that organizational efficiency is fundamentally blocked by a lack of fine-grained, contextual attribution, requiring a system capable of weaving together infrastructure metrics, application traces, and defined pricing into a single, actionable story for leadership.
Bridge Questions: If the cost model definition (input/output split) were determined entirely by workload type rather than fixed rates, how would this change the incentives for model selection? What are the long-term organizational safeguards needed to ensure that visibility is used for optimization and not merely for policing resource consumption? How can organizations institutionalize the habit of using trace data from MLflow as a mandatory prerequisite for any cost review cycle?
From the original · Red Hat Developer Blog
"What did AI cost us this quarter?" That question lands in every platform team's lap eventually, and most teams can't answer it.Read the full story at developers.redhat.com
Sentinel — provisional
No strong signs of machine writing were found in the source article. Provisional estimate, not a finding that a person wrote it.
The text reads like an experienced technical consultant explaining a complex system, demonstrating deep domain knowledge, structured argumentation, and practical application rather than purely synthetic generation.
This looks only at the wording of the original source article, not at this page's AI-written sections. A small local AI model made this estimate. It has not been checked against known human and machine texts, so treat it as provisional. It cannot show who wrote an article.
