Cosmos Operations Center — realtime, predictive, AI-driven
DBA dashboard जो आपको बताता है क्या ठीक करना है ग्राहक के call करने से पहले। Anomaly detection के साथ realtime charts, AI-narrated root cause walk में drill करने के लिए किसी भी spike पर click, saturation से पहले warn करने वाले predictive ETAs, automatically fire करने वाली configurable policies — और full recordable timeline ताकि आप 30 सेकंडों में कल के incident का post-mortem कर सकें।
वह आपको क्या दिखाता है, खोलते ही
4 realtime charts (1s tick)
Consumed RU/s · Throttle 429 events · Latency p99 · Error count। चारों Query Cost ring buffer से हर सेकंड update हो रहे हैं। Memory में Last hour।
लाल dots = anomalies, clickable
Anomaly detector throttle storms, RU saturation, latency spikes, partition skew को होते ही flag करता है। हर anomaly अपने exact timestamp पर pinned एक लाल dot है।
Predictive alerts banner
पिछले 3 मिनट पर Linear regression आपको बताती है "यदि trend जारी रहा तो RU ~14 मिनट में saturate होगा" — breach से पहले actionable, उसके बाद नहीं।
Recording toggle (history-store)
एक क्लिक और हर snapshot उसी MongoDB management instance पर persist होता है जिसे आपने mongostat/mongotop के लिए configure किया है। Job Manager इसे दूसरों के साथ दिखाता है।
हर लाल dot एक AI-narrated Root Cause Walk खोलता है
Anomaly पर क्लिक करें → एक side drawer तीन stops के साथ slide में आता है, हर एक deterministic context (Query Cost ring buffer + cross-scanner caches) + आपकी language में optional AI narration द्वारा backed।
क्या हुआ
सटीक timestamp, metric value, severity, उस मिनट चल रही top query shape, scanner cache से partition state। Deterministic — शून्य AI cost।
क्यों (सबसे probable cause)
AI structured context देखता है और single most probable root cause + data में इसे support करने वाला evidence लिखता है। DBA की language बोलता है। Signature द्वारा Cached।
कैसे ठीक करें
Quick mitigation (5min) + permanent fix (इस हफ्ते), हर एक scanner pane से linked जो apply command का स्वामी है। एक click और आप सही जगह पर हैं।
DBA per analysis AI model चुनता है
नियमित cost incident? gpt-4o-mini ($0.001) उपयोग करें। High-severity production outage? उसी incident के लिए Claude Sonnet ($0.018) पर switch करें। Provider picker drawer के अंदर है — cost upfront दिखाया जाता है, कोई surprises नहीं।
├ OpenAI gpt-4o-mini ($0.001)
├ Anthropic Claude Sonnet ($0.018) ✓
├ Google Gemini ($0.001)
├ Groq Llama ($0.002)
└ Ollama (local · निःशुल्क)
Incident signature × provider × locale (4h TTL) के अनुसार Results cached — समान incident दो बार clicked = शून्य re-charge।
Forecasts जो breach से पहले warn करते हैं
पिछले 3 मिनट की telemetry पर Linear regression — confidence के लिए R²-tagged, ETA द्वारा severity-scaled।
ru-saturation-eta~14 मिनटRU/s 8.2/s पर बढ़ रहा है — trend जारी रहा तो ~14 मिनट में saturate होगा
throttle-trending-upजारी हैThrottle events बढ़ रहे हैं — पिछले 5min में 23, rate +0.04/s²
partition-skew-growing~3.2 घंटेTop partition share बढ़ रहा है (अब 52%) — ~3.2 घंटों में 70% पार करेगा
storage-saturation-eta12 दिनStorage 78% पर — current growth पर 12 दिनों में 90% तक पहुँचने का extrapolated
Policies जो DBA define करता है — स्वचालित रूप से fire करती हैं
हर Cosmos connection की अपनी policies हैं। Triggers (threshold + duration + ns pattern) actions से जुड़े (in-app notify, Slack webhook, generic webhook, pre-stage scale-up, pre-stage path exclusion, AI analyze)।
Throttle storm — 5/min
जब throttle events > 5/min 2min तक sustained, in-app notification + +50% scale-up pre-stages (DBA approval आवश्यक)।
RU saturation — 85%
जब RU/s 5min के लिए observed ceiling के 85% पार करता है, suggested autoscale switch के साथ notify करता है।
Latency p99 — 500ms
जब p99 1min के लिए 500ms पार करता है, notify करता है + slow query path पर AI analysis trigger करता है।
Partition skew — top > 50%
जब एक partition > 50% docs रखती है, एक re-partition plan pre-stages करता है (approval आवश्यक)।
हर fire (policy × ns × 5min bucket) से deduped — कोई alert fatigue नहीं। एक क्लिक से 1h snooze। Audit log हर fire को metric value, trigger करने वाली policy, और queued action के साथ preserve करता है।
30 सेकंडों में कल के incident का Post-mortem
"Start recording" पर एक क्लिक → हर snapshot उसी MongoDB management connection पर persist होता है जिसे mongostat/mongotop पहले से उपयोग करते हैं। Job Manager खोलें → किसी भी "Rec Cosmos Ops" job पर क्लिक करें → unified Timeline scrubber आपको उसी window की MongoDB recordings के साथ cluster behavior replay करने देता है।
और कोई एक scrubber पर Cosmos + MongoDB को correlate नहीं करता।
क्यों यह Datadog / Grafana / Azure Monitor द्वारा copy नहीं किया जा सकता
| Capability | Operations Center | Datadog | Grafana | Azure Monitor |
|---|---|---|---|---|
| Realtime cluster charts | ✅ | ✅ | ✅ | ✅ |
| Spike पर click → root cause walk | ✅ | ⚠ केवल link | ❌ | ⚠ केवल link |
| आपकी language में AI narration | ✅ BYO LLM | ⚠ closed AI | ❌ | ⚠ closed AI |
| Per incident AI model चुनें | ✅ | ❌ | ❌ | ❌ |
| Apply command pre-staged + Shadow-validated | ✅ | ❌ | ❌ | ❌ |
| Cosmos $indexStats / GetPartitionStats correlated | ✅ | ❌ | ❌ | ⚠ |
| MongoDB monitoring के साथ Cross-correlate | ✅ | ❌ | ⚠ यदि दोनों ingested | ❌ |
| मूल्य | $99-499/mo | $15+/host/mo | self-host | प्रति GB ingested |
Templates खोजना बंद करें। Operations Center खोलें और देखें।
Cosmos DB के लिए NoSqlStudio आज़माना मुफ़्त है — कोई card नहीं, कोई signup नहीं। कोई भी Cosmos connection खोलें, Operations Center पर क्लिक करें, charts को जीवंत होते देखें।