← back to case filesCASE FILE · 05 / 09
KubePulse product screenshot
product snapshot ↗
AIOPS · ML SYSTEMS

KubePulse

KubePulse watches Kubernetes pod metrics in real time and predicts memory-related failures before they cause an outage. A hybrid LSTM + LightGBM model trained on Prometheus metrics flags anomalies ahead of the breach, a custom Prometheus exporter feeds the prediction back into the cluster for autoscaling, and Gemini 1.5 Pro generates the exact kubectl remediation command once an anomaly is confirmed.

Stack

AI + ML
TensorFlowLightGBMLangChainGemini 1.5 Pro
Backend + Data
Python
Infrastructure
GoKubernetesPrometheusDocker

The Problem

Hybrid LSTM + LightGBM pipeline predicting Kubernetes pod failures before they happen, with Gemini-powered self-healing remediation.

What I Built

KubePulse watches Kubernetes pod metrics in real time and predicts memory-related failures before they cause an outage. A hybrid LSTM + LightGBM model trained on Prometheus metrics…

The Difference

Hybrid LSTM + LightGBM model on Prometheus metrics cut CrashLoopBackOff events 38% versus threshold-based alerting

Numbers That Matter

38%primary scale
1.5measured result
40%system signal
E2Equality outcome

Engineering Notes ✎

01

Engineering Note 1

Hybrid LSTM + LightGBM model on Prometheus metrics cut CrashLoopBackOff events 38% versus threshold-based alerting

02

Engineering Note 2

Custom Prometheus exporter surfaces a predicted-memory metric that drives HPA autoscaling directly

03

Engineering Note 3

Gemini 1.5 Pro generates targeted kubectl remediation commands on confirmed anomalies, cutting MTTR 40%

Architecture Overview

Pod telemetryAnomaly Engine
LSTM ModelLightGBM ModelGemini Remediator
PrometheusSelf-healing action