Skip to main content

GPU Monitoring & Metrics for MLOps

GPU Hardware Circuitry

GPU Infrastructure Monitoring for MLOps: Essential nvidia-smi Commands & Metrics

When running large-scale AI workloads, memory leaks, thermal throttling, and runaway PyTorch processes can crash inference nodes without warning. System administrators managing AI infrastructure need real-time visibility into GPU VRAM utilization, power draw, and process ownership.


1. Advanced nvidia-smi Query Commands

Instead of running standard nvidia-smi, extract structured CSV data for custom log monitors or automated scripts:

# Stream GPU VRAM usage, temperature, and power consumption every 2 seconds
nvidia-smi --query-gpu=timestamp,name,temperature.gpu,utilization.gpu,utilization.memory,memory.used,memory.free --format=csv -l 2

2. Identifying Specific GPU Process Owners

To identify which Python or Docker container is consuming GPU VRAM on multi-tenant servers:

# List all active GPU compute processes with Process IDs (PIDs)
nvidia-smi pmon -s u -v

Once you locate a runaway process PID, inspect its systemd parent unit or Docker container:

# Inspect process details by PID
ps -up [PID]

# Find associated Docker container ID
cat /proc/[PID]/cgroup | grep docker

3. Setting GPU Power Limits for Stability

To prevent server power supplies from tripping during sudden LLM batch inference spikes, cap the maximum wattage allowed per GPU card:

# Enable persistence mode across reboots
sudo nvidia-smi -pm 1

# Cap maximum GPU power limit to 250 Watts (e.g., for NVIDIA RTX 3090 / A4000)
sudo nvidia-smi -pl 250

Comments

Popular posts from this blog

Production Log Management with

Production Log Management on Linux: Mastering logrotate for High-Uptime Servers Unmanaged log growth is one of the most common causes of unexpected production outages. When application or web server logs fill up host disk partitions, database transactions fail, services crash, and systems become un-writeable. Linux provides a built-in utility called logrotate to automatically compress, rotate, and archive old log files. 1. Understanding logrotate Structure System logrotate configurations are split between two main locations: /etc/logrotate.conf — Global baseline configurations. /etc/logrotate.d/ — Service-specific configuration files (e.g., Nginx, Docker, custom apps). 2. Creating a Custom Application Log Rule Create a rule file for a custom production web service at /etc/logrotate.d/jam-app : /var/log/jam-app/*.log { daily missingok rotate 14 compress delaycompress notifempty create 0640 www-data www-data sharedscripts ...

Building High-Performance AI Infrastructure on Linux: From NVIDIA Drivers to vLLM Production Deployments

  The modern Machine Learning landscape relies entirely on Linux infrastructure. While high-level framework developments like PyTorch, Hugging Face, and LangChain dominate developer discussions, the actual engine serving low-latency inference at scale sits directly on Linux kernels, GPU acceleration drivers, and containerized serving engines. In this guide, we will step through the core architecture required to turn a raw Linux server (Ubuntu/Debian or RHEL/Rocky) into a high-throughput, OpenAI-compatible AI inference endpoint using vLLM and Docker . 1. Prerequisites & Linux Kernel Preparation Before running any Large Language Model (LLM) inference engine, your host operating system must be properly configured with NVIDIA proprietary drivers and the NVIDIA Container Toolkit to allow Docker containers direct access to host GPU acceleration. Step 1.1: Verify GPU Hardware Ensure your host system correctly detects all installed NVIDIA GPUs: Bash # Check PCI bus for recognized NVI...

Automated PostgreSQL Backups to S3 via Bash & Cron

Automated PostgreSQL Backups: A Production Bash & Cron Blueprint Database backup failures are catastrophic when disaster strikes. Relying on manual database exports leads to data loss and unverified recovery points. Setting up an automated pipeline that dumps PostgreSQL tables, compresses the archive using gzip , and syncs the output to remote object storage guarantees data durability. In this guide, we will build an automated shell script for PostgreSQL backups and schedule it via system cron jobs. 1. Prerequisites & AWS CLI Setup First, ensure the AWS CLI tool and PostgreSQL client utilities are installed on your Linux host environment: sudo apt update && sudo apt install -y postgresql-client awscli gzip Configure credentials for an IAM user with Write permissions to your S3 bucket: aws configure 2. Writing the Production Backup Script Create a backup script at /usr/local/bin/pg_backup.sh : sudo nano /usr/local/bin/pg_backup.sh Paste the foll...