🎙️ Daily Podcast (FR) : NViNiO•Podcast™
ADs | ✨ Enhance your Social Media content with NViNiO•AI™ for FREE
The collaboration between Iguazio (acquired by McKinsey) and NVIDIA empowers organizations to build production-grade AI solutions that are not only high-performing and scalable but also agile and ready for real-world deployment.
NVIDIA NIM microservices, critical to these capabilities, are designed to speed up generative AI deployment on any cloud or data center. Supporting a wide range of AI models, including NVIDIA AI foundation, community, and custom models, NIM microservices support seamless, scalable AI inference using industry-standard APIs.
At runtime, NIM selects the optimal inference engine for any combination of foundation model, GPU, and system. NIM containers also provide standard observability data feeds and built-in support for autoscaling with Kubernetes on NVIDIA GPUs.
MLRun is an open source AI orchestration framework that automates the entire AI pipeline, enabling NIM deployment in production environments. This includes all pipeline elements required for enterprise-grade, production-ready applications, including
Batch and real-time data pipelines CI/CD automation Auto-tracking of data lineage Experiment tracking, auto-logging, and model registry Distributed data processing Model training and serving Model monitoring Auto-scaling and serverless architectures for resource optimization SecurityTogether, MLRun and NVIDIA NIM offer a solution for deploying AI applications with optimized performance and orchestration capabilities.
What is MLRun?
MLRun is an open source AI orchestration framework designed to manage ML and generative AI applications throughout their lifecycle. Iguazio, now part of McKinsey & Company’s AI arm, QuantumBlack, built and maintains the open source framework.
It automates processes, such as data preparation, model tuning, customization, validation, and optimization for ML models, large language models (LLMs), and live AI applications over elastic resources. MLRun enables the rapid deployment of scalable real-time serving and application pipelines, providing built-in observability and flexible deployment options that support multi-cloud, hybrid, and on-premises environments.
Based on a variety of use case deployments, MLRun can enable:
Faster time to production Reduced computation costs End-to-end observabilityEnterprises use MLRun to develop AI models at scale in real time, deploy them across infrastructures, mitigate risks associated with generative AI, and drive AI-driven strategies safely and securely in any environment (multi-cloud, on-premises, or hybrid).
The framework is used for many use cases, including real-time agent copilots, call center analysis, chatbot automation, fraud prediction, real-time recommendation engines, and predictive maintenance.
Figure 1. Architecture showing NVIDIA NIM and NeMo microservices integrated in MLRunDeploying a multi-agent financial chatbot with MLRun and NVIDIA NIM
A large bank recently used MLRun to build a multi-agent chatbot that employs intent classification, real-time monitoring, and dynamic resource scaling. This use case demonstrates how financial institutions can deploy AI assistants with NVIDIA NIM inference efficiency and MLRun production-grade oversight, for operational effectiveness and regulatory adherence. In this demo, we show how the solution leverages MLRun to monitor NVIDIA NIM in real time.
The full chatbot architecture uses three distinct AI agents tailored to banking services. The loan agent handles mortgage and loan-related inquiries, such as explaining interest rates for specific mortgage terms. The investment agent provides personalized portfolio advice, analyzing scenarios like renewable energy stock investments. The general agent manages routine customer service tasks, including password resets or transaction history requests, while also directing complex queries to appropriate specialists. These agents operate through an LLM-powered query classification system that routes requests based on intent, with session logging for compliance and a modular design for independent updates without disrupting the entire system.
For quality control, the implementation uses an LLM-as-a-Judge mechanism to monitor interactions in real time. This evaluator validates routing decisions by assessing query-agent relevance, response accuracy, and regulatory compliance. It logs conversations for auditing and fine-tuning purposes while generating performance metrics such as misclassification rates, response quality scores, and compliance violation counts. MLRun operationalizes this monitoring through automated evaluation pipelines, dashboards displaying real-time metrics, and alert systems triggered by critical errors like regulatory breaches.
The success of this solution lies in its ability to integrate advanced AI technologies with operational simplicity. By leveraging NVIDIA NIM containers and combining them with the MLRun orchestration framework, the platform ensures that AI models are both performant and efficient.
Here’s why it works:
Serverless: MLRun simplifies NIM deployment by wrapping the instance in a serverless function and configuring the function with a NIM container image. This supports on-demand scaling for elasticity, monitoring, security, and other aspects of operationalization. Users can deploy LLM NIM microservices as serverless functions with a single click. LLM gateway: A unified interface makes switching between LLMs fast and intuitive. The gateway enables monitoring at different levels: a specific use case, a specific model, a general LLM provider, and a higher level for general usage monitoring such as latency, throughput, and memory. All is done by using labels.
Figure 2. Architectural diagram showcasing MLRun components and how they connect to a variety of use casesConclusion
MLRun and NVIDIA NIM form a powerful synergy for enterprise AI deployment, combining optimized inference with robust operational oversight. NVIDIA NIM delivers GPU-accelerated, containerized microservices for high-performance model execution across environments, while MLRun provides automated orchestration, secure API management, real-time monitoring, and much more. Together, they address critical production challenges, enabling enterprises to deploy scalable AI assistants with cutting-edge capabilities and operational reliability.
To keep going, experiment with MLRun and NIM, and learn more about deployment and model monitoring capabilities in MLRun. See the live demo and further technical explanation in the recording of Iguazio’s MLOps Live series.
To learn more about how NVIDIA supports AI startups, visit the Inception program webpage.
.png)
1 year ago
English (United States) ·
French (France) ·