Skip to main content

langchain easy code

Building production-ready AI agents using LangChain requires robust error handling, provider redundancy, and seamless tool execution. In this guide, we will explore how to set up an advanced LangChain agent with fallback models using ChatGoogleGenerativeAI, Sarvam AI, and LiteLLM to ensure high availability and zero downtime.

Why Implement Model Fallbacks in LangChain Agents?

When deploying AI automation workflows in production, relying on a single Large Language Model (LLM) introduces single-point-of-failure risks—such as rate limits (e.g., token per minute caps), transient network errors, or sudden API outages. By implementing middleware-based fallbacks, your application can automatically switch to secondary providers like Llama 3 or Sarvam-105B without interrupting user experience.

Prerequisites and Environment Setup

Before writing your agent code, make sure you have the necessary API keys configured as environment variables. This ensures your code remains secure and follows best practices for credential management.

import os

# Configure your primary and fallback API keys securely
os.environ["GOOGLE_API_KEY"]     = "your-google-key"
os.environ["SARVAM_API_KEY"]     = "your-sarvam-key"
os.environ["GROQ_API_KEY"]       = "your-groq-key"
os.environ["OPENROUTER_API_KEY"] = "your-openrouter-key"   # Optional
os.environ["TOGETHERAI_API_KEY"] = "your-together-key"     # Optional

Initializing the Primary Gemini Model

For your primary execution engine, Google's Gemini models offer native tool-calling capabilities and generous free tiers. Here is how to initialize gemini-2.5-flash with zero temperature for deterministic outputs:

from langchain_google_genai import ChatGoogleGenerativeAI

primary = ChatGoogleGenerativeAI(
    model="gemini-2.5-flash",        # Use "gemini-2.5-flash-lite" for higher RPM
    temperature=0,
    max_tokens=4096,
    max_retries=2,
)

Setting Up the Fallback Chain with LiteLLM and Sarvam

Next, define your fallback array using Sarvam AI for regional optimization and LiteLLM-wrapped providers for multi-provider resilience:

from langchain_sarvam import ChatSarvam
from langchain_community.chat_models import ChatLiteLLM

fallback_models = [
    ChatSarvam(model="sarvam-105b", temperature=0, max_tokens=4096),
    ChatLiteLLM(model="groq/llama-3.3-70b-versatile", temperature=0, max_tokens=4096),
    ChatLiteLLM(model="openrouter/meta-llama/llama-3.3-70b-instruct:free", temperature=0, max_tokens=4096),
    ChatLiteLLM(model="together_ai/meta-llama/Llama-3.3-70B-Instruct-Turbo-Free", temperature=0, max_tokens=4096),
]

Assembling and Invoking the Deep Agent

Finally, combine your primary model, tools, and fallback middleware into a unified agent workflow:

from langchain.agents.middleware import ModelFallbackMiddleware, ModelRetryMiddleware
from deepagents import create_deep_agent

agent = create_deep_agent(
    model=primary,
    tools=[web_search],
    system_prompt=system_prompt,
    middleware=[
        ModelRetryMiddleware(max_retries=2),        # Handle transient network errors first
        ModelFallbackMiddleware(*fallback_models),  # Switch to backups if limits are hit
    ],
)

# Execute user tasks
result = agent.invoke({"messages": [("user", str(tasks))]})

Conclusion

By combining LangChain middleware, retry mechanisms, and multi-provider failovers, you build resilient AI systems capable of handling production workloads effortlessly. Try integrating this architecture into your next LLM project to maximize uptime!

Comments