AI Governance Specialist

Date: Sep 3, 2026

Location: NAVI MUMBAI, IN

Company: icicisecur

Job Title: AI Model Governance Specialist

Department: Risk Management

Location: Mumbai

Role Overview

We require a seasoned technical person with __ years of experience and good understanding of the AI Model who would assist us in smooth implementation of the AI models into our systems. The person will oversee the risk management, validation, and control lifecycle for Generative AI and Large Language Model (LLM) deployments across enterprise systems. This role bridges technical engineering with governance to ensure AI models—specifically Retrieval-Augmented Generation (RAG) architectures—are safe, compliant, explainable, and reliable.

Key Responsibilities

  • RAG & GenAI Risk Assessment: Audit end-to-end RAG pipelines (knowledge base ingestion, chunking/embedding, vector retrieval, context grounding, and generation) to identify risk vectors such as leakage, prompt injection, context contamination, and retrieval failures.
  • LLM Evaluation & Benchmarking: Design and execute testing frameworks for hallucination rates, toxicity, robustness, model calibration, and fairness/bias.
  • Golden Dataset & Test Design: Curate reference-answer evaluation datasets and implement LLM-as-a-Judge validation pipelines.
  • Governance Frameworks: Establish risk mitigation controls covering data pre-processing, pre-deployment validation gates, and post-deployment observability.
  • Technical Auditing: Independently inspect, evaluate, and validate model outputs

Required Qualifications & Technical Skills

  • Education: B.Tech / M.Tech in Computer Science, Quantitative disciplines, or MCA.
  • Domain Knowledge: Strong quantitative/CS foundation with a thorough understanding of modern AI/ML taxonomy and GenAI workflows.
  • RAG Expertise: Deep understanding of risks associated with vector search, embeddings, relevance scoring, and grounding boundaries.
  • Evaluation Methodologies: Hands-on familiarity with reference-based testing, LLM-as-judge frameworks, red-teaming fundamentals, and semantic similarity evaluations.
  • Platform Awareness: Familiarity with enterprise LLM orchestration and deployment suites (e.g., Amazon Bedrock, Azure OpenAI Service, Vertex AI,etc).

Preferred / Good-to-Have Skills

  • Scripting & Analysis: Working knowledge of Python to query APIs, parse JSON outputs, and independently audit model evaluation scripts.
  • Traditional & Neural NLP Metrics: Practical experience evaluating ROUGE, BLEU, BERTScore, and semantic distance metrics.
  • Control Architecture: Proven ability to construct enterprise AI control frameworks spanning data provenance, model sign-off, rate-limiting, and continuous drift monitoring.
  • Evaluation Tooling: Familiarity with the underlying algorithmic mechanics of off-the-shelf hallucination detectors and automated accuracy engines (e.g., Ragas, TruLens, DeepEval).