Your Name
AI/ML Engineer
City, Country · email@example.com · linkedin.com/in/yourname · yourportfolio.com
Professional summary
AI/ML Engineer with 5 years building and deploying machine learning systems, from recommendation models to LLM-powered search. Shipped a retrieval-augmented support assistant that resolves 35% of tickets automatically, cut inference cost by 60% with quantization and batching, and runs training pipelines on Kubernetes. Skilled in PyTorch, Hugging Face, MLflow, and AWS SageMaker.
Skills
Python · Machine learning · PyTorch / TensorFlow · MLOps · Feature pipelines · Model serving · LLMs / prompt systems · Vector databases · SQL · Docker · Cloud ML · Experiment tracking
Experience
Senior AI/ML Engineer — Company Name
2021 – Present
- Built a retrieval-augmented support assistant with an open-weight LLM, pgvector, and a reranker that resolves 35% of tickets without an agent.
- Cut GPU inference cost by 60% by quantizing models to 8-bit and serving them with vLLM and dynamic batching on Kubernetes.
- Designed an offline evaluation set of 2,000 labeled queries, raising answer faithfulness from 78% to 91% before launch.
- Trained a two-tower recommendation model in PyTorch that increased click-through rate by 8% in an online A/B test.
- Automated retraining and deployment with Airflow, the MLflow model registry, and canary releases, moving model updates from quarterly to weekly.
Projects
Document Q&A with RAG | Python, Hugging Face, LangChain, Qdrant, FastAPI, Docker
- Chunked and embedded 5,000 PDF pages with a sentence-transformers model and stored the vectors in Qdrant.
- Added hybrid BM25 + vector search and a cross-encoder reranker, improving top-5 retrieval recall from 72% to 89%.
- Served the pipeline through a streaming FastAPI endpoint in Docker with p95 response time under 2.5 seconds.