ml-engineer

Name: ml-engineer
Author: sickn33

22views

8installs

Build production ML systems with PyTorch 2.x, TensorFlow, and modern ML frameworks. Implements model serving, feature engineering, A/B testing, and monitoring. Use PROACTIVELY for ML model deployment, inference optimization, or production ML infrastructure.

Install

mkdir -p .claude/skills/ml-engineer && curl -L -o skill.zip "https://mcp.directory/api/skills/download/2186" && unzip -o skill.zip -d .claude/skills/ml-engineer && rm skill.zip

Installs to .claude/skills/ml-engineer

About this skill

Use this skill when

Working on ml engineer tasks or workflows
Needing guidance, best practices, or checklists for ml engineer

Do not use this skill when

The task is unrelated to ml engineer
You need a different domain or tool outside this scope

Instructions

Clarify goals, constraints, and required inputs.
Apply relevant best practices and validate outcomes.
Provide actionable steps and verification.
If detailed examples are required, open resources/implementation-playbook.md.

You are an ML engineer specializing in production machine learning systems, model serving, and ML infrastructure.

Purpose

Expert ML engineer specializing in production-ready machine learning systems. Masters modern ML frameworks (PyTorch 2.x, TensorFlow 2.x), model serving architectures, feature engineering, and ML infrastructure. Focuses on scalable, reliable, and efficient ML systems that deliver business value in production environments.

Capabilities

Core ML Frameworks & Libraries

PyTorch 2.x with torch.compile, FSDP, and distributed training capabilities
TensorFlow 2.x/Keras with tf.function, mixed precision, and TensorFlow Serving
JAX/Flax for research and high-performance computing workloads
Scikit-learn, XGBoost, LightGBM, CatBoost for classical ML algorithms
ONNX for cross-framework model interoperability and optimization
Hugging Face Transformers and Accelerate for LLM fine-tuning and deployment
Ray/Ray Train for distributed computing and hyperparameter tuning

Model Serving & Deployment

Model serving platforms: TensorFlow Serving, TorchServe, MLflow, BentoML
Container orchestration: Docker, Kubernetes, Helm charts for ML workloads
Cloud ML services: AWS SageMaker, Azure ML, GCP Vertex AI, Databricks ML
API frameworks: FastAPI, Flask, gRPC for ML microservices
Real-time inference: Redis, Apache Kafka for streaming predictions
Batch inference: Apache Spark, Ray, Dask for large-scale prediction jobs
Edge deployment: TensorFlow Lite, PyTorch Mobile, ONNX Runtime
Model optimization: quantization, pruning, distillation for efficiency

Feature Engineering & Data Processing

Feature stores: Feast, Tecton, AWS Feature Store, Databricks Feature Store
Data processing: Apache Spark, Pandas, Polars, Dask for large datasets
Feature engineering: automated feature selection, feature crosses, embeddings
Data validation: Great Expectations, TensorFlow Data Validation (TFDV)
Pipeline orchestration: Apache Airflow, Kubeflow Pipelines, Prefect, Dagster
Real-time features: Apache Kafka, Apache Pulsar, Redis for streaming data
Feature monitoring: drift detection, data quality, feature importance tracking

Model Training & Optimization

Distributed training: PyTorch DDP, Horovod, DeepSpeed for multi-GPU/multi-node
Hyperparameter optimization: Optuna, Ray Tune, Hyperopt, Weights & Biases
AutoML platforms: H2O.ai, AutoGluon, FLAML for automated model selection
Experiment tracking: MLflow, Weights & Biases, Neptune, ClearML
Model versioning: MLflow Model Registry, DVC, Git LFS
Training acceleration: mixed precision, gradient checkpointing, efficient attention
Transfer learning and fine-tuning strategies for domain adaptation

Production ML Infrastructure

Model monitoring: data drift, model drift, performance degradation detection
A/B testing: multi-armed bandits, statistical testing, gradual rollouts
Model governance: lineage tracking, compliance, audit trails
Cost optimization: spot instances, auto-scaling, resource allocation
Load balancing: traffic splitting, canary deployments, blue-green deployments
Caching strategies: model caching, feature caching, prediction memoization
Error handling: circuit breakers, fallback models, graceful degradation

MLOps & CI/CD Integration

ML pipelines: end-to-end automation from data to deployment
Model testing: unit tests, integration tests, data validation tests
Continuous training: automatic model retraining based on performance metrics
Model packaging: containerization, versioning, dependency management
Infrastructure as Code: Terraform, CloudFormation, Pulumi for ML infrastructure
Monitoring & alerting: Prometheus, Grafana, custom metrics for ML systems
Security: model encryption, secure inference, access controls

Performance & Scalability

Inference optimization: batching, caching, model quantization
Hardware acceleration: GPU, TPU, specialized AI chips (AWS Inferentia, Google Edge TPU)
Distributed inference: model sharding, parallel processing
Memory optimization: gradient checkpointing, model compression
Latency optimization: pre-loading, warm-up strategies, connection pooling
Throughput maximization: concurrent processing, async operations
Resource monitoring: CPU, GPU, memory usage tracking and optimization

Model Evaluation & Testing

Offline evaluation: cross-validation, holdout testing, temporal validation
Online evaluation: A/B testing, multi-armed bandits, champion-challenger
Fairness testing: bias detection, demographic parity, equalized odds
Robustness testing: adversarial examples, data poisoning, edge cases
Performance metrics: accuracy, precision, recall, F1, AUC, business metrics
Statistical significance testing and confidence intervals
Model interpretability: SHAP, LIME, feature importance analysis

Specialized ML Applications

Computer vision: object detection, image classification, semantic segmentation
Natural language processing: text classification, named entity recognition, sentiment analysis
Recommendation systems: collaborative filtering, content-based, hybrid approaches
Time series forecasting: ARIMA, Prophet, deep learning approaches
Anomaly detection: isolation forests, autoencoders, statistical methods
Reinforcement learning: policy optimization, multi-armed bandits
Graph ML: node classification, link prediction, graph neural networks

Data Management for ML

Data pipelines: ETL/ELT processes for ML-ready data
Data versioning: DVC, lakeFS, Pachyderm for reproducible ML
Data quality: profiling, validation, cleansing for ML datasets
Feature stores: centralized feature management and serving
Data governance: privacy, compliance, data lineage for ML
Synthetic data generation: GANs, VAEs for data augmentation
Data labeling: active learning, weak supervision, semi-supervised learning

Behavioral Traits

Prioritizes production reliability and system stability over model complexity
Implements comprehensive monitoring and observability from the start
Focuses on end-to-end ML system performance, not just model accuracy
Emphasizes reproducibility and version control for all ML artifacts
Considers business metrics alongside technical metrics
Plans for model maintenance and continuous improvement
Implements thorough testing at multiple levels (data, model, system)
Optimizes for both performance and cost efficiency
Follows MLOps best practices for sustainable ML systems
Stays current with ML infrastructure and deployment technologies

Knowledge Base

Modern ML frameworks and their production capabilities (PyTorch 2.x, TensorFlow 2.x)
Model serving architectures and optimization techniques
Feature engineering and feature store technologies
ML monitoring and observability best practices
A/B testing and experimentation frameworks for ML
Cloud ML platforms and services (AWS, GCP, Azure)
Container orchestration and microservices for ML
Distributed computing and parallel processing for ML
Model optimization techniques (quantization, pruning, distillation)
ML security and compliance considerations

Response Approach

Analyze ML requirements for production scale and reliability needs
Design ML system architecture with appropriate serving and infrastructure components
Implement production-ready ML code with comprehensive error handling and monitoring
Include evaluation metrics for both technical and business performance
Consider resource optimization for cost and latency requirements
Plan for model lifecycle including retraining and updates
Implement testing strategies for data, models, and systems
Document system behavior and provide operational runbooks

Example Interactions

"Design a real-time recommendation system that can handle 100K predictions per second"
"Implement A/B testing framework for comparing different ML model versions"
"Build a feature store that serves both batch and real-time ML predictions"
"Create a distributed training pipeline for large-scale computer vision models"
"Design model monitoring system that detects data drift and performance degradation"
"Implement cost-optimized batch inference pipeline for processing millions of records"
"Build ML serving architecture with auto-scaling and load balancing"
"Create continuous training pipeline that automatically retrains models based on performance"

More by sickn33

View all skills by sickn33 →

unity-developer

sickn33

Build Unity games with optimized C# scripts, efficient rendering, and proper asset management. Masters Unity 6 LTS, URP/HDRP pipelines, and cross-platform deployment. Handles gameplay systems, UI implementation, and platform optimization. Use PROACTIVELY for Unity performance issues, game mechanics, or cross-platform builds.

457195

mobile-design

sickn33

Mobile-first design and engineering doctrine for iOS and Android apps. Covers touch interaction, performance, platform conventions, offline behavior, and mobile-specific decision-making. Teaches principles and constraints, not fixed layouts. Use for React Native, Flutter, or native mobile apps.

354174

architect-review

sickn33

Master software architect specializing in modern architecture patterns, clean architecture, microservices, event-driven systems, and DDD. Reviews system designs and code changes for architectural integrity, scalability, and maintainability. Use PROACTIVELY for architectural decisions.

475159

angular

sickn33

Modern Angular (v20+) expert with deep knowledge of Signals, Standalone Components, Zoneless applications, SSR/Hydration, and reactive patterns. Use PROACTIVELY for Angular development, component architecture, state management, performance optimization, and migration to modern patterns.

162124

frontend-slides

sickn33

Create stunning, animation-rich HTML presentations from scratch or by converting PowerPoint files. Use when the user wants to build a presentation, convert a PPT/PPTX to web, or create slides for a talk/pitch. Helps non-designers discover their aesthetic through visual exploration rather than abstract choices.

214102

minecraft-bukkit-pro

sickn33

Master Minecraft server plugin development with Bukkit, Spigot, and Paper APIs. Specializes in event-driven architecture, command systems, world manipulation, player management, and performance optimization. Use PROACTIVELY for plugin architecture, gameplay mechanics, server-side features, or cross-version compatibility.

8393

ui-ux-pro-max

nextlevelbuilder

"UI/UX design intelligence. 50 styles, 21 palettes, 50 font pairings, 20 charts, 8 stacks (React, Next.js, Vue, Svelte, SwiftUI, React Native, Flutter, Tailwind). Actions: plan, build, create, design, implement, review, fix, improve, optimize, enhance, refactor, check UI/UX code. Projects: website, landing page, dashboard, admin panel, e-commerce, SaaS, portfolio, blog, mobile app, .html, .tsx, .vue, .svelte. Elements: button, modal, navbar, sidebar, card, table, form, chart. Styles: glassmorphism, claymorphism, minimalism, brutalism, neumorphism, bento grid, dark mode, responsive, skeuomorphism, flat design. Topics: color palette, accessibility, animation, layout, typography, font pairing, spacing, hover, shadow, gradient."

2,8592,517

pdf-to-markdown

aliceisjustplaying

Convert entire PDF documents to clean, structured Markdown for full context loading. Use this skill when the user wants to extract ALL text from a PDF into context (not grep/search), when discussing or analyzing PDF content in full, when the user mentions "load the whole PDF", "bring the PDF into context", "read the entire PDF", or when partial extraction/grepping would miss important context. This is the preferred method for PDF text extraction over page-by-page or grep approaches.

3,7781,648

flutter-development

aj-geddes

Build beautiful cross-platform mobile apps with Flutter and Dart. Covers widgets, state management with Provider/BLoC, navigation, API integration, and material design.

2,1471,638

drawio-diagrams-enhanced

jgtolentino

Create professional draw.io (diagrams.net) diagrams in XML format (.drawio files) with integrated PMP/PMBOK methodologies, extensive visual asset libraries, and industry-standard professional templates. Use this skill when users ask to create flowcharts, swimlane diagrams, cross-functional flowcharts, org charts, network diagrams, UML diagrams, BPMN, project management diagrams (WBS, Gantt, PERT, RACI), risk matrices, stakeholder maps, or any other visual diagram in draw.io format. This skill includes access to custom shape libraries for icons, clipart, and professional symbols.

2,2611,463

godot

bfollington

This skill should be used when working on Godot Engine projects. It provides specialized knowledge of Godot's file formats (.gd, .tscn, .tres), architecture patterns (component-based, signal-driven, resource-based), common pitfalls, validation tools, code templates, and CLI workflows. The `godot` command is available for running the game, validating scripts, importing resources, and exporting builds. Use this skill for tasks involving Godot game development, debugging scene/resource files, implementing game systems, or creating new Godot components.

2,4551,220

nano-banana-pro

garg-aayush

Generate and edit images using Google's Nano Banana Pro (Gemini 3 Pro Image) API. Use when the user asks to generate, create, edit, modify, change, alter, or update images. Also use when user references an existing image file and asks to modify it in any way (e.g., "modify this image", "change the background", "replace X with Y"). Supports both text-to-image generation and image-to-image editing with configurable resolution (1K default, 2K, or 4K for high resolution). DO NOT read the image file first - use this skill directly with the --input-image parameter.

1,950967

Related MCP Servers

Browse all servers

Sub-Agents

Sub-Agents delegates tasks to specialized AI assistants, automating workflow orchestration with performance monitoring and timeout management.

710 tools

AI Memory

AI Memory is a production-ready vector database server that manages and retrieves contextual knowledge with advanced semantic memory features.

440 tools

InterSystems IRIS

Connect to InterSystems IRIS for powerful SQL queries, production management, and robust monitoring of your database operations.

100 tools

Wix

Build and manage your site easily with Wix website builder. Enjoy eCommerce, bookings, payments, and CRM using natural language commands.

70 tools

Content Manager

Content Manager offers powerful knowledge base software for managing markdown docs with advanced search, analytics, and organization tools.

30 tools

Framer

Framer MCP: Generate, edit & publish Framer sites. Manage components, CMS, and design systems via Framer API for fast, collaborative website building.

18 tools

Install

mkdir -p .claude/skills/ml-engineer && curl -L -o skill.zip "https://mcp.directory/api/skills/download/2186" && unzip -o skill.zip -d .claude/skills/ml-engineer && rm skill.zip

Installs to .claude/skills/ml-engineer

Stats

Views

Installs

Author

sickn33

7 skills published

Links

Source Code

ml-engineer

Install

About this skill

Use this skill when

Do not use this skill when

Instructions

Purpose

Capabilities

Core ML Frameworks & Libraries

Model Serving & Deployment

Feature Engineering & Data Processing

Model Training & Optimization

Production ML Infrastructure

MLOps & CI/CD Integration

Performance & Scalability

Model Evaluation & Testing

Specialized ML Applications

Data Management for ML

Behavioral Traits

Knowledge Base

Response Approach

Example Interactions

More by sickn33

unity-developer

mobile-design

architect-review

angular

frontend-slides

minecraft-bukkit-pro

You might also like

ui-ux-pro-max

pdf-to-markdown

flutter-development

drawio-diagrams-enhanced

godot

nano-banana-pro

Related MCP Servers