B.Tech Computer Eng
Started Computer Engineering B.Tech at Vishwakarma University
Hi, I'm Sumit Shinde. I'm a Computer Engineering student specializing in building end-to-end Machine Learning pipelines, integrating Generative AI/LLM applications, and automating cloud-native deployments.
Bridging machine learning engineering and cloud deployments
I specialize in developing end-to-end Machine Learning workflows, integrating Generative AI solutions, and setting up automated CI/CD pipelines. Currently a Computer Engineering B.Tech student at Vishwakarma University, I focus on transforming data into actionable insights and robust services.
My experience ranges from building Chrome extensions with LightGBM and RAG capabilities to developing NYC taxi demand forecasting systems using EWMA features. I enjoy containerizing applications with Docker and deploying scalable cloud infrastructures on AWS.
RAG, LangChain, LangGraph, Vector DBs, context engineering
Supervised learning, deep learning, NLP, XGBoost, LightGBM
AWS (EC2, S3, Lambda), Docker, GitHub Actions, CI/CD
ETL, Apache Airflow, PostgreSQL, PySpark, Snowflake
Started Computer Engineering B.Tech at Vishwakarma University
Joined Shorat Innovations, performing EDA and dashboard design
Earned the IBM Full Stack Developer Certification from Coursera
Certified in developing AI applications with Python and Flask
“Building automated, robust, and reproducible machine learning workflows to bridge the gap between code and cloud.”
Technical expertise and tools that I work with
Pandas, NumPy, scikit-learn, PyTorch, LightGBM, XGBoost, and statistical modeling.
Retrieval-Augmented Generation (RAG), LangChain, LangGraph, LlamaIndex, OpenAI API, and FAISS.
AWS (EC2, S3, ECR, Lambda, Auto Scaling, Load Balancers) and Google Cloud Platform.
ETL pipelines, Apache Airflow, PySpark, DBT, Snowflake, Airbyte, and PostgreSQL.
GitHub Actions, Jenkins, PyTest for automated testing, Prometheus, Grafana, and API development.
Docker, Kubernetes, MLflow for experiment tracking, DVC for data versioning, and model deployment.
I specialize in designing pipelines that bridge the gap between AI code and scalable production deployments. I use tools like DVC for data versioning, MLflow for tracking experiments, and Docker/AWS for model serving, ensuring robust and reproducible workflows.
Professional Experience and Academic Journey
Industrial training and practical application of data science techniques.
Academic profile and professional technical courses completed.
Completed Diploma in Computer Engineering with 75.66% in 2024.
Currently pursuing B.Tech in Computer Engineering with a CGPA of 7.64. Core coursework includes Algorithms, Database Systems, Artificial Intelligence, and Software Engineering.
Specialized training in building RESTful APIs using Flask and integrating OpenAI models, LLMs, and prompt engineering methods to deliver functional AI endpoints.
Comprehensive curriculum covering cloud development, Git/GitHub, HTML/CSS/JavaScript, containerization with Docker, Kubernetes deployments, and CI/CD pipelines.
Detailed mapping of languages, libraries, and frameworks in my toolkit.
Featured projects highlighting machine learning, NLP, and system deployments
AI/ML Chrome Extension with MLOps pipeline & Flask backend deployed on AWS. Features a LightGBM sentiment model (87% accuracy), T5-small video summarizer, and FAISS RAG chatbot.
End-to-End time-series forecasting system (3M+ trips) using EWMA features and Mini-Batch KMeans. Implemented a Linear Regression model (92.1% accuracy) and Dockerized Streamlit app deployed on AWS via GitHub Actions.
Project feedback and evaluations from academic and internship mentors
Core fields of application and engineering services I offer
Developing predictive models, regression, classification, clustering, and neural networks using scikit-learn, PyTorch, LightGBM, and XGBoost.
Designing context-aware LLM agents, Retrieval-Augmented Generation workflows, semantic search indexes with FAISS, and LangChain applications.
Building CI/CD pipelines via GitHub Actions, versioning datasets with DVC, tracking experiments with MLflow, and deploying via Docker/ECR.
Engineering ETL data pipelines, processing large datasets with PySpark, automating workflows via Airflow, and querying SQL/Snowflake databases.
Creating robust, high-performance REST APIs in Flask and FastAPI to serve models, integrate webhooks, and interface with frontends.
Architecting cloud-native solutions on AWS using EC2, S3, Lambda, Load Balancers, and Auto Scaling groups for maximum scalability.
Common questions about my technical approach, tools, and experience
I specialize in bridging the gap between Machine Learning models and production environments. My interests lie in AI/ML Engineering, Generative AI (building RAG applications and LLM agents), and setting up robust MLOps/CI/CD pipelines.
My primary language is Python, along with SQL and C++. For machine learning, I work extensively with scikit-learn, PyTorch, LightGBM, and XGBoost. For Generative AI and LLM orchestration, I use LangChain, LangGraph, and LlamaIndex.
I use DVC (Data Version Control) to version datasets and pipeline outputs (storing files securely in AWS S3), and MLflow to log experiment metrics, models, and parameters, ensuring complete reproducibility in my workflows.
I deploy containerized applications using Docker and Kubernetes. On AWS, I set up scalable infrastructure using EC2, ECR for container image storage, Load Balancers, CodeDeploy, and Auto Scaling groups, all automated via GitHub Actions.
Yes! I am currently in my third year of B.Tech in Computer Engineering at Vishwakarma University and actively seeking AI/ML, Data Science, or MLOps engineering internship opportunities. I am ready to relocate or work remotely.
Get in touch to discuss internships, collaborations, or project opportunities
Have an interesting project or internship position? Fill out the form or reach out directly via email or phone. I'd love to connect!