Rahul
Veerapur
I help startups build AI, scale data infrastructure, and grow revenue through data-driven decisions.
I help startups and businesses build AI systems, scale data infrastructure, and make better decisions with data. My work spans data engineering, machine learning, analytics, and AI, where I design end-to-end solutions ranging from real-time data pipelines and predictive models to production-ready LLM applications and autonomous AI agents.
I enjoy solving problems that sit at the intersection of software engineering and data science. Whether it's architecting distributed data platforms, deploying machine learning models, or building intelligent systems that automate complex workflows, I'm driven by creating solutions that are scalable, reliable, and built to make a real impact.
Outside of work, I love exploring new places and experiencing different cultures, especially through food. Whether it's discovering a hidden local restaurant, planning my next trip, or trying cuisines I've never had before, I'm always looking for new experiences. Whenever I get the chance, you'll also find me snowboarding, listening to music, gaming with friends, or getting lost in a good book.
My Projects
From natural-language data warehouses to real-time streaming pipelines, every build here solves a real data or AI problem end to end.
Experience
- Multi-Source Integration: Ingested IQVIA NPA and FDA Orange Book data into a drug-level panel dataset spanning 50+ branded and generic compounds for longitudinal prescribing analytics.
- Hybrid SEM + Neural Network: Architected an end-to-end pipeline across 5+ years of cardiovascular drug market data, enabling simultaneous causal pathway estimation and non-linear interaction modeling.
- Prescriber Segmentation: Increased targeting accuracy by 88% via t-SNE/SVD behavioral features enabling demographic and clinical risk stratification.
- Time-Series Forecasting: Achieved 92% R² with a custom LSTM model, outperforming Random Forest baselines on non-linear seasonality in prescription demand.
- Automated ML Pipeline: Built an Airflow + Google Cloud Functions pipeline ingesting Target APIs and Google Trends data for daily stockout probability forecasts with automated retraining triggers.
- Continuous Training: Deployed a CT workflow with XGBoost + PyTorch managed via MLflow, triggered by concept drift detection, sustaining F1-score above 0.90 across quarterly inventory cycles.
- NL-to-SQL Agent: Built a tool-augmented LLM agent with schema-aware SQL generation, automatic repair loops, and validation against production BI views for inventory analysts.
- ETL at Scale: Consolidated 7M+ land parcel records from 15+ sources into Snowflake via Airflow pipelines with fuzzy-matching entity resolution, catching 12% of malformed records before production.
- Predictive Site Selection: Gradient boosting model with 25+ geospatial features achieved 0.81 AUC, reducing initial screening effort by 70%.
- Demand Forecasting: 12 to 24 month regional occupancy and rental yield forecasts (ARIMA/Prophet) correctly predicted 3 of 4 regional demand shifts in FY2023.
- Tenant Churn Prediction: XGBoost churn model on 350+ assets achieved 82% precision, enabling proactive retention that reduced turnover by 18%.
- Asset Segmentation: K-Means/DBSCAN clustering on 900+ logistics assets identified 5 distinct profiles, reducing appraisal variance by 25%.
- Federated Metadata Search: Reduced data discovery time by 80% across teams by enabling search across 50K tables in 1,200 databases via a centralized metadata index.
- Usage Analytics Dashboard: Lowered redundant exploratory queries by 25% by giving teams visibility into dataset usage patterns, connection times, and peak access windows.
- Dataset Recommendations: Cut average employee search time by 30% through relevant-dataset recommendations based on historical query behavior.
Education
Let's Talk
Got a question, an opening, or a project that needs a data or AI engineer? I'd like to hear about it.