Prabhnoor Kaur

I like solving problems using data.

Data Science student at UBC.

By the numbers

10 projects built
$139K+ revenue impact quantified
95.9% recall, deployed ML model
3+ years working with data

tools change. fundamentals don't.

What I work with

In [1]:languages()
Out[1]:
[
  • Python
  • R
  • SQL
  • Java
  • OOP
  • HTML
  • CSS
]
In [2]:machine_learning()
Out[2]:
[
  • Pandas
  • NumPy
  • Scikit-learn
  • XGBoost
  • TensorFlow
  • PyTorch
  • OpenCV
]
In [3]:visualization()
Out[3]:
[
  • Altair
  • Matplotlib
  • ggplot2
  • R Shiny
  • Tableau
  • Power BI
  • Excel
]
In [4]:tools()
Out[4]:
[
  • Git
  • GitHub
  • Docker
  • PostgreSQL
  • MySQL
  • Azure
  • AWS
  • Databricks
  • CI/CD
  • REST APIs
]
In [5]:coordination()
Out[5]:
[
  • Event Planning & Logistics
  • Cross-Team Collaboration
  • Stakeholder Communication
  • Research Communication
  • Safe Data Practices
]

Experience

  1. 2026

    Computational Research Assistant

    May 2026 – Present
    UBC Faculty of Pharmaceutical Sciences
    • Queried the ChEMBL API and ran everything through a multi-step filtering and validation pipeline in R to pull molecular descriptors and plasma protein binding values — got the training dataset to over 95% completeness.
    • Building a Random Forest model to predict milk-to-plasma drug ratios, wrapped in an end-to-end R package and Shiny app so researchers can go from API query to prediction without touching my code.
    • Also built PK Teaching Visualizations — interactive R Shiny apps teaching volume of distribution and drug clearance, including an SVG-animated steady-state infusion model, deployed on shinyapps.io.

    R · Machine Learning · Cheminformatics · ChEMBL API

  2. 2026

    Program Assistant

    May 2026 – Present
    UBC Master of Data Science Program
    • Wrote Python scripts against the Google Calendar API to build, populate, and publish an entire multi-calendar academic session schedule programmatically — turned a season of manual entry into something that just runs.
    • Coordinate program administration and event logistics, running the collaborative workflow through GitHub Issues and pull requests.

    Python · Google Calendar API · Project Coordination · GitHub

  3. 2026

    Machine Learning Intern

    January 2026 – March 2026
    RJH Biosciences
    • Parsed HGVS variant notation into genomic features with regex to build a structured dataset for FLT3 pathogenicity classification, then merged clinical datasets to eliminate duplicate leakage and pushed minority class representation from 7% to ~23%.
    • Built and trained an ML pipeline across 3 models landing 95.9% recall, then shipped it as a live web app for real-time predictions.

    Python · Machine Learning · Genomics · Regex

  4. 2025

    Centre Assistant

    September 2025 – January 2026
    Kumon–Sullivan

    Tutoring · Communication · Performance Tracking

  5. 2025

    WorkLearn Project Specialist

    September 2025 – January 2026
    UBC Chemical and Biological Engineering

    XRF Analysis · Documentation

  6. 2024

    Salaried Notetaker

    October 2024 – June 2025
    Centre for Accessibility, UBC

    Accessibility Support · Technical Notetaking

  7. 2023

    Data Analyst Intern

    February 2023 – April 2024
    Radar International
    • Cleaned, preprocessed, and validated structured datasets of 10,000+ records.
    • Built Excel and Python dashboards translating cleaned datasets into weekly KPI summaries used to make decisions.

    Python · Excel · Data Cleaning · Dashboards

Selected work

A few things I've built, analyzed, and investigated.

Research · Hero Project

MP.prediction

The questionDoes this drug show up in breast milk?

A machine learning pipeline predicting how much of a drug passes into breast milk, built on a 200+ compound dataset cross-validated across ChEMBL, PubChem, and OPERA, with a graph neural network (pkasolver) for pKa prediction.

200+compounds dataset
3APIs cross-validated

The findingTraced a systematic ~30 Ų surface-area discrepancy back to water-of-crystallization effects in source database structures — a data-quality issue relevant to cheminformatics pipelines broadly.

Python · R · RDKit · PyTorch · Graph Neural Networks · REST APIs

→ also what my Computational Research Assistant role is built around
Machine Learning

FLT3 Pathogenicity Classifier

The questionIs this genetic variant pathogenic?

95.9%recall
7%→23%minority class fixed

Parsed HGVS variant notation into genomic features with regex and trained 3 models for FLT3 pathogenicity classification, shipped as a live web app for real-time predictions.

Python · Regex · HGVS parsing · Clinical dataset merging

Live app ↗
Data Analytics · BI

Telecom Customer Churn Dashboard

The questionWhich customers are about to leave, and what's it costing us?

$139K+revenue lost to attrition
51%churn, first-year customers

Engineered churn-flag encoding, tenure cohorts, and CLV from a 7,000+ row telecom dataset, then built a Tableau KPI dashboard surfacing where attrition concentrated — 51% among first-year customers, 42% among fiber-optic users.

Python · Tableau · Feature Engineering · KPI Dashboarding

GitHub ↗
Business Analytics · Datathon

Vancouver City FC — Revenue Growth

The questionWhat happened to revenue?

Ran ETL and analysis on stadium operations, merchandise, and fanbase engagement data, then delivered stakeholder-facing recommendations for growing revenue across three key areas — all within a single datathon.

Python · pandas · NumPy · matplotlib

GitHub ↗
NLP · Statistics · Datathon

NLP & Survival Analysis for HR Retention

The questionWho's about to quit, and what's it costing us?

15–20%projected exit reduction
$2M+savings modeled

Built a pipeline across 5 datasets — VADER sentiment on Google reviews, Kaplan-Meier retention curves, logistic regression and chi-square tests — to model high-performer exit probability. BOLT Datathon Semi-Finalist.

Python · pandas · VADER (NLP) · Kaplan-Meier · Logistic Regression

GitHub ↗
Software Engineering

Financial Risk Calculator

The questionHow do you track risk without spreadsheet chaos?

An object-oriented Java application with a Swing UI for adding, filtering, deleting, saving, and loading financial reports — backed by JSON persistence, JUnit tests, and Git version control.

Java · Swing · JSON · JUnit

GitHub ↗
View all work →

Education

University of British Columbia

Expected 2028

B.Sc., Data Science — GPA 3.90 / 4.33, Dean's List

Relevant coursesModels of Computation · Software Construction · Data Science in R & Python · Data Visualization · Statistical Inference · Calculus I, II & III · Matrix Algebra · Applied Linear Algebra · Machine Learning Fundamentals

Beyond the dataset

What I do when I'm not debugging a dataframe.

It started with a breast cancer dataset in a class project — that spiralled into an ML internship, then a pharma research role. Outside of that:

photo coming soon Calligraphy
photo coming soon Travel
Coffee Addict
Co-Founder & VP Logistics UBC Punjab Students Association
Data Science Student Rep UBC Data Science Club
Corporate Relations Coordinator UBC Computer Science Student Society

Let's build something interesting.

Interested in data, research, visualization, or just talking about a cool problem?