SAHIL-SANGHVI(1)→ TERMINAL(1)

Island Insight

Youth-health engagement analysis for RBC Borealis' Let's Solve It 2026 (Team Island Insight x KHLF).

README

Upopolis Analytics Dashboard

Upopolis dashboard home page

This repository analyzes Google Analytics 4 data for Upopolis, a Kids Health Links Foundation platform. The project turns a stacked GA4 export into a Streamlit dashboard, reproducible notebooks, and generated analysis artifacts for stakeholder reporting.

The current project state is:

  • raw GA4 EDA is complete
  • traffic-quality / bot-suspicion audit is complete
  • bot-aware engagement baseline is complete
  • final bot-aware EDA dashboard is complete
  • aggregate user journey UMAP analysis is complete

The work is intentionally careful about what the available data can and cannot prove. The current GA4 file is aggregate data, not session-level or user-level tracking. That means the dashboard can support population-level engagement, traffic quality, content discovery, and ideal journey analysis. It should not be presented as individual youth disengagement prediction.

Project Goal

The proposal asks for a Youth Engagement Analysis Tool that helps Upopolis and KHLF understand engagement trends, retention signals, program effectiveness, and funder-ready outcomes.

This repository currently delivers the foundation for that tool:

  • a stakeholder-friendly dashboard
  • bot-aware engagement metrics
  • transparent data-quality warnings
  • aggregate content journey mapping
  • plain-English methodology and limitations

Because BigQuery/session-level data is not currently available, the project uses an aggregate journey approach instead of claiming to observe exact user paths.

Data

Main input:

data/full_data.csv

Current snapshot:

June 1, 2023 to April 13, 2026
1,048 days of GA4 history
16 parsed GA4 report sections

The export is a stack of smaller GA4 reports, including daily active users, new users, engagement time, traffic channels, countries, retention cohorts, page views, and event counts. The shared parser lives in:

upopolis/io.py

Dashboard

Run the dashboard from the repo root:

PYTHONPATH=. streamlit run streamlit_app.py

Dashboard pages:

Page Purpose
Home Entry page and dataset summary
Executive Overview Stakeholder summary of the main findings
Methodology Definitions, thresholds, assumptions, and limitations
Raw GA4 EDA Original unfiltered GA4 exploration
Traffic Quality Audit Bot-suspicion and traffic-quality review
Bot Aware Baseline Raw vs bot-aware vs conservative KPI comparison
Final Bot Aware EDA Stakeholder-facing bot-aware EDA
Aggregate User Journey UMAP-based aggregate content journey map

The stakeholder-facing result is the Final Bot Aware EDA plus the Aggregate User Journey page.

Notebook Workflow

Run the notebooks in this order:

1. notebooks/upoplis_intial_eda.ipynb
2. notebooks/bot_traffic_quality_audit.ipynb
3. notebooks/bot_aware_engagement_baseline.ipynb
4. notebooks/upopolis_final_bot_aware_eda.ipynb
5. notebooks/aggregate_user_journey_umap.ipynb

1. Raw GA4 EDA

File:

notebooks/upoplis_intial_eda.ipynb

Purpose:

Establishes the original GA4 baseline: daily users, new users, monthly trends, weekday patterns, acquisition channels, countries, page views, events, and raw retention cohorts.

Key raw numbers:

Active user-days:          41,913
Average daily active users: 39.99
Peak DAU:                    352 on February 11, 2026
Total new users:          36,998
Returning activity share:  11.8%

The raw EDA is useful, but it is not the final interpretation because several patterns suggest traffic-quality risk.

2. Traffic Quality Audit

File:

notebooks/bot_traffic_quality_audit.ipynb

Output folder:

analysis/bot_traffic_quality_audit/

Purpose:

Flags days and patterns that look bot-suspicious or low quality. This is not a confirmed bot detector. It is a transparent warning system based on aggregate GA4 signals.

Current audit summary:

Total days analyzed:       1,048
Tracking warm-up days:        90
Bot-suspicious days:         278
Uncertain days:              178
ARIMA anomaly days:           43
Review-country share:       40.5%
Suspicious page-title share:  6.2%
Passive/basic event share:   99.8%

The first nonzero traffic day is June 15, 2023. The first 90 days after that are treated as a tracking warm-up period, not as bot traffic.

3. Bot-Aware Engagement Baseline

File:

notebooks/bot_aware_engagement_baseline.ipynb

Output folder:

analysis/bot_aware_engagement_baseline/

Purpose:

Compares three versions of the same engagement baseline:

View Meaning
Raw baseline Every day in GA4
Bot-aware baseline Excludes days labelled bot_suspicious
Conservative baseline Keeps only the cleanest days

Headline comparison:

Metric Raw GA4 Bot-aware
Days included 1,048 770
Active user-days 41,913 21,115
Average daily active users 39.99 27.42
Median daily active users 29.0 26.5
Peak DAU 352 83
Total new users 36,998 17,208
Returning activity share 11.8% 18.5%

Interpretation:

Raw GA4 makes the platform look larger and more new-user-heavy. The bot-aware baseline is smaller but more defensible for stakeholder reporting.

4. Final Bot-Aware EDA

File:

notebooks/upopolis_final_bot_aware_eda.ipynb

Purpose:

Recreates the original EDA flow using the bot-aware outputs and clear caveats. This is the safest version to use for reporting.

It covers:

  • headline KPIs
  • daily visitors
  • monthly active-user trends
  • weekday patterns
  • acquisition channels
  • countries
  • trend windows
  • retention cohorts
  • content and event views
  • new vs returning users
  • correlations
  • final findings

Some GA4 sections are still raw rollups because they cannot be filtered at the daily or user level. This applies mainly to channels, events, and cohort retention.

5. Aggregate User Journey UMAP

File:

notebooks/aggregate_user_journey_umap.ipynb

Output folder:

analysis/aggregate_user_journey/

Purpose:

Creates an aggregate content journey map without BigQuery or session-level tracking. It does not show the exact path any one user took. Instead, it maps the Upopolis content ecosystem and compares ideal user journeys against aggregate page attention.

Generated outputs:

analysis/aggregate_user_journey/aggregate_journey_umap.csv
analysis/aggregate_user_journey/journey_stage_summary.csv
analysis/aggregate_user_journey/ideal_journey_summary.csv
analysis/aggregate_user_journey/aggregate_user_journey_summary.json
analysis/aggregate_user_journey/aggregate_journey_map.png
analysis/aggregate_user_journey/journey_stage_views.png
analysis/aggregate_user_journey/ideal_journey_flags.png

Current aggregate journey findings:

Canonical pages mapped:       2,240
Canonical page views:       104,743
Embedding method:              UMAP
Mission-aligned view share:   16.1%
Admin/noise view share:       19.2%

Top journey-stage shares:

Stage View share
Entry / Orientation 31.8%
Admin / Noise 19.2%
Unclassified 13.0%
Browse / Directory 8.9%
Program / Event 7.7%
Professional / Caregiver 5.8%
Account / Login 5.3%
Medical Information 4.9%
Coping / Emotional Support 1.8%
Peer / Community Support 1.7%

Plain-English interpretation:

The site receives attention at the entry and navigation layer, but the pages most closely tied to the ideal support journey receive a smaller share of total views. This is useful for content discovery and program design. It is not proof of individual drop-off.

What Can Be Trusted Right Now

Higher confidence:

  • the GA4 export is parsed consistently
  • raw usage trends are available for the full reporting window
  • traffic-quality risk is real enough to affect reporting
  • bot-aware average DAU is more defensible than raw average DAU
  • single-day traffic peaks should be reviewed before being quoted
  • aggregate content journey mapping can identify discovery gaps

Medium confidence:

  • country mix is useful for review but not proof of real program reach
  • content and page-view analysis is useful but may include crawler activity
  • returning activity share improves after suspicious days are removed
  • UMAP is useful as a visual content map, not as exact journey tracking

Low confidence until better data is available:

  • individual youth disengagement
  • clinical risk flags
  • exact user-level retention
  • exact page-to-page paths
  • confirmed bot classification

Data Limitations

The current work should be described as:

aggregate GA4 analysis with bot-aware traffic-quality review

and:

aggregate content journey mapping

It should not be described as:

confirmed bot detection
individual churn prediction
clinical disengagement risk scoring
session-level journey tracking

The project would need more detailed data for those claims, such as:

  • GA4 BigQuery export
  • session-level events
  • user-level return paths
  • user agent
  • IP or network information
  • full page paths
  • server logs
  • reliable engagement-time instrumentation

Repository Structure

streamlit_app.py                         # Streamlit dashboard entrypoint
pages/                                   # Dashboard pages
upopolis/                                # Parser, metrics, charts, dashboard helpers
notebooks/                               # Reproducible notebook workflow
analysis/bot_traffic_quality_audit/      # Traffic-quality outputs
analysis/bot_aware_engagement_baseline/  # Bot-aware baseline outputs
analysis/aggregate_user_journey/         # UMAP journey outputs
data/full_data.csv                       # Canonical GA4 export
images/Home.png                          # Dashboard home screenshot
tests/                                   # Parser and quality tests

Additional analysis folders under analysis/ and ml/ contain earlier or experimental work. The primary stakeholder workflow is the notebook chain and dashboard described above.

Setup

Create a virtual environment:

python3 -m venv .venv
source .venv/bin/activate

Install the main dependencies:

pip install pandas numpy matplotlib seaborn scipy statsmodels scikit-learn \
    umap-learn jupyter streamlit plotly pytest

If umap-learn is unavailable, the aggregate journey notebook is designed to fall back to a simpler 2D projection, but UMAP is recommended.

Useful Commands

Run tests:

python -m pytest -q

Run parser section summary:

python -m upopolis sections

Run validation:

python -m upopolis validate

Note: validation currently reports raw retention-cohort monotonicity warnings. That is an expected data caveat and is why retention is treated carefully in the dashboard.

Run dashboard:

PYTHONPATH=. streamlit run streamlit_app.py

Open notebooks:

jupyter notebook notebooks/upoplis_intial_eda.ipynb
jupyter notebook notebooks/bot_traffic_quality_audit.ipynb
jupyter notebook notebooks/bot_aware_engagement_baseline.ipynb
jupyter notebook notebooks/upopolis_final_bot_aware_eda.ipynb
jupyter notebook notebooks/aggregate_user_journey_umap.ipynb

Recommended Reporting Language

Use:

The EDA and bot-aware baseline are complete. The current dashboard provides aggregate engagement trends, traffic-quality review, and aggregate content journey mapping from the GA4 data currently available.

Avoid:

The model predicts youth disengagement.

Use:

The aggregate journey map shows where Upopolis content supports or fails to support ideal user journeys.

Avoid:

The journey map shows exactly how users moved through the site.

Current Bottom Line

The project is ready for stakeholder review and aggregate journey discussion.

The next useful decision is not to force a weak churn model. It is to review the bot-aware findings and aggregate user journey map with KHLF / Upopolis, then agree on what meaningful engagement and possible disengagement should mean for this platform.

Once session-level or BigQuery data becomes available, the project can move from aggregate journey mapping toward true user-level retention and disengagement analysis.