View Live Demo View on GitHub

NCR Pollution Intelligence: AI-Powered Pollution Source Identification and Forecasting

Air pollution in Delhi-NCR is not caused by one constant factor. Traffic, construction dust, industrial combustion, agricultural fires, weather, and chemical reactions in the atmosphere can all contribute to poor air quality. NCR Pollution Intelligence is an AI-powered command center designed to understand those changing conditions, identify likely pollution sources, and forecast what the air may look like over the next 24 hours.

The system combines live weather and pollutant observations, historical time-series data, satellite fire information, hyperlocal calibration, machine-learning models, and explainable visualizations. This makes it more than a sensor reader: it is a spatially aware intelligence system for studying how, why, and when pollution changes across the region.

NCR Pollution Intelligence dashboard showing AQI forecasting and pollution source analysis

Project Objective

The central objective is to convert complex environmental data into a clear operational answer. For a selected station in the NCR region, the system estimates the current Air Quality Index, forecasts the next 24 hours, and explains the most probable reason for the observed pollution.

Instead of presenting raw API responses or isolated numbers, the dashboard answers three practical questions:

  1. How polluted is the air right now?
  2. How could air quality change during the next day?
  3. Which source category best explains the chemical fingerprint?

Real-Time Data Ingestion Layer

When a user selects an NCR station and clicks Sync Live Data, the backend starts three complementary data pipelines. These pipelines give the system awareness of both the atmosphere above the station and important events occurring on the ground around it.

1. OpenWeather Pollution and Weather Data

The application first resolves the selected station's geographical coordinates and requests live environmental measurements from the OpenWeather API. The incoming pollution data includes PM2.5, PM10, nitrogen dioxide (NO2), sulfur dioxide (SO2), carbon monoxide (CO), and ozone (O3).

Alongside pollutant concentrations, the system records temperature, humidity, and wind speed. These weather variables matter because they influence dispersion, accumulation, chemical reactions, and the movement of particulate matter.

2. Historical Pollution Timeline

A current reading alone cannot describe the direction of a pollution event. The historical pipeline looks approximately 72 hours into the past and builds a sequential timeline for the selected location. This three-day window allows the forecasting engine to learn recent spikes, daily cycles, persistent episodes, and the relationship between pollution and changing weather.

3. NASA FIRMS Satellite Fire Data

The system also connects to the NASA FIRMS feed for SUOMI NPP VIIRS C2 thermal-anomaly observations. These satellite observations help reveal active farm fires and industrial heat events that may influence regional air quality.

To make the satellite data relevant to the selected station, the backend calculates the distance between the station and each thermal anomaly using the Haversine formula. It then counts fire or heat events within a precise 50-kilometre radius. This produces a local fire-activity signal that can be used by the prediction and source-identification models.

Hyperlocal Calibration Engine

Delhi-NCR is a connected region, but every station has a different environmental profile. A traffic-heavy crossing does not have the same pollution signature as an industrial area or a location closer to agricultural activity.

To account for these differences, the backend uses a scientifically informed STATION_PROFILES dictionary. Each profile applies location-specific mathematical multipliers to the incoming measurements. For example, Bhiwadi can receive stronger PM10 and SO2 calibration because of its industrial character, while Anand Vihar can receive stronger NO2 calibration because it is a major traffic choke point.

This calibration layer prevents the model from treating the entire NCR as one uniform city. It gives the forecasting pipeline a more realistic local context while keeping the data-ingestion architecture reusable for additional stations.

Hybrid Machine-Learning Pipeline

The system uses multiple models because forecasting and source identification are related but different tasks. A time-series model is best suited to predicting future AQI, a regression fallback keeps the dashboard available when history is incomplete, and a classification model translates pollutant relationships into actionable source categories.

Primary Forecasting Engine: LSTM

The primary forecasting model is a Long Short-Term Memory (LSTM) neural network. LSTMs are designed for sequential data because they can retain useful information from earlier time steps while learning which patterns should be forgotten.

The model receives a three-dimensional sequence containing approximately 72 hours of history across 9 environmental features. By studying recent rises and falls in pollution alongside temperature, humidity, wind, and related observations, it estimates the current AQI and forecasts the following 24 hours.

This approach is useful for air quality because pollution has momentum. A stagnant atmosphere can allow pollutants to accumulate, while stronger wind or a change in humidity can alter the next stage of the event. The LSTM attempts to learn those temporal relationships rather than relying only on one instantaneous reading.

Fallback Model: Gradient Boosting Regressor

Live systems must remain useful even when an external service returns incomplete history. If the required 72-hour sequence is unavailable or fails quality checks, the application switches to a Gradient Boosting Regressor (GBR).

The fallback uses engineered features such as the AQI from 24 hours earlier, current environmental readings, and satellite fire radiative power (FRP). Gradient boosting combines many smaller decision-tree models to produce a fast regression estimate. It gives the interface a dependable baseline and also supplies the feature contributions used in the dashboard's SHAP-based explanation.

Source Identification: XGBoost Classification

The source-identification engine uses an XGBoost-style classification approach to organize the likely cause of pollution into five categories. Its decision process examines chemical fingerprints rather than relying on a single pollutant.

These thresholds are interpretable: they let a presenter explain not only the model's label, but also the pollutant relationship that contributed to it.

Dashboard and Intelligence Output

The frontend turns the pipeline into an operations dashboard so that users can understand the result without reading raw JSON, model logs, or API responses.

End-to-End System Flow

  1. The user selects an NCR monitoring station.
  2. The application synchronizes current weather, pollutant, historical, and satellite data.
  3. The backend validates the incoming values and applies the selected station's calibration profile.
  4. The LSTM forecasts the current and next 24 hours when the full historical sequence is available.
  5. The GBR provides a fallback estimate when the historical input is incomplete.
  6. The source-identification model evaluates pollutant ratios and environmental signals.
  7. The dashboard presents AQI, forecast, source probabilities, health advice, feature explanations, and synchronization status.
  8. The inference is recorded in SQLite and can be exported as a PDF report.

How to Present the Project

For a college presentation, begin with the interface and show the live synchronization control. Explain that one action starts weather, historical, and satellite requests in parallel. Then move to the primary AQI forecast and describe how the LSTM learns from the previous 72 hours.

Next, focus on the source-classification probabilities. Use the PM2.5-to-PM10 and SO2-to-NO2 relationships to explain how the system maps a chemical footprint to traffic, smoke, industry, construction dust, or mixed pollution. Finish by exporting the PDF and showing the prediction history to demonstrate that the application is not only visual, but also persistent, explainable, and auditable.

Try It Live and View the Source Code

Open the NCR Pollution Intelligence live demo on Hugging Face to explore the dashboard and its forecasting workflow.

View the complete project source code on GitHub.


Connect with Me
www.tauqueeralam.com
LinkedIn | GitHub

Discussion