Author(s)

Prakash L Pathak , Dr. Pradip D Jadhao

  • ISSN (P): 3139-8464
  • Manuscript ID: 140998
  • Volume: 2
  • Issue: 9
  • Pages: 150–155

Subject Area: Engineering

Abstract

Urban roadside PM2.5 pollution has become one of the more pressing problems facing fast-growing cities, driven by rising traffic and boundary layers that often do not ventilate near-road emissions effectively. In this paper we set out to build and test a hybrid modelling framework that couples a physically based Gaussian line-source dispersion model with a data-driven Random Forest model, in a model-output-statistics (bias-correction) configuration. We want to state clearly, up front, what kind of study this is: because our aim here was methodological — to work through the full modelling, validation and comparison pipeline end to end, and to leave a workflow that is directly reusable once real data are available — we built a synthetic, thirty-day hourly case study for a representative urban arterial corridor. Traffic, meteorological and atmospheric-stability inputs were drawn from patterns and emission factors reported in the published dispersion literature, not from a specific city’s monitoring network. The "observed" concentrations used to validate the models are pseudo-observations, generated by adding representativeness-type noise to the physical model’s own output, rather than independent field measurements.
We compared three configurations - the Gaussian dispersion model on its own, a standalone Random Forest model, and the hybrid combination — on an identical held-out test partition, using six standard air-quality validation statistics: RMSE, MAE, R², mean bias error (MBE), fractional bias (FB), and Willmott’s index of agreement (IOA). The Gaussian model came out ahead on every overall-accuracy statistic (RMSE = 5.10 µg/m³, R² = 0.80, IOA = 0.94); this is expected, since the pseudo-observations were generated from that same model plus noise, so it really sets an upper bound for this exercise rather than an independent benchmark. The standalone Random Forest model trailed on every measure (RMSE = 8.01 µg/m³, R² = 0.50, IOA = 0.79). The hybrid model recovered almost all of the physical model’s accuracy while roughly halving the fractional bias (FB = 0.008 versus 0.015), which suggests that residual correction earns its keep mainly by removing systematic bias rather than by smoothing random noise. A feature-importance analysis pointed to traffic volume, wind direction and mixing height as the strongest predictors of roadside PM2.5 in this setup. Taken together, the results support a practical way of pairing the two modelling philosophies: keep the physical dispersion model as the primary predictive engine, and use machine learning mainly for bias correction, gap-filling and short-term forecasting rather than as a full replacement. We close with some thoughts on what it would take to extend this framework — real monitoring data, multi-pollutant chemistry, street-canyon-resolving CFD — into something usable for actual urban air-quality decisions.

Keywords
Air pollution modelling; Gaussian dispersion; Urban PM2.5; Machine learning; Random Forest; model validation; Hybrid simulation; Traffic emissions