Research Themes +
Energy Technologies Area (ETA) researchers are continually building on the strong scientific foundation we have developed over the past 50 years. We address the world’s most pressing scientific challenges across the buildings, transportation, and industrial sectors. ETA is at the forefront of improving the country's aging electrical grid and innovating distributed energy and storage solutions; developing grid-integrated building systems; and providing the most comprehensive market and data analysis worldwide.
Publications
News +
For media inquiries, please email:

[email protected]
About Us +
The Energy Technologies Area (ETA) is unique in translating fundamental scientific discoveries into scalable technology adoption. Our approach combines an understanding of the marketplace and the role of state and federal regulation and policies. ETA's research drives real-world, practical results that affect and improve the everyday lives of Americans and those across the globe. Saving energy and increasing reliability are key to the foundation of our research, which is driven by techno-economic analysis and in-lab experimentation and discovery.

Predicting childhood lead exposure at an aggregated level using machine learning

Publication Type

Journal Article

Date Published

10/2021

Authors

Lobo, G.P, B Kalyan, Ashok J Gadgil

DOI

https://doi.org/10.1016/j.ijheh.2021.113862

Abstract

Childhood lead exposure affects over 500,000 children under 6 years old in the US; however, only 14 states
recommend regular universal blood screening. Several studies have reported on the use of predictive models to
estimate lead exposure of individual children, albeit with limited success: lead exposure can vary greatly among
individuals, individual data is not easily accessible, and models trained in one location do not always perform
well in another. We report on a novel approach that uses machine learning to accurately predict elevated Blood
Lead Levels (BLLs) in large groups of children, using aggregated data. To that end, we used publicly available zip
code and city/town BLL data from the states of New York (n = 1642, excluding New York City) and Massa-
chusetts (n = 352), respectively. Five machine learning models were used to predict childhood lead exposure by
using socioeconomic, housing, and water quality predictive features. The best-performing model was a Random
Forest, with a 10-fold cross validation ROC AUC score of 0.91 and 0.85 for the Massachusetts and New York
datasets, respectively. The model was then tested with New York City data and the results compared to measured
BLLs at a borough level. The model yielded predictions in excellent agreement with measured data: at a city level
it predicted elevated BLL rates of 1.72% for the children in New York City, which is close to the measured value
of 1.73%. Predictive models, such as the one presented here, have the potential to help identify geographical
hotspots with significantly large occurrence of elevated lead blood levels in children so that limited resources
may be deployed to those who are most at risk.

Journal

International Journal of Hygiene and Environmental Health

Volume

238

Year of Publication

2021