Spatial Clustering of Cancer Risk from Air Toxics Across Census Tracts in California
by: Nicole Vicenti
by: Nicole Vicenti
Air toxics are hazardous air pollutants that can increase cancer risk and pose significant public health concerns, particularly in areas with prolonged exposure (San Joaquin Valley Air Pollution Control District, n.d.). This study analyzes the spatial distribution of cancer risk from air toxics across census tracts in California using data from the Environmental Protection Agency (EPA) EnviroAtlas to identify statistically significant clusters, hot spots, and spatial outliers. Understanding these spatial patterns can help reveal areas of elevated environmental risk and potential geographic inequalities in exposure.
Research Question
Where are statistically significant hot spots, cold spots, and spatial outliers of cancer risk from air toxics located across California census tracts?
Study Area
The study area for this project is the state of California, the most populous state in the United States and one of the most environmentally diverse regions in the country. California frequently experiences some of the worst air quality in the nation due to a combination of factors including high population density, transportation emissions, industrial activity, and geographic features that trap pollutants in valleys such as the Central Valley (Parajuli, 2025). According to the California Air Resources Board, a large proportion of residents are exposed to unhealthy air conditions, making it a critical region for environmental health analysis. Census tracts are used as the unit of analysis to provide a detailed, localized understanding of how cancer risk from air toxics varies spatially across the state.
Data
The primary dataset used in this analysis is cancer risk per million from cumulative air toxics, obtained from the Environmental Protection Agency EnviroAtlas. This dataset represents the estimated lifetime probability of developing cancer due to continuous exposure to air toxics over approximately 70 years. The data is provided as a polygon feature layer at the census tract level for the entire United States and was clipped to the boundary of California prior to analysis (Figure 1).
Metadata:
Name of Dataset: Cancer risk per million due to cumulative air toxics
Publication/Update Date: April 25, 2025
Geometry Type: Polygon Feature Layer
Owner: Environmental Protection Agency (EPA), EnviroAtlas
Description: This dataset estimates cancer risk per million based on cumulative exposure to air toxics from the National Air Toxics Assessment (NATA). Cancer risk is defined as the probability of developing cancer over a lifetime (assumed to be 70 years) under continuous exposure conditions.
Figure 1. Cancer risk per million from cumulative air toxics exposure across census tracts in California.
The dataset was projected in NAD 1983 (2011) California (Teale) Albers (Meters) (WKID 6414), a statewide projected coordinate system designed for California. This projection preserves area relationships and uses meters as the unit of measurement, making it well suited for mapping census tract polygons and conducting distance-based spatial statistical analyses such as Hot Spot Analysis and Cluster and Outlier Analysis.
The variable of analysis used in this study is cancer risk per million due to cumulative air toxics, obtained from the Environmental Protection Agency EnviroAtlas dataset. This variable represents the estimated lifetime probability of developing cancer per one million people as a result of continuous exposure to airborne toxic pollutants over approximately 70 years. It is a continuous numeric variable measured at the census tract level, where higher values indicate greater estimated cancer risk and potential areas of concern for environmental exposure.
Hot Spot Analysis
Hot Spot Analysis was conducted to identify statistically significant spatial clusters of high and low cancer risk values. As an initial exploratory step, the Optimized Hot Spot Analysis tool was used to estimate an appropriate distance band; however, the resulting distance was too small and produced many census tracts with no neighbors, which violates the assumptions of spatial statistical analysis and reduces the reliability of results. To address this, Incremental Spatial Autocorrelation was used to determine a more appropriate distance band based on peak spatial clustering.
The Hot Spot Analysis (Getis-Ord Gi*) tool evaluates each census tract in relation to its neighboring features to identify statistically significant hot spots (high values) and cold spots (low values). A Fixed Distance Band was selected as the conceptualization of spatial relationships, as it better represents how air pollutants disperse across space.
Cluster and Outlier Analysis
The Cluster and Outlier Analysis (Anselin Local Moran’s I) tool was used to identify statistically significant spatial clusters and outliers within the dataset. This method evaluates each census tract relative to its neighbors to classify patterns as high-high clusters, low-low clusters, high-low outliers, or low-high outliers.
The analysis was conducted using a Fixed Distance Band with the same distance determined from the Incremental Spatial Autocorrelation results. Row standardization was applied to the spatial weights matrix to account for variation in the number of neighboring census tracts, ensuring that each feature has equal influence and reducing bias caused by differences in tract size and spatial distribution. A total of 999 permutations were used to increase the robustness of the statistical results.
Assumptions
Several assumptions were made to ensure appropriate analysis of California’s uneven spatial distribution of census tracts. The dataset was already aggregated at the census tract level, which assumes that cancer risk is uniformly distributed within each tract. A Fixed Distance Band of 92,247 meters was selected based on Incremental Spatial Autocorrelation, which identified the distance at which spatial clustering was strongest. This distance was substantially larger than the value suggested by Optimized Hot Spot Analysis (~23,188 meters), which proved too restrictive and resulted in many tracts having no neighbors, particularly in rural areas.
While the selected distance band ensures all census tracts have sufficient neighbors, it may generalize localized patterns in dense urban areas, potentially smoothing smaller clusters of high or low cancer risk.
False Discovery Rate (FDR) correction was applied to adjust for multiple testing and spatial dependency, resulting in more conservative and reliable significance levels. Euclidean distance was used to measure spatial relationships, assuming straight-line distance between features. This approach is appropriate for statewide analysis where general spatial proximity is more important than network-based movement.
Hot Spot Analysis
The Hot Spot Analysis (Getis-Ord Gi*) was conducted on over 8,000 census tracts across California to identify statistically significant clusters of high and low cancer risk from air toxics (Figure 2). The results indicate a strong spatial pattern, with 4,941 tracts identified as hot spots at the 99% confidence level, along with 135 tracts at 95% confidence and 60 at 90% confidence. Cold spots were also present, with 2,470 tracts at 99% confidence, 27 at 95%, and 25 at 90%, while 411 tracts were not statistically significant.
Visually, high-confidence hot spots (99%) are concentrated in major population centers, including the Sacramento region, the Central Valley around Fresno, and large portions of Southern California, particularly the Los Angeles and San Diego metropolitan areas. These clusters suggest that elevated cancer risk from air toxics is spatially concentrated in urban and agriculturally intensive regions where emissions from transportation, industry, and other sources are likely higher. In contrast, cold spots (99% confidence) are predominantly located along the coastal regions and near the eastern border with Nevada, indicating areas of consistently lower cancer risk relative to surrounding tracts.
The dominance of 99% confidence classifications and the relatively small number of moderately significant (90%–95%) or non-significant tracts suggest a strong and consistent spatial clustering pattern across the state. This indicates that cancer risk from air toxics is not randomly distributed but instead exhibits significant geographic concentration. The results highlight clear regional disparities, with inland and urbanized areas experiencing higher risk compared to coastal and less densely populated regions.
Figure 2. Hot Spot Analysis (Getis-Ord Gi*) of cancer risk from air toxics across census tracts in California, showing statistically significant hot and cold spots.
Cluster and Outlier Analysis
The Cluster and Outlier Analysis (Anselin Local Moran’s I) was conducted to identify statistically significant spatial clusters and outliers of cancer risk from air toxics across census tracts in California. This method classifies features based on their value and the values of neighboring features, identifying High-High (HH) clusters, Low-Low (LL) clusters, High-Low (HL) outliers, and Low-High (LH) outliers. Only features with a statistical significance of 95% confidence or greater were included in these classifications (Figure 3).
The results identified 3,969 High-High clusters and 2,351 Low-Low clusters, along with 171 High-Low outliers and 1,164 Low-High outliers. High-High clusters are primarily concentrated in inland regions, including the Central Valley and portions of Southern California, indicating areas where high cancer risk values are surrounded by similarly high values. In contrast, Low-Low clusters are more prevalent along coastal regions, reflecting areas of consistently lower cancer risk.
A notable pattern in the results is the presence of a large number of Low-High outliers, particularly in regions adjacent to major hot spot areas. These represent census tracts with relatively low cancer risk values surrounded by higher-risk neighbors, suggesting localized variation within broader high-risk regions. High-Low outliers are less common and appear more scattered, indicating isolated tracts of higher cancer risk within generally low-risk areas.
Overall, the cluster and outlier patterns largely align with the Hot Spot Analysis results, confirming strong regional clustering of cancer risk. However, the Local Moran’s I analysis provides additional detail by identifying localized inconsistencies within these broader patterns, highlighting areas where cancer risk deviates from surrounding trends.
Figure 3. Cluster and Outlier Analysis (Anselin Local Moran’s I) of cancer risk from air toxics across census tracts in California, identifying spatial clusters and outliers.
The Hot Spot Analysis and Cluster and Outlier Analysis produced largely consistent results, revealing clear spatial patterns in cancer risk from air toxics across California. Both analyses identified statistically significant clusters of high cancer risk concentrated in major urban and inland regions, including Los Angeles, Sacramento, Fresno, and San Diego. These patterns suggest that elevated cancer risk is associated with areas of higher population density and increased emissions from transportation, industrial activity, and other anthropogenic sources. In contrast, areas of consistently lower cancer risk were primarily located along coastal regions and near the Nevada border.
The high proportion of results classified at the 99% confidence level indicates that these spatial patterns are unlikely to be the result of random chance, demonstrating strong spatial autocorrelation in the dataset. While the Hot Spot Analysis highlighted broader regional clustering, the Cluster and Outlier Analysis provided additional insight into localized variation. In particular, the presence of numerous Low-High outliers in areas adjacent to major hot spots suggests that some census tracts exhibit lower cancer risk despite being surrounded by higher-risk neighbors, indicating heterogeneity within high-risk regions. Together, these analyses confirm that cancer risk from air toxics is not randomly distributed, but instead exhibits significant geographic concentration and localized variation. These findings highlight the importance of spatial analysis in identifying environmental health disparities and can inform future planning and policy decisions.
Limitations
Several limitations should be considered when interpreting these results. The use of a Fixed Distance Band required increasing the distance threshold to ensure that all census tracts had sufficient neighboring features, particularly in rural areas. While this improved the reliability of the analysis, it may have generalized spatial patterns in densely populated urban areas by increasing the number of neighbors considered. As a result, smaller, localized clusters may have been smoothed, potentially masking fine-scale variation in cancer risk.
Additionally, the analysis relies on data aggregated at the census tract level, which assumes that cancer risk is uniformly distributed within each tract. This may obscure intra-tract variability and introduce the Modifiable Areal Unit Problem (MAUP). Future research could incorporate finer spatial resolution data or alternative spatial relationship definitions to better capture localized patterns and improve the precision of the analysis.
California Air Resources Board. (n.d.). Health & air pollution. https://ww2.arb.ca.gov/resources/health-air-pollution
Sagar Parajuli. (2025, December 13). Why the Central Valley traps clouds and air pollutants. Extreme Heat and Air Pollution Lab. https://heal.sdsu.edu/why-the-central-valley-traps-clouds-and-air-pollutants/
San Joaquin Valley Air Pollution Control District. (n.d.). What are air toxics? https://ww2.valleyair.org/permitting/air-toxics-program/information-for-the-public/what-are-air-toxics/
United States Environmental Protection Agency. EnviroAtlas. Cancer risk per million due to cumulative air toxics. Accessed: [April, 24, 2026] from https://www.epa.gov/enviroatlas