A Comparative Point Pattern Analysis of Public-Supply Wells in Arizona
by: Nicole Vicenti
by: Nicole Vicenti
Public-supply wells are a critical component of water infrastructure, providing drinking water and supporting household needs for more than one-third of the U.S. population (U.S. Geological Survey, 2019). In arid regions such as Arizona, where surface water is limited and groundwater is heavily relied upon, understanding the spatial distribution of these wells is essential for effective resource management and planning. This study analyzes the spatial pattern of public-supply wells across Arizona to determine whether their distribution is clustered, dispersed, or random.
Null Hypothesis
The spatial distribution of public-supply wells across Arizona is random.
Research Question
Using Quadrat Count, Average Nearest Neighbor, and Ripley's K methods, does the spatial distribution of public-supply wells across Arizona exhibit a clustered, dispersed, or random pattern?
Study Area
The study area for this project encompasses the state of Arizona, located in the southwestern United States. Arizona is characterized by an arid to semi-arid climate, with limited surface water resources and a strong dependence on groundwater for municipal and domestic use. Major population centers such as Phoenix and Tucson rely heavily on groundwater systems, particularly in areas where surface water availability is constrained. The study area is defined using a rectangular boundary that closely encloses the extent of the Arizona state border, providing a consistent spatial frame for point pattern analysis.
Data
The dataset used in this analysis is the Groundwater Site Inventory (GWSI) Sites 2024, maintained by the Arizona Department of Water Resources. This dataset contains point locations and associated attribute information for groundwater wells across Arizona. For this study, only wells classified as “Public Supply” were selected, resulting in a total of 2,946 point features used for analysis (Figure 1).
Name of Dataset: GWSI Sites 2024
Publication/Update Date: April 17, 2026
Geometry Type: Point Feature Layer
Owner: Arizona Department of Water Resources
Description: This dataset represents ADWR’s primary statewide repository for groundwater site locations and related data.
Figure 1. Study area boundary and public-supply well locations used in point pattern analyses, projected in a custom Lambert Conformal Conic coordinate system.
A Lambert Conformal Conic (LCC) projection was used to minimize distortion in distance and area across the study area. Because the dataset spans the entire state of Arizona, the original NAD 1983 UTM Zone 12N projection was less suitable due to Arizona extending beyond a single UTM zone. A custom LCC projection was created with a central meridian of −112°, standard parallels at 32° and 37°, and a latitude of origin of 34°, using meters as the linear unit. These parameters align with Arizona’s geographic extent and help maintain consistent spatial accuracy across the state.
A total of 2,946 public-supply wells were included in this analysis. Each point represents the geographic location of a well used to provide drinking water and household supply to the public.
The extent of analysis was defined using a minimum bounding rectangle that encompasses the full spatial extent of Arizona. The rectangle was created using the extreme latitude and longitude coordinates of the dataset (north, south, east, and west limits) to ensure all points were included. While this approach introduces some empty space outside the true state boundary, it provides a consistent and simplified study area required for Quadrat Count and Ripley’s K analysis.
Quadrat Count Method
The Quadrat Count method is a density-based approach that evaluates spatial patterns by dividing the study area into equal-sized grid cells. In this study, a 10 × 10 grid was applied across the study area, and the number of wells within each quadrat was counted. Each quadrat covered approximately 863,038 acres. The frequency distribution of counts was used to calculate the mean, variance, and the Variance-Mean Ratio (VMR). A VMR greater than 1 indicates clustering, a value less than 1 indicates dispersion, and a value equal to 1 suggests a random distribution.
Average Nearest Neighbor Method
The Average Nearest Neighbor (ANN) method is a distance-based analysis that measures the average distance between each point and its nearest neighboring point using Euclidean distance. The observed mean distance between well points was compared to the expected mean distance of a hypothetical random distribution. The ANN ratio was calculated as the observed mean distance divided by the expected mean distance. Values greater than 1 indicate a dispersed pattern, while values less than 1 indicate clustering.
Ripley's K Method
Ripley’s K function evaluates spatial dependence over multiple distance scales by comparing the observed distribution of points to an expected random distribution. In this analysis, 10 distance bands were used with 999 permutations to generate confidence envelopes. The beginning distance was set to 0, with a distance increment of 2 meters. The study area was defined using the minimum enclosing rectangle. Observed K values above the upper confidence envelope indicate clustering, values below the lower confidence envelope indicate dispersion, and values between the envelopes indicate a random spatial pattern.
Quadrat Count Method
A total of 2,946 public-supply wells were distributed across 100 quadrats (Figure 2), resulting in a mean quadrat count of 29.46. The observed variance was 4,511.83 (Table 1), producing a Variance-to-Mean Ratio (VMR) of 153.15, which is substantially greater than 1. This indicates a strongly clustered spatial pattern of public-supply wells across Arizona. The large difference between the mean and variance suggests that wells are unevenly distributed, with some quadrats containing many wells while others contain few or none. This pattern reflects significant spatial concentration rather than a random or evenly dispersed distribution.
Figure 2. Quadrat grids with public supply well points and frequencies.
Table 1. Quadrat counts and calculation of VMR for the public-supply well pattern.
Average Nearest Neighbor
The Average Nearest Neighbor (ANN) analysis indicates a statistically significant clustered pattern (Figure 3). The ANN ratio of 0.289 is well below 1, confirming that wells are closer together than would be expected under a random distribution. The extremely negative z-score (−73.82) and p-value of 0.00 indicate that this clustering is highly significant and not the result of random chance. These results strongly support a clustered spatial distribution of public-supply wells across the study area.
Ripley's K Method
Ripley’s K analysis indicates significant clustering across all evaluated distance scales (Figure 4). The observed K values consistently plot above the upper confidence envelope, demonstrating that well locations are more spatially concentrated than expected under complete spatial randomness. This pattern persists across increasing distances, indicating that clustering occurs at both local and broader spatial scales. The consistency of this pattern across multiple distance bands reinforces the conclusion that the distribution of public-supply wells is strongly clustered throughout the study area.
Figure 4. Observed K values compared to expected K and upper and lower confidence envelopes derived from the Multi-Distance Spatial Cluster Analysis (Ripley’s K function).
Overall, all three point pattern analysis methods indicate a strongly clustered spatial distribution of public-supply wells across Arizona. The Quadrat Count method, a density-based approach, produced a Variance-Mean Ratio (VMR) of 153.15, which is far greater than 1 and indicates a highly uneven distribution of wells across the study area. This suggests that wells are concentrated in specific regions rather than evenly distributed. The Average Nearest Neighbor (ANN) analysis supports this finding, with an ANN ratio of 0.289, a z-score of −73.82, and a p-value of 0.00, indicating statistically significant clustering and a very low probability that the observed pattern is due to random chance.
Ripley’s K analysis further confirms this clustered pattern across multiple spatial scales. The observed K values consistently exceeded the upper confidence envelope generated from 999 permutations, demonstrating that clustering occurs not only at local distances but also persists at broader spatial scales. Together, these results provide strong and consistent evidence that the spatial distribution of public-supply wells is not random or dispersed, but instead exhibits significant clustering throughout the study area.
The agreement among all three methods strengthens the reliability of this conclusion. While each method evaluates spatial patterns differently—Quadrat Count through density, ANN through nearest-neighbor distances, and Ripley’s K across multiple distance bands—they all identify the same underlying pattern. This convergence suggests that the clustering is a real and meaningful spatial characteristic of the dataset. The observed clustering is likely influenced by real-world factors such as population density, water demand, and the availability of accessible groundwater resources, which tend to concentrate wells in specific regions, particularly urban areas.
Limitations
Several limitations should be considered when interpreting these results. The study area covers a large geographic extent with a substantial number of data points, which can influence the detection and interpretation of spatial patterns. In addition, the study area was defined using a rectangular bounding box rather than the true boundary of Arizona. This introduces areas without wells outside the actual state boundary, which may affect both density-based and distance-based calculations, particularly in the Quadrat Count and Ripley’s K analyses. Furthermore, the selection of quadrat size (10 × 10 grid) influences the Variance-to-Mean Ratio, as different grid sizes can lead to varying interpretations of clustering.
Edge effects also present a limitation for all distance-based methods, including Average Nearest Neighbor and Ripley’s K. Points located near the boundaries of the study area have fewer neighboring points within the analysis window, which can result in overestimated distances and potential bias in the results. This issue may be further amplified by the use of a rectangular study area instead of the true state boundary. Despite these limitations, the results from all three methods consistently indicate that public-supply wells across Arizona exhibit significant clustering.
Arizona Department of Water Resources. (n.d.). GWSI sites 2024. ArcGIS Open Data Portal. https://gisdata2016-11-18t150447874z-azwater.opendata.arcgis.com/datasets/azwater::gwsi-sites-2024/about
U.S. Geological Survey. (2019, March 1). Public supply wells. https://www.usgs.gov/mission-areas/water-resources/science/public-supply-wells