Data lattice point rarefaction method and device, computer equipment and storage medium

By generating gridded regions in meteorological data processing, selecting representative stations, and performing correlation screening and weighted averaging, the accuracy problems of outliers and sparse areas in meteorological data interpolation are solved, achieving efficient and accurate reflection of meteorological characteristics. It is applicable to meteorological data analysis at different scales and in different regions.

CN120873571APending Publication Date: 2025-10-31SHANGHAI INVESTIGATION DESIGN & RES INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510958186.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing meteorological data interpolation methods suffer from low interpolation accuracy and high computational complexity when dealing with outliers and sparse regions. They also fail to effectively reflect meteorological characteristics, especially in long-term series data where inconsistencies exist, affecting the accuracy of climate change analysis.

Method used

By determining the latitude and longitude range and grid density of the target study area, multiple grid regions are generated. Representative stations that meet the preset conditions are selected. Based on the correlation between the representative stations and other stations, a group of significantly related representative stations is selected. Data from these stations is collected and weighted averaged to obtain the representative results of the specified grid regions.

Benefits of technology

It enables the reduction of data bias, improvement of computational efficiency, accurate reflection of regional meteorological characteristics, and reduction of requirements for technical knowledge and computing resources in meteorological data processing at different scales and regions, providing a reliable interpolation solution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120873571A_ABST
    Figure CN120873571A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of big data processing, and discloses a data grid point rarefaction method and device, computer equipment and a storage medium, and the method comprises the steps: determining the longitude and latitude range and the grid point density of a target research region, and generating a plurality of grid point regions based on the longitude and latitude range and the grid point density; determining a plurality of lattice point positions in the specified lattice point area, and selecting lattice points meeting a preset condition from the plurality of lattice point positions as representative sites; according to the method, a representative site group significantly related to a representative site is screened out based on the correlation between the representative site and other sites except the representative site in a specified grid point region, and a representative result of the specified grid point region is obtained based on data of all representative sites in the representative site group. And calculation is carried out through the screened representative station group, so that the overall characteristics of regional meteorological elements can be accurately reflected while the data processing amount is reduced and the calculation efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of big data processing technology, specifically to a method, apparatus, computer equipment, and storage medium for data grid sparsity reduction. Background Technology

[0002] Existing meteorological data station interpolation and sparsification methods typically incorporate data from all stations into the calculation. However, when outliers exist in the data (e.g., too many default values ​​or extreme values), the interpolation results are easily affected, reducing the accuracy of the interpolation.

[0003] On the other hand, many mainstream methods have high requirements for data preprocessing, which places high demands on the theoretical knowledge and technical reserves of data processing personnel, and may increase the introduced errors; or the computational complexity during interpolation is high, consuming a lot of computing resources and time costs, thus raising the threshold for use.

[0004] Some interpolation methods use a fixed interpolation radius of influence (such as the `obj_anal_ic_Wrap` function in the NCAR Command Language), which cannot adapt to the station distribution characteristics of different regions. For example, in areas with dense station data, a fixed radius of influence may result in overly dense interpolation; while in sparsely populated areas, the influence range of a few stations may be excessively amplified, and the interpolation results may not accurately reflect meteorological characteristics.

[0005] When processing long-term series data, interpolation results may not fully consider the temporal coverage and representativeness of the data. If the interpolation method does not properly select sites, it may lead to inconsistencies in the time scale of the site data, resulting in inaccuracies in the analysis of climate change.

[0006] Current sparse methods employ wavelet denoising and reconstructed signal differences to filter data points. Their core lies in handling noise and importance assessment of single-point time series data, but they fail to address the impact of outliers in multiple station datasets on the selection of representative stations and the calculation of final grid point values ​​in spatial interpolation scenarios. Other sparse methods optimize station layout to cover the population, primarily based on station spatial representativeness (concentration similarity) to maximize population coverage. However, they do not address how to dynamically and robustly select the most representative station group that best reflects the meteorological characteristics (including variability and correlation) of each grid point region, nor how to efficiently calculate the final grid point values.

[0007] Existing technologies mainly focus on the instantaneous or static representativeness of a single point or layout, and do not sufficiently emphasize the role of temporal coverage in ensuring the long-term representativeness of regional meteorological characteristics. Summary of the Invention

[0008] In view of this, the present invention provides a data grid sparsity method, apparatus, computer equipment and storage medium to solve the problems of the influence of data outliers and the interpolation accuracy of sparse data regions in the prior art.

[0009] In a first aspect, the present invention provides a data grid sparsity reduction method, the method comprising:

[0010] Determine the latitude and longitude range and grid density of the target study area, and generate multiple grid regions based on the latitude and longitude range and grid density;

[0011] Determine multiple grid locations within a specified grid area, and select grid locations that meet preset conditions as representative stations from these multiple grid locations;

[0012] Within a specified grid area, based on the correlation between the representative site and other sites besides the representative site, a group of representative sites that are significantly correlated with the representative site is selected.

[0013] Collect data from all representative sites within the representative site cluster, and obtain representative results for a specified grid area based on all representative site data.

[0014] This invention provides a data grid sparsity method that determines the latitude and longitude range and grid density of the target study area, and generates multiple grid regions based on the latitude and longitude range and grid density. This method is applicable to research scenarios of different scales and regions. Whether for small-scale local studies or large-scale global analyses, it can accurately define grid regions. After determining the grid locations within the specified grid regions, grid points that meet preset conditions are selected as representative stations. This allows for the selection of stations that best reflect the meteorological characteristics of the region from numerous grid points, effectively capturing the changing patterns of meteorological elements within the region and reducing data bias caused by unreasonable station selection. Based on the correlation between representative stations and other stations, a group of significantly correlated representative stations is obtained. This selection mechanism ensures that the selected stations more accurately reflect the meteorological characteristics of the grid region, reduces the requirements for theoretical knowledge and computational resources for technical personnel, minimizes the introduction of errors, and provides a more reliable interpolation solution for sparse regions. Data from all stations within the representative station group are collected, and representative results for the specified grid region are obtained accordingly. This process achieves efficient integration and processing of large amounts of raw data. Compared to using data from all stations, calculations are performed using a selected group of representative stations. This reduces the amount of data processing and improves computational efficiency, while also accurately reflecting the overall characteristics of regional meteorological elements. It is a method that can accurately and conveniently convert station data into gridded data, solving the problems of outlier data and interpolation accuracy in sparse areas in existing technologies.

[0015] In one optional implementation, the latitude and longitude range and grid density of the target study area are determined, and multiple grid regions are generated based on the latitude and longitude range and grid density, including:

[0016] Acquire the data characteristics of the target study area, and determine the latitude and longitude range and grid density of the target study area based on the data characteristics. The latitude and longitude range includes the latitude and longitude boundaries, and the grid density includes the grid resolution.

[0017] The target study area is divided into multiple grid regions based on latitude and longitude boundaries and grid resolution.

[0018] This invention provides a data grid sparsity method that determines the latitude and longitude range and grid density by acquiring data characteristics of the target study area, ensuring that the divided grid regions fully match the actual distribution of meteorological data in that area. By dividing the grid regions based on latitude and longitude boundaries and grid resolution, the target study area is systematically divided into grids. This clear grid region division also facilitates the comparison and integration of data from different regions, providing convenience for in-depth analysis of the differences and relationships between meteorological elements in different regions.

[0019] In one optional implementation, determining the locations of multiple grid points within a specified grid region includes:

[0020] Calculate the latitude and longitude coordinates of the center point of the specified grid area, and determine the position of each grid point based on the latitude and longitude coordinates of the center point.

[0021] This invention provides a data grid sparsity method. Latitude and longitude coordinates are a globally universal and accurate geographic positioning method. The method calculates the latitude and longitude coordinates of the center point of a specified grid area and uses these coordinates to determine the location of each grid point, accurately marking the location of each grid point on the Earth's surface. Whether targeting a small area or a global grid area, the uniqueness and accuracy of each grid point's location are guaranteed. This ensures a close correspondence between meteorological data and actual geographical locations, enabling more accurate analysis of the spatial distribution and variation patterns of meteorological elements based on precise grid locations in meteorological research and disaster monitoring.

[0022] In one optional implementation, selecting grid points that meet preset conditions from multiple grid point locations as representative stations includes:

[0023] Collect all available time series data within a specified grid area, and perform data preprocessing on all available time series data to remove missing and outlier values;

[0024] Calculate the standard deviation of the preprocessed time series data, and select the grid point with the largest standard deviation and the time series data coverage that meets the preset requirements as the representative site.

[0025] The present invention provides a data grid sparsity method, which uses a three-in-one screening mechanism to select the station with the largest standard deviation and the time coverage that meets the preset requirements as the initial representative station in each grid area by calculating the standard deviation of the station data, thereby ensuring the representativeness of the data fluctuation characteristics.

[0026] In one optional implementation, within a specified grid area, a group of representative sites that are significantly related to the representative sites are selected based on the correlation between the representative sites and other sites, including:

[0027] Calculate multiple correlation coefficients between a representative station and other stations other than the representative station within a specified grid area;

[0028] All sites that meet the preset significance test threshold are selected from multiple correlation coefficients to form a representative site group that is significantly correlated with the representative site.

[0029] The present invention provides a data grid sparsity method that uses correlation coefficient analysis to analyze the correlation between representative stations and other stations, and combines adaptive significance thresholds to dynamically screen highly correlated stations, avoiding the problem of overly dense or sparse interpolation caused by fixed influence radius, and is especially suitable for areas with uneven station distribution.

[0030] In one optional implementation, data from all representative sites within a representative site group are collected, and representative results for a specified grid area are obtained based on all representative site data, including:

[0031] Collect all time series data that meet the preset significance test threshold within a representative site group as representative site data;

[0032] Calculate the weighted average of all representative site data that meet the preset significance test threshold as the representative result of the specified grid area.

[0033] In one optional implementation, the weighted average of representative station data that meets a preset significance test threshold is calculated as the representative result for the specified grid area, including:

[0034] Obtain the size range of the specified grid area, and then use inverse distance weighting or simple arithmetic mean to perform a weighted average on the representative station data based on the size range of the specified grid area, and obtain the weighted average as the representative result of the specified grid area.

[0035] This invention provides a data grid sparsity method that uses dynamic thresholds to reduce the number of redundant stations involved in the calculation, and combines simple averaging calculation (instead of weighted averaging) to significantly reduce computational complexity; it supports dynamic adjustment of significance test criteria based on meteorological element type (such as temperature / precipitation), and users can customize the threshold range.

[0036] In a second aspect, the present invention provides a data grid sparsity reduction device, the device comprising:

[0037] The grid region generation module is used to determine the latitude and longitude range and grid density of the target study area, and generate multiple grid regions based on the latitude and longitude range and grid density.

[0038] The representative site filtering module is used to determine the locations of multiple grid points in a specified grid area and select grid points that meet preset conditions from the multiple grid point locations as representative sites.

[0039] The representative site cluster calculation module is used to filter out representative site clusters that are significantly related to the representative site within a specified grid area based on the correlation between the representative site and other sites besides the representative site.

[0040] The representative result calculation module is used to collect data from all representative sites within the representative site group and obtain representative results for a specified grid area based on all representative site data.

[0041] Thirdly, the present invention provides a computer device, comprising: a memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to perform the data grid sparsification method described in the first aspect or any corresponding embodiment thereof.

[0042] Fourthly, the present invention provides a computer-readable storage medium storing computer instructions for causing a computer to perform the data grid sparsification method described in the first aspect or any corresponding embodiment thereof.

[0043] Fifthly, the present invention provides a computer program product, including computer instructions for causing a computer to execute the data grid sparsification method described in the first aspect or any corresponding embodiment thereof. Attached Figure Description

[0044] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0045] Figure 1 This is a flowchart illustrating a data grid sparsity reduction method according to an embodiment of the present invention;

[0046] Figure 2This is a flowchart illustrating another data grid sparsity method according to an embodiment of the present invention;

[0047] Figure 3 This is a flowchart illustrating another data grid sparsity method according to an embodiment of the present invention;

[0048] Figure 4 This is a structural block diagram of a data grid sparsification device according to an embodiment of the present invention;

[0049] Figure 5 This is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. Detailed Implementation

[0050] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0051] According to an embodiment of the present invention, a data grid sparsity method embodiment is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0052] This embodiment provides a data grid sparsity reduction method, which can be used in the aforementioned computer equipment. Figure 1 This is a flowchart of a data grid sparsity reduction method according to an embodiment of the present invention, such as... Figure 1 As shown, the process includes the following steps:

[0053] Step S101: Determine the latitude and longitude range and grid density of the target study area, and generate multiple grid regions based on the latitude and longitude range and grid density.

[0054] For example, taking climate characteristics as an example, long-term meteorological observation data of the target study area, such as historical records of elements like temperature, precipitation, and wind speed, are collected to analyze the spatiotemporal distribution characteristics of the data. For instance, the fluctuations in data across different seasons and years, as well as the differences in different locations within the region, are observed. If the meteorological data in the region fluctuates greatly, a smaller grid resolution can be set, such as 0.5°×0.5°; if the data is relatively stable, a larger resolution can be used, such as 2°×2°. Based on the analysis results, the latitude and longitude boundaries (minimum longitude, maximum longitude, minimum latitude, maximum latitude) and grid resolution are determined to clarify the scope of the study area and the density of the grid points.

[0055] Using Geographic Information System (GIS) technology or programming algorithms, a regular grid is generated on the target study area based on defined latitude and longitude boundaries and grid resolution. Using a two-dimensional plane coordinate system as a foundation, starting from the intersection of the minimum longitude and minimum latitude, horizontal and vertical grid lines are sequentially divided according to the grid resolution, forming multiple grid regions of uniform size. Each grid region can be identified by its row and column number within the grid and its corresponding latitude and longitude range.

[0056] Step S102: Determine multiple grid locations in the specified grid area, and select grid locations that meet preset conditions as representative stations.

[0057] Specifically, for each specified grid region, the latitude and longitude coordinates of its geometric center point are calculated. This can be achieved by obtaining the coordinates of the midpoints of the four boundary lines of the grid region and then calculating the average of these midpoint coordinates. Using this center point as a reference, and considering the grid resolution, multiple grid point positions are determined within the grid region according to certain rules (such as equal spacing). For example, if the grid resolution is 1°×1°, a grid point can be determined every 0.2° within the region, forming a regular grid array.

[0058] Collect meteorological station data within a certain radius (e.g., 50 km radius around the grid point) near each grid point, including station observation records and data completeness. Preset conditions may include data completeness (e.g., a data missing rate of less than 20%), observation duration (at least 10 years of continuous observation data), and data variability (by calculating the standard deviation of station data and selecting stations with larger standard deviations to reflect the changing characteristics of meteorological elements within the region). Based on these preset conditions, screen the stations corresponding to the grid points, selecting those that meet the conditions and best represent the meteorological characteristics of the grid point region as representative stations.

[0059] Step S103: Within the specified grid area, based on the correlation between the representative site and other sites besides the representative site, select a group of representative sites that are significantly related to the representative site.

[0060] Specifically, correlation analysis methods such as Pearson correlation coefficient were used to compare the time series data of meteorological elements from representative stations with the data from other non-representative stations within the grid area. For each non-representative station, the correlation coefficient between its data series and that of the representative station was calculated.

[0061] The correlation threshold is set according to the type of meteorological element and research needs. For elements that are close to a normal distribution (such as temperature), a lower significance level (such as 0.1) can be set, with a corresponding correlation coefficient threshold of approximately 0.3-0.4. For elements that are not normally distributed (such as precipitation), a more stringent significance level (such as 0.05 or 0.01) is used, and the correlation coefficient threshold may be above 0.5-0.6. The calculated correlation coefficients are compared with the preset thresholds, and stations with correlation coefficients greater than the preset thresholds are retained. These stations, along with the representative stations, constitute a representative group of significantly correlated stations.

[0062] Step S104: Collect data from all representative sites within the representative site group, and obtain representative results for the specified grid area based on all representative site data.

[0063] Specifically, meteorological data from all stations within a representative station group are collected from meteorological data storage systems or related databases to ensure that the time range of the data is consistent. For missing data, interpolation methods (such as linear interpolation and spline interpolation) can be used to supplement them, and abnormal data are identified and corrected (such as removing data that exceeds 3 times the standard deviation).

[0064] Depending on the size of the grid area and the research needs, an appropriate method is selected to calculate representative results. For example, a weighted average is used to calculate the representative meteorological values ​​for the grid area based on meteorological data.

[0065] The data grid sparsity method provided in this embodiment determines the latitude and longitude range and grid density of the target study area, and generates multiple grid regions based on the latitude and longitude range and grid density. It is applicable to research scenarios of different scales and regions. Whether for small-scale local studies or large-scale global analysis, it can accurately set grid regions. After determining the grid locations in the specified grid regions, grids that meet preset conditions are selected as representative stations. This allows for the selection of stations that best reflect the meteorological characteristics of the region from numerous grids, effectively capturing the changing patterns of meteorological elements within the region and reducing data bias caused by unreasonable station selection. Based on the correlation between representative stations and other stations, a group of significantly correlated representative stations is obtained. This selection mechanism ensures that the selected stations can more accurately reflect the meteorological characteristics of the grid region, reduces the requirements for theoretical knowledge and computing resources of technical personnel, minimizes the introduction of errors, and provides a more reliable interpolation solution for sparse regions. Data from all stations within the representative station group are collected, and representative results for the specified grid region are obtained accordingly. This process achieves efficient integration and processing of a large amount of raw data. Compared to using data from all stations, calculations are performed using a selected group of representative stations. This reduces the amount of data processing and improves computational efficiency, while also accurately reflecting the overall characteristics of regional meteorological elements. It is a method that can accurately and conveniently convert station data into gridded data, solving the problems of outlier data and interpolation accuracy in sparse areas in existing technologies.

[0066] This embodiment provides a data grid sparsity reduction method, which can be used in the aforementioned computer equipment. Figure 2 This is a flowchart of a data grid sparsity reduction method according to an embodiment of the present invention, such as... Figure 2 As shown, the process includes the following steps:

[0067] Step S201: Determine the latitude and longitude range and grid density of the target study area, and generate multiple grid regions based on the latitude and longitude range and grid density.

[0068] Specifically, step S201 includes:

[0069] Step S2011: Obtain the data features of the target study area, and determine the latitude and longitude range and grid density of the target study area based on the data features. The latitude and longitude range includes the latitude and longitude boundaries, and the grid density includes the grid resolution.

[0070] Specifically, in this embodiment, users can customize the latitude and longitude coordinates and grid density of the grid points to be generated according to research needs. For example, based on the climate characteristics and data requirements of the research area, the latitude and longitude range and grid density of the target research area can be set (i.e., setting standards), including latitude and longitude boundaries (minimum longitude Lon).min Longitude of Lon max Minimum latitude Lat min, Maximum latitude Lat max ) and grid resolution (e.g., ΔLon degrees longitude × ΔLat degrees latitude, or equivalent kilometer spacing, such as 100 kilometers).

[0071] Step S2012: Divide the target study area into multiple grid regions based on latitude and longitude boundaries and grid resolution.

[0072] Specifically, based on set standards, a regular grid is generated within a specified latitude and longitude range, determining the location of each grid point. For example, if the latitude and longitude range is selected as global, and the grid resolution is 1°, longitude × 1°, and latitude × 1°, then the globe can be divided into 360 × 180 = 64,800 grids (180°W - 180°E, 90°N - 90°S). Each grid cell (grid area) is uniquely identified by the latitude and longitude coordinates of its center point.

[0073] Step S202: Determine multiple grid locations in the specified grid area, and select grid locations that meet preset conditions as representative stations.

[0074] Specifically, step S202 includes:

[0075] Step S2021: Calculate the latitude and longitude coordinates of the center point of the specified grid area, and determine the position of each grid point based on the latitude and longitude coordinates of the center point.

[0076] Specifically, latitude and longitude coordinates are a globally universal geographic coordinate system. By determining the latitude and longitude coordinates of the center point of a grid cell, each grid area can be precisely located on the Earth's surface. Just as each grid cell on a map is marked with a unique coordinate system, regardless of its location in the world, its position can be accurately found as long as the latitude and longitude of its center point are known. For example, in meteorological data processing, each grid area is associated with corresponding meteorological data. Using the latitude and longitude coordinates of the center point to uniquely identify the grid area facilitates the accurate mapping of collected meteorological data to specific grid points. For instance, when acquiring data such as temperature and precipitation at a certain location, the data can be accurately recorded on the corresponding grid point based on the latitude and longitude of the center point of the grid cell to which that location belongs, facilitating subsequent analysis and calculations.

[0077] Step S2022: Collect all available time series data within the specified grid area, and perform data preprocessing on all available time series data to remove missing values ​​and outliers.

[0078] Specifically, time series data includes, but is not limited to, meteorological data and power grid data.

[0079] Collect data from all available meteorological stations within the selected grid area and perform data cleaning to remove missing and outlier values. Specifically, remove records with an excessively high percentage of missing values ​​(e.g., a single station with a missing value rate > 33%). Perform preliminary outlier detection on the remaining data (e.g., delete data exceeding 3 standard deviations).

[0080] Step S2023: Calculate the standard deviation of the preprocessed time series data, and select the grid point with the largest standard deviation and the time series data coverage that meets the preset requirements as the representative site.

[0081] Specifically, for the meteorological element time series data of each station i within the preprocessed grid area, its standard deviation SD is calculated. i The formula is as follows:

[0082]

[0083] Where, n i Let X be the number of observations at station i. ij For the j-th observation, Let be the average value of station i.

[0084] The site with the largest standard deviation and the required data coverage (covering at least two-thirds of the time period) is selected as the representative site for that grid point. This selection ensures that the variability characteristics of the region at various time scales (such as interdecadal and interannual) can be reflected.

[0085] It should be noted that the time coverage constraint requires that the data from each site cover ≥2 / 3 of the study period, and supports extended constraints (such as data integrity for the first / last 10% of the time period) to ensure consistency of the time scale.

[0086] Step S203: Within the specified grid area, based on the correlation between the representative site and other sites besides the representative site, filter out a group of representative sites that are significantly correlated with the representative site. For details, please refer to [link to details]. Figure 1 Step S103 of the illustrated embodiment will not be described again here.

[0087] Step S204: Collect data from all representative sites within the representative site group, and obtain representative results for the specified grid area based on all representative site data. For details, please refer to [link to relevant documentation]. Figure 1 Step S104 of the illustrated embodiment will not be described again here.

[0088] The data grid sparsity method provided in this embodiment determines the latitude and longitude range and grid density by acquiring the data characteristics of the target study area, ensuring that the divided grid areas fully match the actual distribution of meteorological data in that area. Grid areas are divided based on latitude and longitude boundaries and grid resolution, resulting in an orderly gridded segmentation of the target study area. This clear grid area division facilitates the comparison and integration of data from different regions, providing convenience for in-depth analysis of the differences and connections between meteorological elements in different regions. Through a three-pronged selection mechanism, within each grid area, the standard deviation of the station data is calculated, and the station with the largest standard deviation and the time coverage meeting preset requirements is selected as the initial representative station, ensuring the representativeness of data fluctuation characteristics.

[0089] This embodiment provides a data grid sparsity reduction method, which can be used in the aforementioned computer equipment. Figure 3 This is a flowchart of a data grid sparsity reduction method according to an embodiment of the present invention, such as... Figure 3 As shown, the process includes the following steps:

[0090] Step S301: Determine the latitude and longitude range and grid density of the target study area, and generate multiple grid regions based on the latitude and longitude range and grid density. For details, please refer to [link to relevant documentation]. Figure 2 Step S201 of the illustrated embodiment will not be described again here.

[0091] Step S302: Determine multiple grid point locations within the specified grid area, and select grid points that meet preset conditions from these locations as representative stations. For details, please refer to [link to relevant documentation]. Figure 2 Step S202 of the illustrated embodiment will not be described again here.

[0092] Step S303: Within the specified grid area, based on the correlation between the representative site and other sites besides the representative site, select a group of representative sites that are significantly related to the representative site.

[0093] Specifically, step S303 includes:

[0094] Step S3031: Calculate multiple correlation coefficients between the representative station and other stations other than the representative station within the specified grid area.

[0095] Specifically, after determining representative sites, a correlation analysis is performed, which involves comparing the selected representative sites with other non-representative sites within the grid area. The Pearson correlation coefficient r is used to represent the correlation, as shown in the following formula:

[0096]

[0097] in, and , where are the average values ​​of the data sequences from two different observation stations, and N is the number of observations.

[0098] Step S3032: Select all sites that meet the preset significance test threshold from multiple correlation coefficients as the representative site group that is significantly correlated with the representative site.

[0099] Specifically, based on the correlation analysis results, a dynamic threshold is set, with a significance level of 90% as the basis (degrees of freedom are based on effective degrees of freedom). This threshold is dynamically adjusted to ensure significant correlations exist between the selected stations. The significance level α is dynamically adjusted according to the meteorological element type: for elements with near-normal distributions (such as temperature), a basic α (e.g., 0.10) can be used; for non-normally distributed elements (such as precipitation), a more stringent α (e.g., 0.05 or 0.01) is used; α can also be manually specified based on regional station density or specific needs.

[0100] Screening for highly correlated site groups: Retain all sites that are significantly correlated with the representative site (such as sites whose correlation coefficients pass the 90% significance level test). These sites are statistically significantly correlated with the initial representative site of the grid area and constitute the representative site group of the grid, which will participate in subsequent calculations.

[0101] Step S304: Collect data from all representative sites within the representative site group, and obtain representative results for the specified grid area based on all representative site data.

[0102] Specifically, step S304 includes:

[0103] Step S3041: Collect all time series data that meet the preset significance test threshold within the representative site group as representative site data.

[0104] Specifically, data from all representative sites within the representative site group that meet the significance test are collected to ensure consistent time series coverage of the data.

[0105] Furthermore, based on the previously set significance test thresholds (correlation coefficient threshold and p-value threshold), the time series data of each station within the representative station group are filtered. All station data that meet the condition of "correlation coefficient greater than the set threshold and p-value less than the corresponding significance level (e.g., 0.05)" are retrieved from the database or file system storing meteorological data.

[0106] The p-value threshold is an important indicator used to determine whether the correlation between data is significant. The p-value is a probability value that reflects the probability of observing the current data result or a more extreme result if the null hypothesis (usually assuming there is no correlation between the two variables) is true.

[0107] In correlation analysis, setting a p-value threshold aims to determine which correlations between sites are genuine and not due to random factors. When the calculated p-value between two sites is less than the pre-set p-value threshold (e.g., 0.05 or 0.01), the null hypothesis can be rejected, indicating a significant correlation between the two sites. Conversely, if the p-value is greater than the threshold, the null hypothesis cannot be rejected, meaning a significant correlation cannot be established.

[0108] For example, setting the p-value threshold to 0.05 means that if the p-value calculated for a site's correlation with a representative site is 0.03 (less than 0.05), the correlation between that site and the representative site is statistically significant, and that site will be retained for subsequent analysis. Conversely, if the p-value is 0.06 (greater than 0.05), the correlation between that site and the representative site is not significant, and the site may be excluded. By setting an appropriate p-value threshold, sites that are significantly correlated with the representative site can be effectively screened, thereby improving the reliability of the data and the validity of the analysis results.

[0109] Step S3042: Calculate the weighted average of all representative site data that meet the preset significance test threshold as the representative result of the specified grid area.

[0110] In some optional implementations, step S3042 above includes:

[0111] Step a: Obtain the size range of the specified grid area, and based on the size range of the specified grid area, use inverse distance weighting or simple arithmetic mean to perform a weighted average on the representative station data to obtain the weighted average as the representative result of the specified grid area.

[0112] For example, the data from each representative station can be weighted and averaged to obtain the final meteorological value for a specified grid area.

[0113] The formula for calculating the weighted average is as follows:

[0114]

[0115] Among them, Z g To provide representative results for a specified grid area, the values ​​can be meteorological element values, Z. i Let w be the meteorological element value of the i-th representative station. i The weights are assigned to representative sites, where M is the number of representative sites with a preset significance test threshold. The weights can be set according to actual conditions. If the grid spatial range is large (e.g., 5° longitude × 5° latitude, or 1000 km × 1000 km), inverse distance weighting (IDW) can be used, with weights w...i Calculated based on the geographical distance from station i to the grid center; if the grid area is small, a simple arithmetic mean can be used.

[0116] The IDW method is a geographic distance-based interpolation method that calculates the interpolated value of the target point by weighting the values ​​of the observation points with the inverse of the distances between the observation points and the target point. The basic formula is as follows:

[0117]

[0118] Where Z(x,y) is the predicted value of the target point, z i Let d be the value of the observation point. i denoted as , where is the distance between the target point and the observation point, and n is the number of observation points involved in the interpolation.

[0119] The data grid sparsity method provided in this embodiment uses correlation coefficient analysis to analyze the correlation between representative stations and other stations, and combines it with an adaptive significance threshold to dynamically select highly correlated stations. This avoids the problem of overly dense or sparse interpolation caused by a fixed influence radius, and is particularly suitable for areas with uneven station distribution. The use of a dynamic threshold reduces the number of redundant stations involved in the calculation, and combined with a simple averaging calculation (replacing the weighted average), significantly reduces computational complexity. It supports dynamically adjusting the significance test criteria based on meteorological element type (such as temperature / precipitation), and users can customize the threshold range.

[0120] As one or more specific application embodiments of the present invention, the data grid sparsity method provided by the present invention will be further described in detail as follows:

[0121] 1. Determine the latitude and longitude range and the grid density:

[0122] In this embodiment, the user can customize the latitude and longitude coordinates and density of the grid points to be generated according to research needs. The process includes the following steps:

[0123] 1) Setting Standards: Based on the climate characteristics and data requirements of the study area, set the latitude and longitude range and grid density of the target area, including latitude and longitude boundaries (minimum longitude Lon). min Longitude of Lon max Minimum latitude Lat min Maximum latitude Lat max ) and grid resolution (e.g., ΔLon degrees longitude × ΔLat degrees latitude, or equivalent kilometer spacing, such as 100 kilometers).

[0124] 2) Based on the set standards, generate a regular grid within the specified latitude and longitude range and determine the location of each grid point. For example, if the latitude and longitude range is selected as global, and the grid resolution is 1°, longitude × 1°, and latitude × 1°, then the globe can be divided into 360 × 180 = 64,800 grids (180°W - 180°E, 90°N - 90°S). Each grid cell (grid area) is uniquely identified by the latitude and longitude coordinates of its center point.

[0125] 2. Site selection based on the largest standard deviation:

[0126] Within each designated grid area, perform the following steps to select representative sites:

[0127] 1) Data Preprocessing: Collect data from all available meteorological stations within the selected grid area and perform data cleaning to remove missing and outlier values. Specifically, remove records with an excessively high percentage of missing values ​​(e.g., a single station with a missing value rate > 33%). Perform preliminary outlier detection on the remaining data (e.g., delete data exceeding 3 standard deviations).

[0128] 2) For the meteorological element time series data of each station i within the preprocessed grid area, calculate its standard deviation SD. i The formula is as follows:

[0129]

[0130] Where, n i Let X be the number of observations at station i. ij For the j-th observation, Let be the average value of station i.

[0131] 3) Select the station with the largest standard deviation and the required data coverage (covering at least 2 / 3 of the time period) as the representative station for that grid point. This selection ensures that the variability characteristics of the region at various time scales (such as interdecadal and interannual) can be reflected.

[0132] 3. Correlation Analysis: After identifying representative sites, a correlation analysis is performed. The specific steps are as follows:

[0133] 1) Calculate correlation: Perform correlation analysis between the selected representative site and other sites within the grid area, using the Pearson correlation coefficient r formula:

[0134]

[0135] in, and , where are the average values ​​of the data sequences from two different observation stations, and N is the number of observations.

[0136] 2) Dynamic Threshold Optimization: Based on the correlation analysis results, a dynamic threshold is set, using a 90% significance level as a baseline (degrees of freedom are based on effective degrees of freedom). This threshold is dynamically adjusted to ensure significant correlations exist between the selected stations. The significance level α is dynamically adjusted according to the meteorological element type: For elements with near-normal distributions (such as temperature), a basic α (e.g., 0.10) can be used. For non-normally distributed elements (such as precipitation), a more stringent α (e.g., 0.05 or 0.01) is used. α can also be manually specified based on regional station density or specific needs.

[0137] 3) Screening highly correlated site groups: Retain all sites that are significantly correlated with the representative site (such as sites whose correlation coefficients pass the 90% significance level test). These sites are statistically significantly correlated with the initial representative site of the grid area and constitute the representative site group of the grid, which will participate in subsequent calculations.

[0138] 4. Data averaging:

[0139] 1) Collect data from all representative sites within the representative site group that meet the significance test, ensuring consistent time series coverage of the data.

[0140] 2) Data averaging: A weighted average is calculated from the data from each representative station to obtain the final meteorological value for the target grid point. The formula for calculating the weighted average is:

[0141]

[0142] Among them, Z g To provide representative results for a specified grid area, the values ​​can be meteorological element values, Z. i Let w be the meteorological element value of the i-th representative station. i The weights are assigned to representative sites, where M is the number of representative sites with a preset significance test threshold. The weights can be set according to actual conditions. If the grid spatial range is large (e.g., 5° longitude × 5° latitude, or 1000 km × 1000 km), inverse distance weighting (IDW) can be used, with weights w... i Calculated based on the geographical distance from station i to the grid center; if the grid area is small, a simple arithmetic mean can be used.

[0143] The spatial distribution analysis results of the main mode of the number of days with high temperatures above 35°C in the Northern Hemisphere during summer are obtained by using the built-in function obj_anal_ic_Wrap of NCL (NCAR Command Language, a meteorological data processing language) and the data grid sparsification method provided in this embodiment to interpolate the station data to the grid and then performing Empirical Orthogonal Function Analysis (EOF).

[0144] In this embodiment of the invention, a 2.5°×2.5° resolution is used as the target grid point. The station with the largest standard deviation within each grid point is selected, and the station whose correlation coefficient passes the 90% significance test is used as the representative station. The average value of the data of the representative station for each grid point is used as the result value of that grid point.

[0145] Although the two methods exhibit similar overall distribution patterns, significant differences exist in specific regions. The interpolation scheme using the NCL built-in function `obj_anal_ic_Wrap`, due to its uniform interpolation influence radius, incorrectly amplifies the impact on a few stations in sparsely populated areas. These stations only represent a portion of the surrounding area within a specific timeframe, not the majority of the target area. Furthermore, when the number of effective stations in the target area is small, the interpolation scheme using `obj_anal_ic_Wrap` unreasonably amplifies the impact on stations within the target area. For densely populated areas, the results of this invention are similar to those using the NCL built-in function `obj_anal_ic_Wrap`, but the method of this invention more accurately reflects the actual situation in that area.

[0146] The above demonstrates the advantages of the method provided by the embodiments of the present invention in terms of result accuracy. In addition, the method provided by the embodiments of the present invention has a small data processing volume and a low professional threshold. It can be considered that this method is a suitable solution for sparsifying long-term station meteorological data with uneven spatial distribution into gridded data.

[0147] The data grid sparsity method provided in this embodiment introduces an adaptive threshold optimization mechanism to address the differences in various meteorological element types (such as temperature and precipitation). This process includes:

[0148] 1) Site data screening: Select site data that covers at least 2 / 3 of the time period. This can be adjusted flexibly according to the research content, such as adding data that contains both the first 10% and the last 10% of the time series, to meet the requirements of climate change research.

[0149] 2) Significance Test: Different significance test criteria are set according to the characteristics of different types of meteorological data. For example, the distribution characteristics of precipitation data may differ from those of temperature data. Therefore, using targeted significance test criteria can improve the reliability of the results when conducting correlation analysis.

[0150] 3) Dynamic Adjustment: During data analysis, the significance test criteria (e.g., confidence level 80%–95%) are dynamically adjusted based on actual conditions and model feedback to adapt to regional station density differences and ensure the adaptability of interpolation. If gridded data is not required, the defined region can be compared to a grid, and the station data within that region can be processed to obtain representative results for that region.

[0151] Based on the above analysis, the data grid sparsity reduction method provided in this embodiment of the invention has the following significant advantages:

[0152] 1) Enhanced robustness against outliers: Through a pre-implemented "standard deviation screening + time coverage constraint" mechanism, stations with poor data quality or insufficient time representativeness are proactively excluded within each grid region before interpolation calculation. The standard deviation screening mechanism automatically excludes stations with data anomalies (such as excessive default values ​​or extreme values), reducing the sensitivity of the interpolation results to noise. Existing technologies (such as Optimal Interpolation, OI) rely on complex preprocessing and remain susceptible to extreme value interference. Compared to existing technologies that judge the difference between wavelet denoising and reconstructed signals on single-point time series, this invention more directly avoids the participation of inferior data in the final calculation of regional representative values.

[0153] 2) Significantly improved computational efficiency: The use of dynamic thresholds reduces the number of redundant sites involved in the calculation, and combined with simple averaging (replacing weighted averaging), the computational complexity is greatly reduced; it overcomes the shortcomings of OI interpolation method, which requires the construction of a covariance matrix, and Kriging method, which requires fitting a semi-variogram function, resulting in high computational resource consumption.

[0154] 3) Better adaptability in sparse areas: By adaptively adjusting the influence range of stations through correlation analysis and dynamic thresholds, it avoids the problem of overly dense or sparse interpolation caused by a fixed influence radius, making it particularly suitable for areas with uneven station distribution. In sparse station areas, stations that are far apart but still have significant statistical correlations are allowed to be retained, making the regional representativeness more reasonable.

[0155] 4) Time consistency guarantee: It is mandatory to require that the time coverage of the site is ≥2 / 3 to ensure the consistency of the time scale of the interpolation results, while existing methods often ignore the differences in time representativeness.

[0156] 5) Greater parameter flexibility: It supports dynamic adjustment of significance test criteria based on meteorological element type (such as temperature / precipitation), and users can customize the threshold range, while traditional methods (such as inverse distance weighting, IDW, and Kriging) rely on fixed parameters.

[0157] This embodiment also provides a data grid sparsity reduction device for implementing the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0158] This embodiment provides a data grid sparsity reduction device, such as... Figure 4 As shown, it includes:

[0159] The grid region generation module 401 is used to determine the latitude and longitude range and grid density of the target study area, and generate multiple grid regions based on the latitude and longitude range and grid density.

[0160] The representative site filtering module 402 is used to determine the locations of multiple grid points in a specified grid area and select grid points that meet preset conditions from the multiple grid point locations as representative sites.

[0161] The representative site group calculation module 403 is used to filter out representative site groups that are significantly related to the representative site within a specified grid area based on the correlation between the representative site and other sites besides the representative site.

[0162] The representative result calculation module 404 is used to collect data from all representative sites within the representative site group and obtain representative results for a specified grid area based on all representative site data.

[0163] In some optional implementations, the grid region generation module 401 includes:

[0164] The latitude and longitude range and grid density determination unit is used to acquire data features of the target study area and determine the latitude and longitude range and grid density of the target study area based on the data features. The latitude and longitude range includes the latitude and longitude boundaries, and the grid density includes the grid resolution.

[0165] Grid region division unit, used to divide the target study area into multiple grid regions based on latitude and longitude boundaries and grid resolution.

[0166] In some alternative implementations, the representative site filtering module 402 includes:

[0167] A grid point location determination unit is used to calculate the latitude and longitude coordinates of the center point of a specified grid point area, and determine the location of each grid point based on the latitude and longitude coordinates of the center point. In some optional embodiments, the representative site filtering module 402 further includes:

[0168] The data preprocessing unit is used to collect all available time series data within a specified grid area and perform data preprocessing on all available time series data to remove missing and outlier values.

[0169] The representative site selection unit is used to calculate the standard deviation of the preprocessed time series data and select the grid points with the largest standard deviation and the time series data coverage that meet the preset requirements as representative sites.

[0170] In some alternative implementations, the representative site cluster calculation module 403 includes:

[0171] The correlation coefficient calculation unit is used to calculate multiple correlation coefficients between a representative station and other stations other than the representative station within a specified grid area.

[0172] The representative site group screening unit is used to select all sites that meet the preset significance test threshold from multiple correlation coefficients as the representative site group that is significantly correlated with the representative site.

[0173] In some optional implementations, the representative result calculation module 404 includes:

[0174] The representative site data collection unit is used to collect all time series data that meet the preset significance test threshold within the representative site group as representative site data.

[0175] The representative result calculation unit is used to calculate the weighted average of all representative site data that meet the preset significance test threshold as the representative result of the specified grid area.

[0176] In some optional implementations, the representative result calculation unit includes:

[0177] The weighted average calculation subunit is used to obtain the size range of the specified grid area, and to perform a weighted average on the representative station data based on the size range of the specified grid area using inverse distance weighting or simple arithmetic mean, so as to obtain the weighted average as the representative result of the specified grid area.

[0178] Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.

[0179] In this embodiment, the data grid sparsity device is presented in the form of a functional unit. Here, a unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.

[0180] This invention also provides a computer device having the above-described features. Figure 4 The data grid sparsification device shown.

[0181] Please see Figure 5 , Figure 5 This is a schematic diagram of the structure of a computer device provided in an optional embodiment of the present invention, such as... Figure 5 As shown, the computer device includes one or more processors 10, memory 20, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components communicate with each other via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the computer device, including instructions stored in or on memory to display graphical information of a GUI on external input / output devices (such as display devices coupled to the interfaces). In some alternative implementations, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple computer devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 5 Take a processor 10 as an example.

[0182] Processor 10 may be a central processing unit, a network processor, or a combination thereof. Processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CAMP), a field-programmable gate array (FPGA), a general-purpose array logic (GDA), or any combination thereof.

[0183] The memory 20 stores instructions executable by at least one processor 10 to cause the at least one processor 10 to perform the method shown in the above embodiments.

[0184] The memory 20 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the computer device. Furthermore, the memory 20 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some alternative embodiments, the memory 20 may optionally include memory remotely located relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0185] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid-state drive; the memory 20 may also include a combination of the above types of memory.

[0186] The computer device also includes an input device 30 and an output device 40. The processor 10, memory 20, input device 30, and output device 40 can be connected via a bus or other means. Figure 5 Taking the example of a connection between China and Israel via a bus.

[0187] Input device 30 can receive input numerical or character information, and generate key signal inputs related to user settings and function control of the computer device, such as a touchscreen, keypad, mouse, trackpad, touchpad, joystick, one or more mouse buttons, trackball, joystick, etc. Output device 40 may include display devices, auxiliary lighting devices (e.g., LEDs), and haptic feedback devices (e.g., vibration motors). The aforementioned display devices include, but are not limited to, liquid crystal displays, light-emitting diodes, displays, and plasma displays. In some alternative embodiments, the display device may be a touchscreen.

[0188] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods shown in the above embodiments.

[0189] A portion of this invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the invention through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.

[0190] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A data grid sparsity reduction method, characterized in that, The method includes: Determine the latitude and longitude range and grid density of the target study area, and generate multiple grid regions based on the latitude and longitude range and grid density; Determine multiple grid locations within a specified grid area, and select grid locations that meet preset conditions as representative stations from these multiple grid locations; Within a specified grid area, based on the correlation between the representative site and other sites besides the representative site, a group of representative sites that are significantly correlated with the representative site is selected. Collect data from all representative sites within the representative site group, and obtain representative results for a specified grid area based on all representative site data.

2. The method according to claim 1, characterized in that, The process of determining the latitude and longitude range and grid density of the target study area, and generating multiple grid regions based on the latitude and longitude range and grid density, includes: Data features of the target study area are acquired, and the latitude and longitude range and grid density of the target study area are determined based on the data features. The latitude and longitude range includes the latitude and longitude boundaries, and the grid density includes the grid resolution. The target study area is divided into multiple grid regions based on latitude and longitude boundaries and grid resolution.

3. The method according to claim 1, characterized in that, Determining the positions of multiple grid points within a specified grid region includes: Calculate the latitude and longitude coordinates of the center point of the specified grid area, and determine the position of each grid point based on the latitude and longitude coordinates of the center point.

4. The method according to claim 1, characterized in that, Selecting grid points that meet preset conditions from the plurality of grid point locations as representative stations includes: Collect all available time series data within a specified grid area, and perform data preprocessing on all available time series data to remove missing and outlier values; Calculate the standard deviation of the preprocessed time series data, and select the grid point with the largest standard deviation and the time series data coverage that meets the preset requirements as the representative site.

5. The method according to claim 1, characterized in that, The step of filtering out a representative group of sites that are significantly related to the representative sites within a specified grid area based on the correlation between the representative sites and other sites includes: Calculate multiple correlation coefficients between a representative station and other stations other than the representative station within a specified grid area; All sites that meet the preset significance test threshold are selected from multiple correlation coefficients to form a representative site group that is significantly correlated with the representative site.

6. The method according to claim 1, characterized in that, The step of collecting data from all representative sites within the representative site group and obtaining representative results for a specified grid area based on all representative site data includes: Collect all time series data that meet the preset significance test threshold within the representative site group as representative site data; Calculate the weighted average of all representative site data that meet the preset significance test threshold as the representative result of the specified grid area.

7. The method according to claim 6, characterized in that, The calculation of the weighted average of all representative site data that meet the preset significance test threshold as the representative result of the specified grid area includes: Obtain the size range of the specified grid area, and then use inverse distance weighting or simple arithmetic mean to perform a weighted average on the representative station data based on the size range of the specified grid area, and obtain the weighted average as the representative result of the specified grid area.

8. A data grid sparsity reduction device, characterized in that, The device includes: The grid region generation module is used to determine the latitude and longitude range and grid density of the target study area, and generate multiple grid regions based on the latitude and longitude range and grid density. The representative site filtering module is used to determine the locations of multiple grid points in a specified grid area, and select grid points that meet preset conditions from the multiple grid point locations as representative sites. The representative site cluster calculation module is used to filter out representative site clusters that are significantly related to the representative site within a specified grid area based on the correlation between the representative site and other sites besides the representative site. The representative result calculation module is used to collect data from all representative sites within the representative site group and obtain representative results for a specified grid area based on all representative site data.

9. A computer device, characterized in that, include: A memory and a processor are communicatively connected, the memory stores computer instructions, and the processor executes the computer instructions to perform the data grid sparsification method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a computer to perform the data grid sparsification method according to any one of claims 1 to 7.