A machine learning-based geographic data collection method and system

CN121958443BActive Publication Date: 2026-08-28JIANGSU LIANYUNGANG GEOLOGY ENG RECONNAISSANCE INST
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610078074.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-21
Publication Date
2026-08-28
Estimated Expiration
2046-01-21

AI Technical Summary

Technical Problem

[0005]因此,本发明提供了一种基于机器学习的地理数据采集方法解决现有技术因持续强化历史采集偏差而导致地理数据长期分布不均的问题

Benefits of technology

[0016] The beneficial effects of this invention are as follows: by acquiring and analyzing historical collection records of the target area to quantify the collection frequency distribution, historically low-coverage sub-regions are identified, and deviation correction guidance factors are generated for candidate locations based on their spatial distribution and the change characteristics of associated geographic elements. When responding to collection task requests, the guidance factors and the information gain scores given by the pre-trained change prediction model are combined to generate an optimized collection task sequence, drive the device to perform collection, and update the prediction model using new data. Based on the updated model, the geographic status of each sub-region is re-inferred to generate the latest map. By introducing a deviation correction mechanism, the efficiency of information acquisition and the balance of spatial coverage are optimized in path planning, improving historical collection deviations and enhancing the comprehensiveness and long-term application value of the geographic database.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121958443B_ABST
    Figure CN121958443B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on machine learning geographic data acquisition method and system, it is related to geographic data acquisition field, including, the historical acquisition record set of target geographic area is obtained, the acquisition frequency distribution of different sub-regional in target geographic area is obtained to the historical acquisition record set statistics;According to acquisition frequency distribution, the historical low-coverage sub-region of acquisition frequency lower than first threshold value is identified;Based on the spatial distribution characteristics of historical low-coverage sub-region and associated geographic feature variation characteristics, generate deviation correction guide factor for each candidate collection position;In response to current collection task request, according to deviation correction guide factor and the information gain prediction score of candidate collection position obtained from pre-trained geographic feature variation prediction model.The application introduces deviation correction mechanism, cooperates optimization information acquisition efficiency and spatial coverage balance in path planning, improves historical acquisition deviation, improves the comprehensiveness and long-term application value of geographic database.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of geographic data acquisition, and in particular to a geographic data acquisition method and system based on machine learning. Background Technology

[0002] In the field of machine learning-based geographic data acquisition, existing technologies primarily employ active learning strategies to optimize acquisition paths. Typical methods train geographic feature change prediction models to assess the information value of different regions, prioritizing acquisition of areas with a high probability of predicted change. For example, remote sensing image change detection technology based on convolutional neural networks can identify urban expansion hotspots, guiding acquisition equipment to efficiently acquire the latest geographic data. These methods effectively improve the targeting of data acquisition, avoid redundant observations in stable areas, and have achieved good application results in geographic information updates and environmental monitoring.

[0003] Existing methods may amplify spatial coverage biases formed by historical data collection activities during continuous optimization. Due to limitations such as data collection equipment deployment and transportation conditions, the actual collection frequency is often unevenly distributed in space. Existing technologies, with the primary goal of maximizing information gain, tend to repeatedly access historical hotspots while paying insufficient attention to areas with lower collection frequencies but potential importance. Although this strategy is highly efficient in the short term, it may lead to spatial representation biases in geographic databases in the long term, affecting the comprehensiveness and fairness of data applications, especially in application scenarios that require balanced data across the entire domain. Summary of the Invention

[0004] In view of the aforementioned existing problems, the present invention is proposed.

[0005] Therefore, this invention provides a machine learning-based geographic data acquisition method to solve the problem of uneven distribution of geographic data in the long term caused by the continuous reinforcement of historical acquisition bias in existing technologies.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: In a first aspect, the present invention provides a geographic data acquisition method based on machine learning, which includes: acquiring a set of historical acquisition records for a target geographic area; and statistically analyzing the set of historical acquisition records to obtain the acquisition frequency distribution of different sub-regions within the target geographic area. Based on the sampling frequency distribution, historically low-coverage sub-regions with sampling frequencies below the first threshold were identified; Based on the spatial distribution characteristics of historically low-coverage sub-regions and the associated changes in geographical elements, bias correction guidance factors are generated for each candidate collection location. In response to the current data collection task request, a data collection task sequence is generated based on the bias correction guidance factor and the information gain prediction score of the candidate data collection locations obtained from the pre-trained geographic feature change prediction model. Control the currently available acquisition devices to move according to the acquisition task sequence, perform geographic data acquisition operations, obtain the current batch of raw geographic data, and update the pre-trained geographic element change prediction model; The status of geographic elements in each sub-region within the target geographic area is re-inferred to generate an updated geographic element map.

[0007] As a preferred embodiment of the machine learning-based geographic data acquisition method of the present invention, obtaining the historical collection record set of the target geographic area includes the following steps: Define and set the boundary coordinates of the target geographic area and the start and end times of the preset historical period; Based on the boundary coordinates of the target geographic area and the start and end times of the preset historical period, all matching historical collection records are queried from the storage unit to form a set of historical collection records.

[0008] As a preferred embodiment of the machine learning-based geographic data acquisition method of the present invention, the following steps are included: Statistical analysis of historical data collection sets to obtain the data collection frequency distribution of different sub-regions within the target geographic area: The target geographic area is divided according to the set rules to form a division scheme of different sub-regions within the target geographic area; Based on the division scheme of different sub-regions within the target geographical area, the collection location of each historical collection record in the historical collection record set is statistically analyzed to determine the sub-region to which the historical collection record belongs, thus obtaining the attribution result of the historical collection record. Based on the attribution result of each historical collection record in the historical collection record set, the total number of collections in different sub-regions within the target geographic area within the preset historical period is calculated using the accumulation method; Based on the frequency distribution of data collection in different sub-regions within the target geographic area, the frequency distribution of data collection in different sub-regions within the target geographic area is obtained by comparing the total number of data collections in different sub-regions within the target geographic area with the total duration of a preset historical period.

[0009] As a preferred embodiment of the machine learning-based geographic data acquisition method of the present invention, the method includes the following steps: Identifying historically low-coverage sub-regions with acquisition frequencies below a first threshold based on the acquisition frequency distribution: Obtain the sampling frequency distribution of different sub-regions within the target geographic area, and extract the sampling frequency value of each different sub-region within the target geographic area from the sampling frequency distribution of different sub-regions within the target geographic area; Based on a preset first threshold, the collection frequency value of different sub-regions within each target geographic region is compared with the first threshold, and the collection frequency value is collected if it is lower than the first threshold. A list of historically low-coverage sub-regions is formed based on different sub-regions within the target geographic area where the collection frequency value is lower than the first threshold.

[0010] As a preferred embodiment of the machine learning-based geographic data acquisition method of the present invention, the method includes the following steps: generating bias correction guidance factors for each candidate acquisition location based on the spatial distribution characteristics of historically low-coverage sub-regions and the associated geographic element change characteristics. Read the list of historical low-coverage sub-regions and extract the geographic range identifier and corresponding collection frequency value for each historical low-coverage sub-region in the list. Based on the geographic extent identifier of each historical low-coverage sub-region, the clustering, dispersion and geometric center location of the historical low-coverage sub-region within the target geographic region are analyzed to form the spatial distribution characteristics of the historical low-coverage sub-region. Query the historical collection record set to obtain historical geographic element data arranged in time series corresponding to each historical low-coverage sub-region; Time series analysis is performed on the historical geographic element data corresponding to each historically low-coverage sub-region to extract indicators describing the intensity or frequency of changes in geographic elements, forming related geographic element change characteristics. Based on the spatial distribution characteristics of historically low-coverage sub-regions and the changes in associated geographic elements, fusion rules are defined; Based on the fusion rules, a comprehensive judgment is made on the spatial proximity relationship between candidate collection locations and historical low-coverage sub-regions, as well as the changes in associated geographical elements. Each candidate collection location is assigned a quantitative result that represents the value of actively correcting historical coverage deviations, thus obtaining a deviation correction guidance factor.

[0011] As a preferred embodiment of the machine learning-based geographic data acquisition method of the present invention, in response to the current acquisition task request, an acquisition task sequence is generated based on the deviation correction guidance factor and the information gain prediction score of the candidate acquisition location obtained from the pre-trained geographic feature change prediction model, including the following steps: Receive the current acquisition task request constrained by the status of currently available acquisition devices; The feature information of each candidate collection location is input into the pre-trained geographic feature change prediction model, and the information gain prediction score of each candidate collection location is output. Based on the deviation correction guidance factor and the information gain prediction score, a comprehensive utility value for each candidate acquisition location is constructed according to a preset weight. Based on the comprehensive utility value of all candidate acquisition locations, and combined with the movement cost and task capacity of currently available acquisition devices, the acquisition task sequence of currently available acquisition devices is planned with the goal of maximizing the total comprehensive utility value.

[0012] As a preferred embodiment of the machine learning-based geographic data acquisition method of the present invention, the method includes the following steps: controlling the currently available acquisition devices to move according to the acquisition task sequence, performing geographic data acquisition operations, and obtaining the current batch of raw geographic data. Based on the acquisition positions and order assigned to the currently available acquisition devices in the acquisition task sequence, generate the movement path instructions for the currently available acquisition devices; The movement path command is sent to the navigation control unit of the currently available acquisition device, and the navigation control unit drives the currently available acquisition device to move to each acquisition position in sequence. Once the currently available data acquisition device reaches the acquisition location specified in the acquisition task sequence, the acquisition action set in the acquisition task sequence is triggered, and the sensors on the currently available data acquisition device are operated to perform geographic data acquisition operations to obtain the current batch of raw geographic data.

[0013] As a preferred embodiment of the machine learning-based geographic data acquisition method of the present invention, updating the pre-trained geographic feature change prediction model includes the following steps: Using the original geographic data of the current batch and the corresponding collection location and time, construct a training sample set for updating the geographic element change prediction model; The training sample set is input into the pre-trained geographic feature change prediction model, and the incremental learning algorithm is executed to update the internal parameters of the pre-trained geographic feature change prediction model, resulting in the updated geographic feature change prediction model.

[0014] As a preferred embodiment of the machine learning-based geographic data acquisition method of the present invention, the method involves: re-inferring the geographic feature status of each sub-region within the target geographic region to generate an updated geographic feature map, including the following steps: The latest representative feature data of each sub-region within the target geographic area are input into the updated geographic element change prediction model, and the updated geographic element change prediction model outputs the geographic element status inference results of each sub-region within the target geographic area. The updated geographic feature map is generated by integrating the inferred geographic feature status results of each sub-region within the target geographic region with the geographic feature information confirmed in the current batch of original geographic data.

[0015] Secondly, the present invention provides a geographic data acquisition system based on machine learning, including a data processing module, which acquires a set of historical acquisition records for a target geographic area, performs statistics on the set of historical acquisition records, and obtains the acquisition frequency distribution of different sub-regions within the target geographic area. The identification module identifies historically low-coverage sub-regions with collection frequencies below a first threshold based on the collection frequency distribution. The deviation correction factor generation module generates deviation correction guidance factors for each candidate collection location based on the spatial distribution characteristics of historical low-coverage sub-regions and the related changes in geographic elements. The sequence generation module, in response to the current data collection task request, generates a data collection task sequence based on the deviation correction guidance factor and the information gain prediction score of the candidate data collection locations obtained from the pre-trained geographic feature change prediction model. The execution module controls the currently available acquisition devices to move according to the acquisition task sequence, performs geographic data acquisition operations, obtains the current batch of raw geographic data, and updates the pre-trained geographic element change prediction model. The inference module re-infers the status of geographic elements in each sub-region within the target geographic area and generates an updated geographic element map.

[0016] The beneficial effects of this invention are as follows: by acquiring and analyzing historical collection records of the target area to quantify the collection frequency distribution, historically low-coverage sub-regions are identified, and deviation correction guidance factors are generated for candidate locations based on their spatial distribution and the change characteristics of associated geographic elements. When responding to collection task requests, the guidance factors and the information gain scores given by the pre-trained change prediction model are combined to generate an optimized collection task sequence, drive the device to perform collection, and update the prediction model using new data. Based on the updated model, the geographic status of each sub-region is re-inferred to generate the latest map. By introducing a deviation correction mechanism, the efficiency of information acquisition and the balance of spatial coverage are optimized in path planning, improving historical collection deviations and enhancing the comprehensiveness and long-term application value of the geographic database. Attached Figure Description

[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a flowchart of a machine learning-based geographic data acquisition method.

[0019] Figure 2 This is a schematic diagram of a machine learning-based geographic data acquisition system. Detailed Implementation

[0020] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0021] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0022] Secondly, the term "one embodiment" or "example" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the invention. The appearance of an embodiment in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that mutually excludes other embodiments.

[0023] Reference Figures 1-2 This is one embodiment of the present invention, which provides a geographic data acquisition method based on machine learning, including the following steps: S1. Obtain the historical collection records of the target geographic area.

[0024] S1.1 Define and set the boundary coordinates of the target geographic area and the start and end times of the preset historical period.

[0025] Furthermore, the boundary coordinates of the target geographic area are usually obtained and set by receiving the latitude and longitude sequence of the geofence polygon vertices input by the user or by specifying a known administrative division code, such as specifying the administrative boundary of Beijing or giving a rectangular area defined by four latitude and longitude points. This defines a clear and unique operating range for all subsequent spatial analyses, ensuring that frequency statistics and spatial identification have a unified geographic benchmark.

[0026] Specifically, the start and end times of the preset historical period are defined by receiving the start and end date and time strings input by the user or by selecting a preset time window template. This time range setting defines the range of historical data to be analyzed from a time dimension, so that the calculation of the collection frequency is based on a stable and comparable time span. The operation of defining and setting the boundary coordinates of the target geographic area and the start and end times of the preset historical period establishes a clear spatial and temporal analysis framework for the entire method, avoiding inconsistent or incomparable statistical results due to ambiguity in the range.

[0027] S1.2. Based on the boundary coordinates of the target geographical area and the start and end times of the preset historical period, query all matching historical collection records from the storage unit to form a historical collection record set.

[0028] Furthermore, based on the pre-set boundary coordinates of the target geographic area and the start and end times of the preset historical period, a structured query language query is initiated to the database or data warehouse storing historical collection records. This query statement includes spatial filtering conditions and temporal filtering conditions. The spatial filtering conditions use spatial relational functions to determine whether the collection location coordinates in the historical collection records are within the polygon range defined by the boundary coordinates of the target geographic area. The temporal filtering conditions determine whether the collection timestamps in the historical collection records are between the start and end times of the preset historical period.

[0029] Specifically, the database uses spatial indexes to accelerate spatial range queries and time indexes to accelerate time range filtering. After executing the query, it returns all historical collection records that simultaneously meet the spatial and time conditions. The returned historical collection records are organized in memory or temporary storage according to a unified format to form a historical collection record set. Based on the boundary coordinates of the target geographic area and the start and end times of the preset historical period, it queries the storage unit to retrieve all matching historical collection records to form the historical collection record set. This operation accurately extracts a subset of data that is spatially and temporally completely related to the current analysis target from the potentially massive global historical data. This ensures that the content of the historical collection record set not only comprehensively covers all relevant activities in the target geographic area within the preset historical period, but also excludes interference data from irrelevant areas.

[0030] S2. Statistically analyze the historical collection records to obtain the collection frequency distribution of different sub-regions within the target geographical area.

[0031] S2.1 Divide the target geographic area according to the set rules to form a division scheme of different sub-regions within the target geographic area.

[0032] Furthermore, dividing the target geographic area according to the set rules refers to using predefined spatial segmentation criteria to discretize the continuous geographic space to form a division scheme of different sub-regions within the target geographic area. This can be based on a uniform geographic grid, such as dividing the entire area into square or hexagonal units of equal size, or based on natural or human geographic features, such as dividing based on natural boundaries like rivers and ridges or artificial boundaries like administrative divisions and postal codes. It can also be based on the road network topology to divide the area into different traffic zones.

[0033] Specifically, the continuous spatial analysis problem is transformed into a statistical problem for a finite number of well-defined sub-regions, forming a scheme for dividing different sub-regions within the target geographic area. The macro-geographic region is deconstructed into micro-level, manageable analytical units, with each sub-region serving as an independent observation sample. This allows the spatial distribution differences of data collection activities to be quantified and compared. For example, in urban traffic data collection, dividing by traffic zones can analyze the collection intensity of different road networks, while in environmental monitoring, dividing by watersheds can assess the data coverage of different water bodies.

[0034] S2.2. Based on the division scheme of different sub-regions within the target geographical area, count the collection location of each historical collection record in the historical collection record set, determine the sub-region to which the historical collection record belongs, and obtain the attribution result of the historical collection record.

[0035] Furthermore, the inclusion relationship between spatial points and polygons is determined. For each historical acquisition record in the historical acquisition record set, its acquisition location coordinates are extracted. Then, the coordinates are traversed or the spatial index is used to quickly query which specific sub-region polygon within the different sub-region division schemes of the target geographic area the coordinate point falls within. This is done efficiently using the ray method in computational geometry or the standard library function for points within polygons. After each historical acquisition record is determined, it is marked on the historical acquisition record or its index is associated with the corresponding sub-region identifier.

[0036] Specifically, all discrete historical data collection events are mapped onto a predefined, structured spatial framework, realizing the transformation from original coordinate points to sub-regional units with clear affiliations. Establishing this mapping relationship is a prerequisite for subsequent aggregation and statistical analysis. It enables the massive, geographically dispersed historical data collection records to be categorized according to spatial location, organizing unordered data collection events into an ordered data structure grouped by sub-regions. S2.3 Based on the attribution result of each historical collection record in the historical collection record set, use the accumulation method to calculate the total number of collections in different sub-regions within the target geographic area within the preset historical period.

[0037] Furthermore, initialize a counting array or mapping structure, whose keys or indices correspond to each sub-region in the sub-region division scheme of different sub-regions within the target geographic area. The initial value is set to zero. Then, traverse the historical collection record set. For each historical collection record, determine the corresponding sub-region based on its already marked belonging result, and increment the count value of the sub-region by one. After the traversal is completed, the final count value corresponding to each sub-region in the counting array is the total number of collections for the sub-region within the preset historical period.

[0038] Specifically, the cumulative calculation method is an efficient and direct aggregation statistical technique. By sequentially accessing and accumulating data, the number of individual events belonging to the same category is summarized into a category-level total. This transforms the spatial mapping relationship into an intuitive and measurable frequency indicator, allowing the activity level of each sub-region during the historical period to be quantified using simple counting. The total number of data collections is the most basic measure reflecting the intensity of historical data collection.

[0039] The expression for the total number of data collections is: ; in, For the first Total number of data collections for each sub-region The first in the historical collection of records One historical collection record, This refers to the index number of the historical data collection record within the historical data collection record set. The first division within the target geographic region Sub-regions The total number of records in the historical data collection set. This refers to the index number of different sub-regions within the target geographic region.

[0040] S2.4 Based on the sampling frequency distribution of different sub-regions within the target geographic area, the sampling frequency distribution of different sub-regions within the target geographic area is obtained by comparing the total number of samplings of different sub-regions within the target geographic area with the total duration of the preset historical period.

[0041] Furthermore, the frequency distribution of data collection in different sub-regions within the target geographic area is calculated by dividing the total number of data collections in the target geographic area by the total duration of a preset historical period. This is a standardized operation. The total duration of the preset historical period is a constant, such as the total length of time in days, months, or years. For each sub-region in the sub-region division scheme within the target geographic area, the total number of data collections in the sub-region is divided by the total duration of the preset historical period. The quotient is the average data collection frequency per unit time for the sub-region. This division operation is repeated for all sub-regions to obtain a set of data collection frequency values ​​corresponding to each sub-region, i.e., the frequency distribution of data collection in different sub-regions within the target geographic area.

[0042] Specifically, the original counts were normalized by introducing a time dimension, eliminating the incomparability that may arise from different analysis periods when simply using the total number of counts. For example, the collection intensity of an area that was collected ten times in a month is completely different from that of an area that was collected ten times in a year. The collection frequency distribution obtained by calculating the ratio allows for a fair comparison of the collection activity between different sub-regions on a unified time scale, thereby revealing the relative density differences of collection activities in space more realistically and scientifically.

[0043] S3. Based on the sampling frequency distribution, identify historical low-coverage sub-regions where the sampling frequency is lower than the first threshold.

[0044] S3.1 Obtain the sampling frequency distribution of different sub-regions within the target geographic area, and extract the sampling frequency value of each different sub-region within the target geographic area from the sampling frequency distribution of different sub-regions within the target geographic area.

[0045] Furthermore, data is read from the completed storage results. The data is represented as a structured collection, where each element is associated with a different sub-region within the target geographic region and its corresponding time-normalized collection frequency value. The collection frequency value of each different sub-region within the target geographic region is extracted from the collection frequency distribution of different sub-regions. To traverse this structured collection, each sub-region identifier and its corresponding collection frequency value are read and temporarily stored sequentially.

[0046] Specifically, separating data preparation from decision-making allows frequency comparison operations to be performed based on clear and independent input data, avoiding repeated searches in complex data structures, improving processing efficiency and code clarity. Obtaining and extracting these values ​​provides direct and unambiguous data input for subsequent threshold-based filtering, ensuring the accuracy and repeatability of the identification process.

[0047] S3.2. Based on a preset first threshold, compare the collection frequency values ​​of different sub-regions within each target geographical region with the first threshold, and collect data where the collection frequency value is lower than the first threshold.

[0048] Furthermore, the first threshold is a predefined reference value or a value generated according to certain rules. It represents the critical criterion for determining whether a sub-region belongs to low coverage. The comparison operation is performed on each collection frequency value one by one, checking whether the value is less than the first threshold. For collection frequency values ​​below the first threshold, whenever a collection frequency value is determined to be less than the first threshold, the identifiers of different sub-regions within the target geographic area corresponding to the collection frequency value are recorded in an intermediate result set or list.

[0049] Specifically, a simple dichotomy rule is applied to discretize and classify the continuous spatial frequency distribution, dividing all sub-regions into two categories: those below a threshold and those above a threshold. By introducing a unified and objective threshold standard, the relatively vague concept of low coverage is transformed into an automatically executed and explicit mathematical judgment, making the identification process fully automated and standardized, avoiding subjective assumptions. For example, the first threshold can be set as the median of the frequency values ​​collected from all sub-regions, thereby identifying the latter half of the region where the activity level is below the median; or it can be set as the average frequency minus a standard deviation, to identify cold spot regions that significantly deviate from the average level.

[0050] S3.3. Based on different sub-regions within the target geographic area where the collection frequency value is lower than the first threshold, form a list of historically low-coverage sub-regions.

[0051] Furthermore, the identifiers of different sub-regions within the target geographic area whose corresponding collection frequency values ​​are lower than the first threshold, together with the spatial definition information of these sub-regions in the sub-region division scheme within the target geographic area, such as boundary coordinates or polygon descriptions, and optionally their corresponding collection frequency values ​​themselves, are organized into a structured data list. This list is the list of historical low-coverage sub-regions, where each entry clearly represents a spatial unit that has been determined to have insufficient historical coverage, thus forming the list of historical low-coverage sub-regions.

[0052] Specifically, scattered identifiers generated through logical judgment are aggregated and encapsulated into a complete entity. When analyzing spatial distribution characteristics or generating bias correction guidance factors, this list can be directly used as the operational object, eliminating the need for complex threshold judgments and data retrieval. This list, as the final output, clearly and concisely identifies those relatively neglected parts of the entire target geographic area during historical data collection, providing clear spatial target guidance for the entire method to achieve proactive and targeted coverage bias correction.

[0053] S4. Based on the spatial distribution characteristics of historically low-coverage sub-regions and the associated changes in geographical elements, bias correction guidance factors are generated for each candidate collection location.

[0054] S4.1 Read the list of historical low-coverage sub-regions and extract the geographic range identifier and corresponding collection frequency value of each historical low-coverage sub-region in the list.

[0055] Furthermore, each entry in the list of historically low-coverage sub-regions is traversed, and two key fields are read and parsed from each entry: one is a geographic extent identifier used to uniquely determine the spatial location and shape of the sub-region, such as a polygon coordinate sequence or a feature ID referenced to the spatial database; the other is the collection frequency value of the sub-region in the judgment, reflecting the degree of its historical insufficient coverage. The decision results output from the previous step are transformed into basic data units containing spatial and quantitative attributes that can be processed in this step.

[0056] Specifically, clear input data specifications were established to ensure that subsequent spatial analysis and feature extraction operations were based on accurate, one-to-one correspondences of spatial entities and their associated quantitative indicators. Extracting geographic extent identifiers provided the geometric basis for subsequent spatial relationship calculations, while extracting collection frequency values ​​provided data support for potential weighted or priority ranking, thus laying a precise data preparation foundation for generating bias correction guidance factors.

[0057] S4.2 Based on the geographic extent identifier of each historical low-coverage sub-region, analyze the clustering, dispersion and geometric center location of the historical low-coverage sub-region within the target geographic region to form the spatial distribution characteristics of the historical low-coverage sub-region.

[0058] Furthermore, the analysis provides a macroscopic description of the spatial patterns of low-coverage regions. Clustering analysis can calculate spatial autocorrelation indices, such as the Moran index, for historically low-coverage sub-regions, or identify whether these sub-regions form distinct clusters or bands in space using clustering algorithms. Dispersion analysis can calculate the standard deviation elliptic parameters of the center points of historically low-coverage sub-regions to describe their distribution range and directionality. Geometric center location can be calculated by determining the mean or median of the geometric centers of all historically low-coverage sub-regions. These analyses utilize existing methods in computational geometry and spatial statistics to form spatial distribution characteristics of historically low-coverage sub-regions, elevating the list of isolated, point-like low-coverage sub-regions to an understanding of the overall spatial pattern of low-coverage phenomena.

[0059] Specifically, it is recognized that bias correction should not only target isolated points, but should consider the overall spatial structure of low-coverage areas. For example, if low-coverage areas are highly concentrated in a certain quadrant of a city, the correction strategy may require holistic strengthening of exploration in that quadrant; if low-coverage areas are discretely distributed, the correction strategy may require more dispersed collection points. This shift from an individual to a pattern perspective allows the subsequently generated bias correction guidance factors to contain richer spatial strategy information, guiding collection actions not only to cover individual weak points, but also to effectively address weak surfaces or weak zones defined by the spatial pattern, thereby optimizing coverage balance at a higher level.

[0060] S4.3 Query the historical collection record set to obtain historical geographic element data arranged in time series corresponding to each historical low-coverage sub-region.

[0061] Furthermore, a link is established between low-coverage areas and their historical dynamics. For each historical low-coverage sub-area in the list, a spatial range query is performed in the historical data collection set based on the spatial range defined by its geographic range identifier. All historical data collection records whose collection locations fall within the historical low-coverage sub-area are selected. These selected historical data collection records are then sorted in ascending order according to their collection timestamp field, forming a subset of historical data collection records related to a specific historical low-coverage sub-area, arranged in a time series. From these records, geographic feature data obtained from each collection can be further extracted or correlated, such as land cover type, attribute information, or change markers.

[0062] Specifically, a dedicated historical data timeline is constructed for each low-coverage area, linking spatial low-coverage identification with temporal data evolution. This allows the analysis to not only identify areas with low coverage but also to understand what changes have occurred in these areas historically. This provides crucial context for achieving more intelligent bias correction, because the urgency of data freshness requirements differs between a historically frequently changing low-coverage area and a historically extremely stable low-coverage area, which directly affects the priority judgment of subsequent bias correction.

[0063] S4.4 Perform time-series analysis on the historical geographic element data corresponding to each historically low-coverage sub-region, extract indicators describing the intensity or frequency of changes in geographic elements, and form related geographic element change characteristics.

[0064] Furthermore, by mining dynamic patterns from historical timelines, time-series analysis can include calculating the differences between geographic feature data collected at adjacent time points, such as by comparing the magnitude of changes in image feature vectors or attribute values; or counting the number of significant changes to geographic features within a preset historical period; it can also analyze the distribution of change types, such as additions, disappearances, or attribute modifications. The extracted indicators may be the average change per unit time, the frequency of changes, or the length of time since the most recent change, forming correlated geographic feature change characteristics and assigning a dynamic attribute label to each historically low-coverage sub-region. This feature transcends static spatial location and frequency, revealing the inherent activity or instability within the region.

[0065] Specifically, the value judgment of geographic data collection is extended from simply identifying areas that have not been visited to areas that have not been visited and may be changing or about to change. For example, two sub-regions with similarly low historical collection frequencies may exhibit different characteristics: one shows continuous building changes, while the other has maintained farmland characteristics for a long time. Obviously, the former has a more urgent need for data updates. Combining this change characteristic with the spatial low-coverage characteristic allows the generated bias correction guidance factor to distinguish between different types of low coverage: one is silent, potentially unimportant low coverage, and the other is active, potentially high-value low coverage. This guides collection resources to be more accurately directed to areas with insufficient historical coverage and significant dynamics, thereby achieving dual optimization of coverage correction and information value acquisition, greatly improving the intelligence level and long-term benefits of proactive collection strategies.

[0066] S4.5. Define fusion rules based on the spatial distribution characteristics of historically low-coverage sub-regions and the associated changes in geographical elements.

[0067] Furthermore, the core strategy is to integrate multi-dimensional information into a single decision-making basis. A fusion rule is a logical or mathematical expression that specifies how to combine inputs from spatial distribution characteristics and inputs from geographic feature change characteristics. For example, a fusion rule might stipulate that for any candidate sampling location, a spatial correlation degree is first calculated based on its spatial relationship with the clustering centers or discrete patterns identified in the spatial distribution characteristics of historical low-coverage sub-regions. A change urgency degree is then calculated based on the geographic feature change characteristics of its neighboring historical low-coverage sub-regions. The spatial correlation degree and change urgency degree are then combined using weighted summation or product. Another more complex fusion rule might consider the morphology of spatial distribution. For example, it might stipulate that if low-coverage areas are clustered, priority should be given to enhancing the detection of the periphery of the clustered areas to define the scope; if they are discrete, priority should be given to detecting the center of each discrete point. This is a configurable strategy framework.

[0068] Specifically, the macro-level spatial pattern understanding and the micro-level regional dynamic attributes are organically unified into a decision generation mechanism. By adjusting the parameters or form of the fusion rules, different bias correction strategies can be implemented, such as focusing more on filling spatial gaps, or focusing more on capturing dynamic changes, or seeking the best balance between the two. This makes the whole method highly flexible and adaptable, and can customize the optimal bias correction orientation for the geographic data collection needs of different application scenarios.

[0069] S4.6. Based on the fusion rules, a comprehensive judgment is made on the spatial proximity relationship between the candidate collection location and the historical low-coverage sub-region, as well as the change characteristics of the associated geographical elements. Each candidate collection location is assigned a quantitative result that represents the value of actively correcting historical coverage deviations, and a deviation correction guidance factor is obtained.

[0070] Furthermore, for each candidate acquisition location to be evaluated, according to the requirements of the fusion rules, the spatial proximity relationship between the location and each sub-region in the list of historical low-coverage sub-regions is first calculated. For example, the inverse distance to the boundary of the nearest historical low-coverage sub-region is calculated, or its classification score (whether it is located inside, on the edge, or outside the cluster) is calculated based on spatial distribution characteristics. The geographic feature change characteristic values ​​of the historical low-coverage sub-regions most relevant to the candidate location are obtained. Following the logic defined by the fusion rules, these calculated spatial relationship indicators and change characteristic values ​​are comprehensively calculated. This calculation produces a single scalar value, namely the bias correction guidance factor. The value of the factor directly represents the expected comprehensive value of performing acquisition at the location for actively and effectively correcting historical coverage bias.

[0071] Specifically, the abstract goal of correcting deviations is transformed into a quantifiable and comparable value score for each specific location through the bridge of integration rules. The output results are direct and clear, and deeply integrate considerations of both spatial pattern and spatiotemporal dynamics. The generated deviation correction guidance factor serves as a core decision variable.

[0072] S5. In response to the current data collection task request, generate a data collection task sequence based on the deviation correction guidance factor and the information gain prediction score of the candidate data collection locations obtained from the pre-trained geographic feature change prediction model.

[0073] S5.1 Receive the current acquisition task request based on the status constraints of the currently available acquisition devices.

[0074] Furthermore, receiving a current acquisition task request with constraints on the status of currently available acquisition devices refers to obtaining an external instruction that triggers the acquisition planning process. The instruction explicitly includes the schedulable resource constraints for executing this acquisition task. The constraints on the status of currently available acquisition devices specifically cover the real-time status information of each device in the set of acquisition devices currently available to execute the task, such as the device's current location coordinates, remaining power or fuel, device health status, current load, and the maximum operating range or acquisition capacity limited by the device type. The current acquisition task request may also include task-level constraints, such as the latest expected completion time of the task, overall budget constraints, or specific areas of interest that need to be prioritized for coverage, ensuring that the generated acquisition task sequence is not only optimal in terms of information value but also feasible at the actual physical execution level.

[0075] Specifically, by anchoring abstract planning problems to concrete, constrained physical world and operational conditions, the planning process is transformed from purely theoretical calculations into decision support oriented towards actual execution. This avoids generating theoretically optimal but unexecutable plans. Receiving and clarifying the current state constraints of available acquisition equipment is the primary prerequisite for generating an operable and executable sequence of acquisition tasks.

[0076] S5.2 Input the feature information of each candidate collection location into the pre-trained geographic feature change prediction model, and output the information gain prediction score of each candidate collection location.

[0077] Furthermore, machine learning models are used to predictively assess geographic dynamics. The pre-trained geographic feature change prediction model is a model that has already been trained using historical data, such as a prediction model based on temporal convolutional neural networks or long short-term memory networks. The feature information of each candidate collection location usually includes the spatial coordinates of the location, surrounding environmental features, a summary of historical collection records, and time-related contextual information such as season and day of the week. After these feature information are constructed into a format that meets the model input requirements and input, the pre-trained geographic feature change prediction model will perform forward inference and output one or more prediction results. The information gain prediction score is a scalar value interpreted from the model output. It quantifies the incremental information value that is expected to be brought by performing a collection at a candidate collection location. It usually reflects the probability of changes in geographic features in the future or the degree of uncertainty of expected changes. This elevates the value assessment of data collection from purely historical statistics to the prediction level based on machine learning.

[0078] Specifically, in the dynamically changing geographical world, the value of data collection lies not only in filling historical gaps and capturing changes that are happening or about to happen, but also in proactively estimating the potential freshness of information or the reduction in information entropy at each location through pre-trained models for predicting changes in geographical features. This allows data collection decisions to actively target locations most likely to generate new knowledge, thereby greatly improving the efficiency of information acquisition in a single data collection operation and directing limited data collection resources to areas with the highest expected information value.

[0079] S5.3 Based on the deviation correction guidance factor and the information gain prediction score, construct the comprehensive utility value of each candidate acquisition location according to the preset weights.

[0080] Furthermore, the bias correction guidance factor reflects the long-term strategic value of correcting uneven historical coverage and promoting spatial fairness; while the information gain prediction score reflects the short-term tactical value of acquiring fresh information and meeting immediate data needs. The preset weight is one or a set of predefined parameters used to weigh the relative importance of these two values ​​in the final decision. The specific way to construct the comprehensive utility value can be a simple linear weighted sum, that is, the comprehensive utility value equals the bias correction guidance factor multiplied by the first weight plus the information gain prediction score multiplied by the second weight; or it can be a more complex nonlinear function, such as a product form or a function based on multi-attribute utility theory. Constructing the comprehensive utility value according to the preset weights realizes the scalarization of multi-objective decision-making.

[0081] Specifically, a configurable and transparent fusion mechanism unifies the often-tense yet crucial goals of fairness and efficiency into a quantifiable and comparable framework. For example, when initially building the basic geographic database, a higher weight may be assigned to the bias correction guidance factor to quickly achieve full coverage; while during routine high-frequency updates and maintenance, a higher weight may be assigned to the information gain prediction score to quickly track changes. The flexibility of preset weights allows the entire method to adapt to different business stages and strategic priorities. By calculating the comprehensive utility value for each candidate collection location, a unified value ranking list is generated, providing clear and consistent input for the next step of global optimization under complex constraints. This ensures that the final planned collection task sequence can simultaneously consider the balance of coverage and the freshness of information.

[0082] S5.4 Based on the comprehensive utility value of all candidate acquisition locations, and combined with the movement cost and task capacity of the currently available acquisition devices, the acquisition task sequence of the currently available acquisition devices is planned with the goal of maximizing the total comprehensive utility value.

[0083] Furthermore, mobility cost refers to the resources consumed by a data acquisition device when moving it from one location to another, typically related to distance, time, or energy consumption. Task capacity constraints include the maximum number of locations each device can access in a single task, total working time limits, etc. The planning process can be modeled as a combinatorial optimization problem, such as the team orientation problem with capacity constraints or its variants. The objective function for solving the optimization problem is to maximize the sum of the combined utility values ​​of all candidate data acquisition locations selected into the task sequence, while satisfying constraints such as each device's movement path not exceeding its cost budget and the total number of locations accessed not exceeding its task capacity. The solution method can employ exact algorithms such as integer programming, or heuristic algorithms such as greedy algorithms and genetic algorithms, to plan the data acquisition task sequence of currently available data acquisition devices, deeply integrating discrete point-like value assessment with continuous path planning and resource allocation.

[0084] Specifically, under the hard constraints of device mobility and task carrying capacity, the goal is to find one or more paths that maximize the total utility value of the access points along the path. This requires the planning algorithm to consider both the value of the points and the cost of the lines. For example, two points with high total utility values ​​that are far apart may not be efficiently accessed by the same device in a single task; while a point with slightly lower utility values ​​that is near a high-value point may be included in the same path, thereby improving overall efficiency. Sequence planning based on optimization theory ensures that, under realistic constraints, a truly executable data collection action plan that maximizes total utility from a global perspective can be generated. This translates value judgments into specific instruction sequences that drive the actions of physical devices, completing a closed loop from data value recognition to optimal resource scheduling.

[0085] S6. Control the currently available data acquisition devices to move according to the data acquisition task sequence, perform geographic data acquisition operations, and obtain the current batch of raw geographic data.

[0086] S6.1 Generate movement path instructions for the currently available acquisition devices based on the acquisition positions and order assigned to the currently available acquisition devices in the acquisition task sequence.

[0087] Furthermore, the logical task list is transformed into physical, executable navigation instructions. The data acquisition task sequence explicitly defines the list of acquisition location coordinates that each currently available acquisition device needs to visit sequentially. Generating movement path instructions requires planning a continuous path connecting these points based on these location coordinates, considering the feasibility and efficiency of device movement. This can be done by calling path planning algorithms, such as the A* algorithm, Dijkstra's algorithm, or a route planning service that considers road networks, to calculate a coherent trajectory for each device that starts from the current location or starting point, passes through all specified acquisition locations in sequence, and may eventually return to the destination or stop at the last location. This trajectory consists of a series of path point coordinates, turning instructions, and speed suggestions, which constitute the movement path instructions.

[0088] Specifically, the transformation from a discrete set of points to continuous action was completed. The acquisition plan selected which points to collect and how to efficiently connect these points. Through integrated path planning, the movement path instructions ensured that the acquisition device could traverse all target points in the shortest time, shortest distance, or lowest energy consumption. This maximized the execution efficiency of the task sequence at the physical level. Generating movement path instructions is a key link between high-level decision-making and low-level control. It enables the abstract task sequence to be understood and executed by specific navigation execution units, ensuring the precise implementation of the entire acquisition operation in time and space.

[0089] S6.2. Send the movement path command to the navigation control unit of the currently available acquisition device. The navigation control unit drives the currently available acquisition device to move to each acquisition position in sequence.

[0090] Furthermore, the movement path command is transmitted via wireless communication link or internal bus to the navigation control unit integrated on the currently available acquisition device. The navigation control unit is a hardware and software combined controller. After receiving the movement path command, it parses it into lower-level control signals and drives the currently available acquisition device to move sequentially to each acquisition position. This involves the navigation control unit comprehensively processing position feedback from positioning sensors, attitude information from inertial measurement units, and movement path commands. Through control algorithms such as PID control or model predictive control, it generates real-time control commands for actuators such as drive motors and steering servos, enabling the device to stably and accurately arrive at each target acquisition position sequentially along the planned path.

[0091] Specifically, it achieves cross-domain instruction transmission and automated execution, seamlessly integrating high-level path instructions optimized on the server side with low-level control technology embedded in the device that can handle real-time dynamics. The navigation control unit acts as a translator and executor, transforming static path instructions into dynamic control behaviors that adapt to changes in the real-time environment, such as local obstacle avoidance when encountering temporary obstacles, while still adhering to the overall sequential access to the target. This ensures that even in non-ideal environments, the core intent of the data acquisition task can be reliably realized by sequentially accessing the designated locations. By sending movement path instructions and having them executed by the navigation control unit, the leap from planning to action is completed, enabling the theoretical solutions derived from optimization algorithms to become a reality in the physical world.

[0092] S6.3 After the currently available acquisition device reaches the acquisition location specified in the acquisition task sequence, the acquisition action set in the acquisition task sequence is triggered, and the sensor on the currently available acquisition device is operated to perform geographic data acquisition operation to obtain the current batch of raw geographic data.

[0093] Furthermore, the triggering mechanism can be based on location judgment, such as automatically triggering when the navigation control unit confirms that the error between the device pose and the target acquisition location is less than the set tolerance; or based on instructions, with the navigation control unit sending a trigger signal upon arrival. The acquisition actions set in the acquisition task sequence define the specific parameters and requirements for this acquisition, such as which sensors need to be activated, acquisition duration, shooting angle, and sampling frequency. The system operates the sensors on the currently available acquisition devices to perform geographic data acquisition operations. Based on these settings, control instructions are sent to the corresponding sensors to start them working. For example, controlling a panoramic camera to take a set of multi-angle photos, controlling a LiDAR to perform a three-dimensional scan, or controlling a spectrometer to acquire spectral data in a specific band. After the sensors complete the acquisition, they package the generated raw data stream along with information such as timestamps and location tags to form the current batch of raw geographic data, realizing task-aware, precisely triggered automated data acquisition.

[0094] Specifically, by precisely binding the data collection actions to spatial locations and pre-setting them with parameters, data collection is no longer blind or manually controlled. Instead, it is automatically executed in the right place and in the right way according to task requirements. For example, for areas requiring high-precision 3D modeling, LiDAR can be pre-triggered for fine scanning; for areas requiring only visual confirmation, camera shooting can be pre-triggered. By automating these pre-set actions, not only is the accuracy and consistency of data collection guaranteed, but operational efficiency is also greatly improved. Obtaining the original geographic data for the current batch marks the completion of the entire physical data collection cycle. The new data, rich in information, will be used for subsequent model updates and map reconstruction, thus initiating a new cycle of data value realization.

[0095] S7. Update the pre-trained geographic element change prediction model.

[0096] S7.1. Using the original geographic data of the current batch and the corresponding collection location and time, construct a training sample set for updating the geographic element change prediction model.

[0097] Furthermore, the newly acquired raw observation data is transformed into structured training samples that can be used for model learning. The current batch of raw geographic data consists of unprocessed measurements directly output by the sensors, such as image pixel matrices, point cloud coordinates, or spectral curves. The corresponding acquisition location and time provide the spatial and temporal context for each data point. The process of constructing the training sample set first includes data preprocessing, such as denoising, registration, and correction of the current batch of raw geographic data to improve data quality. Based on the acquisition location and time, the state of the corresponding location at a previous time point is extracted from existing historical geographic feature maps or databases as labels or comparison benchmarks. For example, for an acquisition location, the image acquired this time is paired with the image of the previous version at the same location to form a set of old-state-new-state sample pairs, which implicitly contain change information; or, the acquisition time is used to label the current data as the true state label under the latest timestamp.

[0098] Specifically, the training sample set for updating the geographic element change prediction model achieves automatic conversion from raw data to supervised learning signals. It utilizes the inherent spatiotemporal correlation of geographic data and automatically generates training samples with labels of change or latest status by aligning newly collected data with historical archives in a spatiotemporal manner. This avoids costly manual annotation and enables the model to evolve using data collected in the field each time. For example, by comparing new and old images, it can automatically identify newly built areas as positive samples and stable areas as negative samples, ensuring that the geographic element change prediction model can continuously learn from the latest real-world observations, enabling the model's predictive ability to keep pace with the times and adapt to the dynamic evolution of the geographic environment.

[0099] S7.2 Input the training sample set into the pre-trained geographic element change prediction model, execute the incremental learning algorithm, update the internal parameters of the pre-trained geographic element change prediction model, and obtain the updated geographic element change prediction model.

[0100] Furthermore, the core mechanism for achieving continuous model adaptation is as follows: The pre-trained geographic feature change prediction model has initial intrinsic parameters, allowing the model to learn new knowledge from newly arriving training sample sets (i.e., samples constructed from the current batch of data) without forgetting old knowledge. The specific operations of the incremental learning algorithm include: using the training sample set to perform forward propagation on the pre-trained geographic feature change prediction model to calculate the predicted output; calculating the loss between the predicted output and the true label or comparison target; calculating the gradient of the loss relative to the model's intrinsic parameters through the backpropagation algorithm; and finally using an optimization algorithm such as a variant of stochastic gradient descent to make minor adjustments to the model's intrinsic parameters based on the calculated gradient. This process may include strategies to prevent catastrophic forgetting, such as elastic weight consolidation or using a replay buffer, to update the intrinsic parameters of the pre-trained geographic feature change prediction model. This enables the model to learn and evolve online and continuously, recognizing that the geographic environment is constantly changing. A static, one-time trained prediction model will gradually become outdated over time. Through incremental learning, the model can immediately transform each fresh data obtained from field collection into an improvement in its predictive ability.

[0101] Specifically, it creates an enhanced closed loop of data collection, learning, prediction, and re-collection: the model guides the collection of valuable data, and the new data, in turn, improves the model, making the next prediction and guidance more accurate. For example, the model may initially be poor at predicting the change patterns of a new type of infrastructure, but after collecting relevant data and updating through incremental learning, its accuracy in predicting similar changes will improve. The updated geographic feature change prediction model means that the model's knowledge state is synchronized with the latest state of the real world, thus providing a more accurate and reliable information gain prediction score for the next round of data collection planning, driving the entire system to continuously evolve towards higher efficiency and accuracy.

[0102] S8. Re-infer the status of geographic elements in each sub-region within the target geographic area and generate an updated geographic element map.

[0103] S8.1 Input the latest representative feature data of each sub-region within the target geographic area into the updated geographic element change prediction model, and the updated geographic element change prediction model outputs the geographic element status inference results of each sub-region within the target geographic area.

[0104] Furthermore, the latest version of the model is used to assess the consistency of the entire region. The latest representative feature data of each sub-region within the target geographic area refers to a set of quantitative information reflecting the current status of each sub-region. This data may come from various sources, including but not limited to: summaries of the most recently successfully collected raw geographic data within the sub-region after feature extraction; basic attributes recorded in historical geographic feature maps of the sub-region; static or quasi-static descriptive information related to the sub-region obtained from external sources; and auxiliary information reflecting the spatiotemporal context of the sub-region. This feature data is constructed into a format that meets the input requirements of the updated geographic feature change prediction model and is batch-processed. The model performs forward propagation calculations on the features of each sub-region. The model output is the geographic feature status inference result, which is typically a multi-dimensional vector or structured data representing the model's prediction of the probability distribution, existence status, attribute values, or change confidence of geographic feature categories in the sub-region at the current time.

[0105] Specifically, it achieves a generalization from targeted, sparse field data collection to comprehensive, dense state inference. Utilizing an updated geographic element change prediction model that has absorbed the latest field observation knowledge through incremental learning, it serves as a powerful inference engine to estimate the state of sub-regions that have not been directly collected recently. Based on local change patterns and spatial correlations learned from field data collection points, it infers the state of the entire region. For example, after observing new features of road renovation in several sub-regions, the model may infer that neighboring sub-regions with similar road ages and traffic flow characteristics may also be undergoing renovations. This makes the geographic element state inference results not only based on historical data but also incorporate the patterns revealed by the latest, localized real-world observations, thus providing a more accurate and timely snapshot of the overall state than simply relying on old maps or sparse new data.

[0106] S8.2. Integrate the geographic feature status inference results of each sub-region within the target geographic area with the geographic feature information confirmed in the current batch of original geographic data to generate an updated geographic feature map.

[0107] Furthermore, this is the final step in synthesizing authoritative geographic products from diverse, multi-source, and multi-resolution information. The geographic feature information confirmed in the current batch of raw geographic data is specific geographic facts with high confidence obtained through automated data processing and interpretation, such as the outlines of newly constructed buildings automatically identified and vectorized from the latest acquired images, and road boundary lines accurately extracted from lidar point clouds. The fusion process needs to handle two different types of information: the geographic feature state inference result is a feature-based soft classification probability distribution or uncertain prediction, which has full coverage but inference uncertainty; the geographic feature information confirmed in the current batch of original geographic data is hard evidence derived from direct observation, which has high local reliability but incomplete spatial coverage. The fusion strategy can adopt a confidence-based weighted update: for sub-regions that have undergone field collection, the geographic feature information confirmed in the current batch of original geographic data is used first and directly, because these are conclusive observation facts; for sub-regions that have not undergone field collection, the geographic feature state inference result output by the updated geographic feature change prediction model is used as the best estimate of its state. In areas where the two overlap, consistency checks and evidence fusion can be performed to generate an updated geographic feature map, achieving the complementary advantages and organic unity of hard evidence and soft inference.

[0108] Specifically, a hierarchical information update authority system was established: direct observation data has the highest authority and is used for precise local updates; while model inference, which incorporates the latest observational knowledge, is used to fill in the observation gaps and achieve globally consistent updates. For example, in a city map, road realignment information in newly collected areas is definitively updated, while the road status in other uncollected areas remains the latest predicted state based on model inference. Model inference may also probabilistically refresh the building age attributes of all areas based on a certain urban renewal pattern exhibited in the newly collected areas. This fusion method ensures the timeliness and accuracy of the map in verified areas, while maintaining the timeliness and logical consistency of the entire map in unverified areas through intelligent inference. Thus, a new version of the geographic feature map that includes the latest field verification information and intelligent inference filling is generated in an efficient manner, providing a reliable, fresh, and complete spatial data foundation for all downstream applications.

[0109] This embodiment also provides a machine learning-based geographic data acquisition system, including: a data processing module, which acquires a set of historical acquisition records for a target geographic area, performs statistics on the set of historical acquisition records, and obtains the acquisition frequency distribution of different sub-regions within the target geographic area; The identification module identifies historically low-coverage sub-regions with collection frequencies below a first threshold based on the collection frequency distribution. The deviation correction factor generation module generates deviation correction guidance factors for each candidate collection location based on the spatial distribution characteristics of historical low-coverage sub-regions and the related changes in geographic elements. The sequence generation module, in response to the current data collection task request, generates a data collection task sequence based on the deviation correction guidance factor and the information gain prediction score of the candidate data collection locations obtained from the pre-trained geographic feature change prediction model. The execution module controls the currently available acquisition devices to move according to the acquisition task sequence, performs geographic data acquisition operations, obtains the current batch of raw geographic data, and updates the pre-trained geographic element change prediction model. The inference module re-infers the status of geographic elements in each sub-region within the target geographic area and generates an updated geographic element map.

[0110] This embodiment also provides a computer device applicable to the case of a machine learning-based geographic data acquisition method, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the machine learning-based geographic data acquisition method proposed in the above embodiment.

[0111] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.

[0112] This embodiment also provides a storage medium storing a computer program that, when executed by a processor, implements the machine learning-based geographic data acquisition method proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0113] In summary, this invention acquires and analyzes historical data collection records of the target area to quantify the frequency distribution of data collection, thereby identifying historically low-coverage sub-regions. Based on their spatial distribution and the changing characteristics of associated geographic elements, it generates deviation correction guidance factors for candidate locations. When responding to data collection task requests, it integrates the guidance factors with the information gain score given by the pre-trained change prediction model to generate an optimized data collection task sequence, driving the device to perform data collection and updating the prediction model with new data. Based on the updated model, it re-infers the geographic status of each sub-region to generate the latest map. By introducing a deviation correction mechanism, it synergistically optimizes information acquisition efficiency and spatial coverage balance in path planning, improves historical data collection deviations, and enhances the comprehensiveness and long-term application value of the geographic database.

[0114] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A geographic data acquisition method based on machine learning, characterized in that: This includes obtaining a set of historical data collection records for the target geographic area, performing statistics on the historical data collection records, and obtaining the data collection frequency distribution of different sub-regions within the target geographic area; Based on the sampling frequency distribution, historically low-coverage sub-regions with sampling frequencies below the first threshold were identified; Based on the spatial distribution characteristics of historically low-coverage sub-regions and the associated changes in geographic elements, bias correction guidance factors are generated for each candidate collection location, including the following steps: Read the list of historical low-coverage sub-regions and extract the geographic range identifier and corresponding collection frequency value for each historical low-coverage sub-region in the list. Based on the geographic extent identifier of each historical low-coverage sub-region, the clustering, dispersion and geometric center location of the historical low-coverage sub-region within the target geographic region are analyzed to form the spatial distribution characteristics of the historical low-coverage sub-region. Query the historical collection record set to obtain historical geographic element data arranged in time series corresponding to each historical low-coverage sub-region; Time series analysis is performed on the historical geographic element data corresponding to each historically low-coverage sub-region to extract indicators describing the intensity or frequency of changes in geographic elements, forming related geographic element change characteristics. Based on the spatial distribution characteristics of historically low-coverage sub-regions and the changes in associated geographic elements, fusion rules are defined; Based on the fusion rules, a comprehensive judgment is made on the spatial proximity relationship between candidate collection locations and historical low-coverage sub-regions, as well as the changes in associated geographical elements. Each candidate collection location is assigned a quantitative result that represents the value of actively correcting historical coverage deviations, thereby obtaining a deviation correction guidance factor. In response to the current data collection task request, a data collection task sequence is generated based on the bias correction guidance factor and the information gain prediction score of the candidate data collection locations obtained from the pre-trained geographic feature change prediction model. Control the currently available acquisition devices to move according to the acquisition task sequence, perform geographic data acquisition operations, obtain the current batch of raw geographic data, and update the pre-trained geographic element change prediction model; The status of geographic elements in each sub-region within the target geographic area is re-inferred to generate an updated geographic element map.

2. The geographic data acquisition method based on machine learning as described in claim 1, characterized in that: Obtaining the historical data collection set for the target geographic area includes the following steps: Define and set the boundary coordinates of the target geographic area and the start and end times of the preset historical period; Based on the boundary coordinates of the target geographic area and the start and end times of the preset historical period, all matching historical collection records are queried from the storage unit to form a set of historical collection records.

3. The geographic data acquisition method based on machine learning as described in claim 2, characterized in that: By statistically analyzing the historical data collection set, the frequency distribution of data collection in different sub-regions within the target geographic area is obtained, including the following steps: The target geographic area is divided according to the set rules to form a division scheme of different sub-regions within the target geographic area; Based on the division scheme of different sub-regions within the target geographical area, the collection location of each historical collection record in the historical collection record set is statistically analyzed to determine the sub-region to which the historical collection record belongs, thus obtaining the attribution result of the historical collection record. Based on the attribution result of each historical collection record in the historical collection record set, the total number of collections in different sub-regions within the target geographic area within the preset historical period is calculated using the accumulation method; Based on the frequency distribution of data collection in different sub-regions within the target geographic area, the frequency distribution of data collection in different sub-regions within the target geographic area is obtained by comparing the total number of data collections in different sub-regions within the target geographic area with the total duration of a preset historical period.

4. The geographic data acquisition method based on machine learning as described in claim 3, characterized in that: Based on the sampling frequency distribution, historically low-coverage sub-regions with sampling frequencies below a first threshold are identified, including the following steps: Obtain the sampling frequency distribution of different sub-regions within the target geographic area, and extract the sampling frequency value of each different sub-region within the target geographic area from the sampling frequency distribution of different sub-regions within the target geographic area; Based on a preset first threshold, the collection frequency value of different sub-regions within each target geographic region is compared with the first threshold, and the collection frequency value is collected if it is lower than the first threshold. A list of historically low-coverage sub-regions is formed based on different sub-regions within the target geographic area where the collection frequency value is lower than the first threshold.

5. The geographic data acquisition method based on machine learning as described in claim 1, characterized in that: In response to the current data collection task request, a data collection task sequence is generated based on the bias correction guidance factor and the information gain prediction scores of candidate data collection locations obtained from a pre-trained geographic feature change prediction model, including the following steps: Receive the current acquisition task request constrained by the status of currently available acquisition devices; The feature information of each candidate collection location is input into the pre-trained geographic feature change prediction model, and the information gain prediction score of each candidate collection location is output. Based on the deviation correction guidance factor and the information gain prediction score, a comprehensive utility value for each candidate acquisition location is constructed according to a preset weight. Based on the comprehensive utility value of all candidate acquisition locations, and combined with the movement cost and task capacity of currently available acquisition devices, the acquisition task sequence of currently available acquisition devices is planned with the goal of maximizing the total comprehensive utility value.

6. The geographic data acquisition method based on machine learning as described in claim 5, characterized in that: Control the currently available data acquisition devices to move according to the data acquisition task sequence, execute geographic data acquisition operations, and obtain the raw geographic data of the current batch, including the following steps: Based on the acquisition positions and order assigned to the currently available acquisition devices in the acquisition task sequence, generate the movement path instructions for the currently available acquisition devices; The movement path command is sent to the navigation control unit of the currently available acquisition device, and the navigation control unit drives the currently available acquisition device to move to each acquisition position in sequence. Once the currently available data acquisition device reaches the acquisition location specified in the acquisition task sequence, the acquisition action set in the acquisition task sequence is triggered, and the sensors on the currently available data acquisition device are operated to perform geographic data acquisition operations to obtain the current batch of raw geographic data.

7. The geographic data acquisition method based on machine learning as described in claim 6, characterized in that: Updating a pre-trained model for predicting changes in geographic features involves the following steps: Using the original geographic data of the current batch and the corresponding collection location and time, construct a training sample set for updating the geographic element change prediction model; The training sample set is input into the pre-trained geographic feature change prediction model, and the incremental learning algorithm is executed to update the internal parameters of the pre-trained geographic feature change prediction model, resulting in the updated geographic feature change prediction model.

8. The geographic data acquisition method based on machine learning as described in claim 7, characterized in that: The status of geographic features in each sub-region within the target geographic region is re-inferred to generate an updated geographic feature map, including the following steps: The latest representative feature data of each sub-region within the target geographic area are input into the updated geographic element change prediction model, and the updated geographic element change prediction model outputs the geographic element status inference results of each sub-region within the target geographic area. The updated geographic feature map is generated by integrating the inferred geographic feature status results of each sub-region within the target geographic region with the geographic feature information confirmed in the current batch of original geographic data.

9. A machine learning-based geographic data acquisition system, based on the machine learning-based geographic data acquisition method according to any one of claims 1 to 8, characterized in that: This includes a data processing module, which acquires a set of historical data collection records for the target geographic area, performs statistical analysis on the historical data collection records, and obtains the data collection frequency distribution of different sub-regions within the target geographic area. The identification module identifies historically low-coverage sub-regions with collection frequencies below a first threshold based on the collection frequency distribution. The deviation correction factor generation module generates deviation correction guidance factors for each candidate collection location based on the spatial distribution characteristics of historical low-coverage sub-regions and the related changes in geographic elements. The sequence generation module, in response to the current data collection task request, generates a data collection task sequence based on the deviation correction guidance factor and the information gain prediction score of the candidate data collection locations obtained from the pre-trained geographic feature change prediction model. The execution module controls the currently available acquisition devices to move according to the acquisition task sequence, performs geographic data acquisition operations, obtains the current batch of raw geographic data, and updates the pre-trained geographic element change prediction model. The inference module re-infers the status of geographic elements in each sub-region within the target geographic area and generates an updated geographic element map.

Citation Information

Patent Citations

  • Three-dimensional entity and element dynamic updating method based on terrain-level real scene

    CN119540486A