Surveying and mapping data quality supervision method and system based on machine learning
By using machine learning-based process topology models and generative adversarial networks, key quality control points in surveying and mapping data are identified and automatically replaced, solving the problems of lagging traditional supervision and reliance on manual labor, and achieving efficient and intelligent data quality management.
Patent Information
- Application Number
- CN202511475584.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-16
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-10-16
AI Technical Summary
Surveying and mapping data quality issues include missing data, large errors, redundancy, and poor consistency, which affect the effectiveness and reliability of practical applications. Traditional data quality inspection processes rely on manual labor, are passive and lagging, and cannot achieve intelligent supervision throughout the entire life cycle.
A process topology model is built based on machine learning, and graph neural inference networks are used to identify key quality control points. Dynamic quality supervision is carried out by combining time series analysis and generative adversarial networks, and the optimal decision candidate version is generated for automatic replacement.
It has enabled accurate identification and dynamic monitoring of key quality control points, improved monitoring efficiency, reduced the rate of missed reports and false reports, and formed a fully automated quality supervision system.
Smart Images

Figure CN120975652A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data supervision, in particular to a surveying and mapping data quality supervision method and system based on machine learning. BACKGROUND
[0002] Surveying and mapping data covers geographic information, spatial position, topography and other information, and is widely used in urban planning, land management, transportation construction, resource exploration and other fields. However, with the rapid increase in the amount of surveying and mapping data, data quality problems have become increasingly prominent. Fluctuations in the quality of surveying and mapping data, including data missing, large errors, redundancy and poor consistency, seriously affect the effectiveness and reliability of surveying and mapping data in practical applications. In the modern surveying and mapping big data ecosystem, how to transform the traditional data quality checking process that relies on manual, passive and lagging into an automated, proactive and intelligent supervision system throughout the entire life cycle of data, which needs to go beyond simple technical error correction, become a business intelligence platform serving organizational management and supervision, can quantitatively manage the performance of data production processes, support resource optimization, predict potential risks, and adaptively evaluate whether the data is suitable for specific use according to different business application scenarios, so as to improve data quality assurance as the core strategic asset of the organization. Therefore, a surveying and mapping data quality supervision method and system based on machine learning are proposed. SUMMARY
[0003] The purpose of the present application is to provide a surveying and mapping data quality supervision method and system based on machine learning, which supervises the quality of surveying and mapping data through machine learning.
[0004] To achieve the above purpose, the present application provides the following technical solutions: The surveying and mapping data quality supervision method based on machine learning comprises: Real-time reception of surveying and mapping data stream, synchronous acquisition of quality attribute information representing business process execution, construction of process topology model according to the quality attribute information, use of analysis engine to quantify the frequency of compliance deviation, and identification of quality critical control points; Based on the quality critical control points, the life cycle of surveying and mapping data in the business process is tracked, the dynamic quality attribute of the surveying and mapping data stream is automatically determined by using time series analysis method, the dynamic quality attribute is defined by business rules, and a quality risk prediction model is established to generate a quality risk index for the surveying and mapping data stream according to the dynamic quality attribute; trigger a business compliance decision-making process based on the quality risk index, input the quality risk index as a business strategy variable into a quality decision-making model to adjust the decision range and decision degree of the surveying and mapping data, generate n decision-making candidate versions, calculate the difference between each decision-making candidate version and the original data, and select the optimal candidate version for automatic replacement.
[0005] The quality attribute information is used to construct a process topology model. The quality attribute information representing the execution of the business process includes execution personnel, execution equipment, job environment parameters, and processing software versions. Each processing link of the quality attribute information is represented as a node in the graph, the flow relationship of the surveying and mapping data in the execution of the business process is represented as a directed edge in the graph, the risk weight is assigned to a specific combination of node attributes through machine learning based on historical logs, and a process topology graph integrating multi-dimensional attributes and associated risks is constructed. A process topology model is constructed based on the process topology graph and an analysis engine, the analysis engine is composed of a graph neural inference network, the graph neural inference network is used to learn the nonlinear propagation and accumulation mode of the risk weight in the process topology graph structure, quantify the occurrence frequency of compliance deviation, and identify the quality key control point.
[0006] The process of identifying the quality key control point is specifically: The process topology model is used to perform a business result-oriented sensitivity analysis for each node in the process topology graph. The sensitivity analysis simulates the impact of the quality fluctuation of the business output of the node on the business value of the final delivery result of the entire process when the quality fluctuation conforms to the historical law in the model deduction, and quantifies the impact as a business importance index. The occurrence frequency of historical compliance deviation associated with each node is calculated as a management index representing the stability of the node based on historical audit logs. The business importance index and the stability index are combined by weighting, the comprehensive management priority score of each node is calculated, and the node with the highest comprehensive management priority score is determined as the quality key control point.
[0007] The process of automatically determining the dynamic quality attribute of the surveying and mapping data stream using the time series analysis method is specifically: For the identified quality key control point, the dynamic quality attribute is extracted from the surveying and mapping data stream based on the business rules using the time series analysis method. The time series analysis method uses a multi-scale adaptive time series modeling framework, combines graph convolution network and spatiotemporal collaborative modeling, and captures long-term and short-term dependencies of the surveying and mapping data stream. The business data includes survey data generation time, transmission time, processing time, equipment usage, and operator; The dynamic quality attributes include data integrity, data precision, data consistency, data transmission delay, data processing delay, and survey equipment stability.
[0008] The process of establishing a quality risk prediction model to generate a quality risk index for the survey data stream based on the dynamic quality attributes is as follows: The quality risk prediction model uses gradient boosting trees to generate a quality risk index for the survey data stream based on the dynamic quality attributes. The gradient boosting trees construct a series of decision trees by gradually weighting the weak learners and adjust the weight of each decision tree leaf node to the quality risk index. Cross-validation is used for training and verification, and the deviation between the model prediction results and the actual quality fluctuations is continuously fed back through reinforcement learning to optimize the prediction results.
[0009] The process of using cross-validation for training and verification is as follows: The hierarchical blocking cross-validation based on business dimensions is used. The historical logs and corresponding true quality labels are grouped based on quality attribute information representing business process execution, including execution personnel, execution equipment, and job environment parameters, and divided into data blocks. In each iteration of cross-validation, the complete data block is used as the validation set, and the remaining blocks are used as the training set.
[0010] The quality risk index triggers the business compliance decision-making process as follows: The quality risk index triggers the business compliance decision-making process. The quality risk index is used as a business strategy variable to input a quality decision-making model to adjust the decision-making range and degree of survey data, generating n decision-making candidate versions. The quality decision-making model is a generative adversarial network. The generator maps the quality risk index to a modulation vector and injects it into the network level to control the range and degree of repair. By sampling the input random noise vector n times, n decision-making candidate versions are generated.
[0011] The difference between each decision-making candidate version and the original data is calculated, and the optimal candidate version is selected for automatic replacement as follows: This is accomplished by the business value impact assessment module and the optimal decision selection module. The business value impact assessment module calculates the weighted difference between each decision-making candidate version and the original data, where the weight is determined by the business importance index calculated by the process topology model for each data point. When modified, it will produce a higher difference score. The optimal decision selection module executes a multi-objective optimization algorithm, calculates a comprehensive service cost for each candidate version, and selects the lowest cost as the optimal version, wherein the comprehensive service cost comprises a residual quality risk index of the candidate version after correction, a weighted difference degree calculated by the service value influence evaluation module, a data fidelity cost, and a service execution cost corresponding to a data processing operation required for generating the version.
[0012] The machine learning-based surveying and mapping data quality supervision system comprises: A data processing module: real-time receiving of surveying and mapping data flow, and synchronous acquisition of quality attribute information representing service process execution; A key identification module: constructing a process topology model according to the quality attribute information, using an analysis engine to quantify the occurrence frequency of compliance deviation, and identifying quality key control points; A quality risk module: based on the quality key control points, tracking the life cycle of surveying and mapping data in the service process by data blood relationship, using a time series analysis method to automatically determine the dynamic quality attribute of the surveying and mapping data flow, the dynamic quality attribute being defined by a service rule, and establishing a quality risk prediction model to generate a quality risk index for the surveying and mapping data flow according to the dynamic quality attribute; A quality correction module: based on the quality risk index, triggering a service compliance decision-making process, inputting the quality risk index as a service strategy variable into a quality decision-making model to adjust the decision range and decision degree of surveying and mapping data, generating n decision-making candidate versions, calculating the difference degree between each decision-making candidate version and the original data, and selecting the optimal candidate version for automatic replacement.
[0013] Compared with the prior art, the present application has the following advantages: 1、The present application constructs a process topology model integrating multi-dimensional quality attributes such as execution personnel, equipment and environment, and uses a graph neural inference network for analysis, combines business importance indicators and historical stability indicators, and realizes accurate and dynamic identification of quality key control points.
[0014] 2、The application carries out deep analysis on the data flow of the key control point by adopting a multi-scale adaptive time sequence modeling framework, and combines gradient boosting tree and reinforcement learning to construct a quality risk prediction model capable of self-optimization, which breaks through the traditional quality evaluation method based on static threshold, can capture the dynamic evolution trend and complex mode of data quality, especially the introduction of the reinforcement learning feedback mechanism enables the model to learn from the actual effect of historical prediction and continuously optimize the prediction strategy, thereby significantly improving the accuracy of the risk index and the adaptability to the variable operation environment, and effectively reducing the false negative rate and false positive rate of quality problems.
[0015] 3、The application generates multiple repair candidate versions by introducing a generative adversarial network, and establishes a comprehensive business cost function including risk reduction degree, data fidelity cost and business execution cost to select the optimal version, realizing automatic and refined closed-loop disposal. The method avoids the one-size-fits-all mode of traditional automatic repair, converts the data repair problem into a multi-objective optimization problem facing business value, ensures that each automatic replacement decision is the optimal choice made after comprehensively balancing repair effect, influence on downstream business and self-cost, and finally forms a whole-process automatic quality supervision system of intelligent perception-accurate prediction-optimized decision. BRIEF DESCRIPTION OF DRAWINGS
[0016] Figure 1 It is a flowchart of the surveying and mapping data quality supervision method based on machine learning of the application; Figure 2 It is a data flowchart of the surveying and mapping data quality supervision method based on machine learning of the application; Figure 3 It is a structural diagram of the surveying and mapping data quality supervision system based on machine learning of the application. DETAILED DESCRIPTION
[0017] The technical solutions in the embodiments of the application will be described clearly and completely below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, rather than all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor fall within the protection scope of the application.
[0018] Embodiment one:
[0019] Please refer to Figure 1 The application provides a surveying and mapping data quality supervision method based on machine learning, and the technical solutions are as follows: The surveying and mapping data quality supervision method based on machine learning comprises: Real-time receive mapping data stream, synchronous acquisition of quality attribute information representing business process execution, construction of process topology model according to the quality attribute information, the process topology model uses an analysis engine to quantify the frequency of compliance deviation, and identifies the quality key control point; Based on the quality key control point, the data bloodline tracking mapping data in the business process is combined to determine the dynamic quality attribute of the mapping data stream automatically by using the time series analysis method, the dynamic quality attribute is defined by the business rule, and a quality risk prediction model is established, and a quality risk index is generated for the mapping data stream according to the dynamic quality attribute; Based on the quality risk index, trigger business compliance decision process, input the quality risk index as business strategy variable into quality decision model to adjust the decision range and decision degree of mapping data, generate n decision candidate versions, calculate the difference between each decision candidate version and original data, and select the optimal candidate version for automatic replacement.
[0020] The process topology model is constructed according to the quality attribute information: The quality attribute information representing business process execution includes execution personnel, execution equipment, job environment parameters and processing software version; Each processing link of the quality attribute information is represented as a node in the graph, the mapping data stream in the flow relationship of the business process execution is represented as a directed edge in the graph, the risk weight is assigned to the specific combination of node attributes through machine learning according to the historical log, and the process topology graph integrating multi-dimensional attributes and associated risks is constructed; Referring to Figure 2 , based on the process topology graph, train the analysis engine to construct the process topology model, the analysis engine is composed of graph neural inference network, the graph neural inference network is used to learn the nonlinear propagation and accumulation mode of the risk weight in the process topology graph structure, quantify the frequency of compliance deviation, and identify the quality key control point.
[0021] At the same time of real-time receiving mapping data stream, the quality attribute information related to each processing link of the data stream is synchronously acquired and recorded, the quality attribute information is multi-dimensional, including: execution personnel information (such as operator ID, technical level, historical operation excellent rate), execution equipment information (such as total station model, GNSS receiver ID, unmanned aerial vehicle sensor parameter, data processing workstation configuration), job environment parameters (such as temperature, humidity, satellite visibility, network bandwidth during field collection) and processing software information (such as version number of data solution software, plug-in version of data processing software).
[0022] Each specific processing link (such as data collection, data preprocessing, feature extraction or manual revision) is abstractly defined as a node in the graph, and the flow path of surveying and mapping data between these links is defined as a directed edge in the graph, thereby forming a flowchart reflecting the actual business logic; then a large number of historical production log data is called, which includes the above-mentioned multi-dimensional quality attribute information and the final quality result of each link (such as data rework record, error in quality inspection report, etc.). Using these historical data, a machine learning model (such as gradient boosting decision tree or deep neural network) is trained, which is used to learn the correlation between a specific combination of node attributes and the probability of quality problems. The output of the model, i.e. the "risk weight" of a specific attribute combination (such as "A employee" + "B equipment" + "C environment" + "D software"), for example, if the historical data shows that the rework rate of a certain combination is significantly higher, a higher risk weight value is assigned to this combination.
[0023] After the calculation of the risk weight is completed, these weight information is integrated as a key attribute to the corresponding node of the flow topology graph; After this step, a static business flowchart is transformed into a dynamic flow topology graph that integrates multi-dimensional attributes and associated risks. The enhanced topology graph not only depicts the flow path of data, but also quantitatively reveals the potential risk level of each link in the process under certain conditions. In order to understand the nonlinear propagation and accumulation pattern of risk in complex processes, this embodiment uses a graph neural inference network to build the final process topology model. A large number of risk process topology graphs generated in history are used as a training sample set to train the graph neural inference network. The training goal is to let the network learn the conduction rule of risk weight in the entire graph structure, for example, to identify how the risk of certain upstream nodes is amplified or suppressed by downstream specific links. After training is completed, this network constitutes an analysis model that can perform risk prediction; when a new surveying and mapping data stream is generated, a corresponding real-time risk process topology graph is generated and input into the trained model for inference. The model outputs the cumulative risk score of each node or the predicted frequency of compliance deviation. By analyzing these output values, the frequency of compliance deviation can be quantified, and those nodes that have the greatest impact on overall quality and the most significant risk accumulation effect are determined as quality key control points.
[0024] By changing the traditional passive, experience-based quality inspection to an active, data-driven risk prediction and supervision mode. By using the graph neural inference network to deeply learn the nonlinear conduction and accumulation effect of multi-dimensional factors such as personnel, equipment, environment, etc. in complex business processes, this method can automatically and accurately identify quality key control points, enabling the optimization of quality control resources, significantly improving supervision efficiency and data production quality, and realizing intelligent fine management.
[0025] For each node in the process topology graph, a business outcome-oriented sensitivity analysis is performed using the constructed process topology model, which is simulated in the model inference by artificially introducing a quality fluctuation conforming to its historical law for a specific node (such as "data preprocessing"), for example, according to historical data statistics, simulate an 80th percentile precision deviation of the output data of this link; then, under this simulation condition, continue to execute the model reasoning, observe and calculate the negative impact of this local fluctuation on the business value of the final delivery result of the whole process (such as the precision level, integrity score or project rating of the final data), and quantify the impact degree, and the value obtained is the business importance indicator of the node. Repeat this process for all nodes in the graph to obtain the business importance of each node.
[0026] In parallel, for the purpose of evaluating the stability of the quality of each process link, statistical analysis is performed based on historical audit logs; historical quality audit and compliance check records related to surveying and mapping production are retrieved, and the actual occurrence frequency of historical compliance deviations (such as errors, defects, rework, etc.) associated with each node in the process topology graph is accurately calculated; for example, the manual revision node has performed a total of 1000 times in the past year, and 50 times have produced quality problems that need to be reworked, so the deviation frequency is 5%, and this statistical frequency value is defined as a management indicator representing the stability of the node. The higher the indicator value, the more unstable the historical performance of the node.
[0027] The business importance indicator and stability indicator are combined by weighting to calculate a comprehensive management priority score for each node in the graph; the weight coefficient (for example, business importance accounts for 70%, and stability accounts for 30%) can be preset according to specific management goals to reflect different emphasis on the two types of risks, impact and problems; after completing the score calculation of all nodes, the scores are sorted, and the node with the highest comprehensive management priority score (or the nodes ranked in the top few) is finally determined as the quality key control point under the current business process. These control points are those that will cause the most serious consequences if problems occur, and are the highest priority for quality regulation and resource investment.
[0028] By combining the business importance of the node (i.e. the potential impact on the final result) and the historical stability (i.e. the frequency of historical problems), the fatal risk points with significant impact but occasional and the chronic hidden danger points with small impact but frequent are accurately distinguished, so that the final key control points not only reflect the probability of risk occurrence, but also reflect the severity of the consequences, thereby guiding the quality management resources to focus on the links that have a truly decisive impact on the overall quality in the most efficient way.
[0029] With the identified quality key control points as targets, the target-oriented and focused analysis of the mapping data stream fragments entering, flowing out and being processed inside these key nodes requires the extraction of dynamic quality attributes defined by pre-set business rules; for example, a business rule can define data integrity as the number of point clouds or the missing rate of key fields before and after data frame transmission; another rule can define data processing delay as the difference between the timestamps of data entering and leaving a certain key control point, which provides clear targets for subsequent timing analysis; To achieve accurate capture of dynamic quality attributes, a multi-scale adaptive timing modeling framework is adopted, which combines graph convolution network and spatio-temporal collaborative modeling technology; first, each index of multi-source business data (such as mapping data generation time, processing time, device ID, operator ID, etc.) accompanying the data stream is regarded as a node in the graph, and the edges of the graph are constructed according to the prior association between these indexes or the learned potential correlation; the graph convolution network is used to propagate information on the graph structure, thereby capturing the complex spatial dependency between different business data indexes on each time slice (for example, the change pattern of processing time when a specific operator uses a specific device combination); Based on the spatial dependency captured by the graph neural network, a spatio-temporal collaborative modeling unit is introduced to learn the evolution law of the entire graph state in the time dimension, which enables the model to understand not only the mutual influence of factors at the same time, but also the trend of these influences over time. This multi-scale framework can analyze data in different time windows simultaneously, thereby effectively capturing long-term dependencies (such as slow degradation of device performance) and short-term dependencies (such as instantaneous transmission delay caused by network congestion) of the mapping data stream; finally, a series of quantitative and dynamically changing index sequences are output, which are automatically determined and labeled as corresponding dynamic quality attributes according to the aforementioned business rules, such as data integrity, data accuracy, data consistency, data transmission delay, data processing delay and mapping device stability, etc. These outputs provide input for subsequent risk index generation, and through spatio-temporal collaborative modeling, deep insights into data quality from static slices to dynamic evolution are achieved, and the immediate correlation and long-term evolution law of multiple influencing factors are comprehensively understood; By applying multi-scale timing analysis only at quality key control points, not only is the analysis efficiency greatly improved, but also the potential associations and dynamic evolution laws between various business data are deeply mined using graph convolution and spatio-temporal collaborative modeling, making quality monitoring no longer limited to isolated index threshold judgments, but able to predict potential quality decline risks based on the overall health status and long-term trends of the data stream, achieving higher-dimensional real-time quality insights and prediction capabilities; After obtaining the multi-dimensional dynamic quality attributes of the survey data stream, a quality risk prediction model is established to integrate these discrete attributes into a single and intuitive quality risk index. Gradient boosting tree is used as the core quality risk prediction model, which is constructed in an iterative and progressive manner. First, a preliminary and relatively simple decision tree (i.e., weak learner) is generated based on the training data. Then, the residual or gradient between the prediction result of the tree and the true quality label is calculated. A new decision tree is established, which is no longer aimed at directly predicting the quality label, but at fitting the residual generated in the previous step. In this way, each subsequent added decision tree corrects the cumulative errors of all previous trees. This process is repeated to build a series of decision trees through step-by-step weighted learning, and by adaptively adjusting the weights of the leaf nodes of each decision tree, a powerful integrated model is finally formed. This model can effectively capture the complex nonlinear relationship between dynamic quality attributes (such as data completeness, processing delay, device stability, etc.) and the final quality risk, and output a continuous numerical value, i.e., the quality risk index. To ensure the robustness and generalization ability of the model, cross-validation method is used in the initial training and verification process. The prepared historical data set (containing multi-dimensional dynamic quality attributes and their corresponding true quality results) is divided into several mutually exclusive subsets (i.e., "folds"). In each training round, one subset is selected as the validation set, and the remaining subsets are used as the training set to train the gradient boosting tree model. This process is repeated until each subset has been used as a validation set once. By comprehensively evaluating the model's performance on all validation sets, a reliable estimate of its prediction performance can be obtained, and the model's hyperparameters can be optimized accordingly, effectively avoiding overfitting to specific training data and ensuring high accuracy when facing new, unseen data. Further reinforcement learning mechanism is introduced for continuous optimization. In this framework, the quality risk index generated by the model can be regarded as an action. After the data stream undergoes subsequent processing or manual quality inspection, its actual quality fluctuation (e.g., finally rated as qualified, rework, or discarded) is revealed. The deviation between the model's predicted risk index and the actual result is used as a feedback signal (i.e., reward or punishment). For example, if the model predicts low risk but the data is ultimately judged as unqualified, the model will receive a negative feedback. This feedback signal is used to adjust the internal parameters or weights of the gradient boosting tree model, with the goal of enabling the model to make decisions that maximize cumulative rewards in future predictions. Through this continuous "prediction-verification-feedback-optimization" closed loop, the model can continuously learn from new real cases and achieve dynamic and adaptive optimization of prediction results. The multi-dimensional and dynamic quality attributes are efficiently converged into a single and intuitive quality risk index by the gradient boosting tree model, which simplifies the complex quality state and provides a clear basis for immediate decision-making; by introducing the continuous feedback and optimization mechanism of reinforcement learning, the limitations of the traditional prediction model that is fixed after training are completely changed, the model can continuously learn from the deviation between its prediction and actual results in a dynamic production environment, realize long-term and adaptive precision improvement, and ensure the long-term effectiveness and reliability of quality risk assessment.
[0030] To solve the problem of autocorrelation in production data, which means that data points from the same business entity (such as the same person or equipment) often have inherent relevance, and random division may lead to information leakage, a hierarchical blocked cross-validation based on business dimensions is constructed. The complete historical data set (including logs and real quality labels) is grouped according to one or more key quality attribute information that represents the execution of the business process (such as the execution personnel, execution equipment, or job environment parameters). For example, if the execution personnel is selected as the grouping dimension, all data generated by operator A will be classified into one data block, and all data generated by operator B will be classified into another block, and so on. The entire data set is divided into several independent data blocks bound to business entities. After completing the block division, the basic unit for dividing the training set and the validation set in each iteration of cross-validation is no longer a single data point, but the entire data block. Specifically, in one iteration, one or more complete data blocks are selected as the validation set, and the remaining other blocks collectively constitute the training set. For example, in a 5-fold cross-validation with the execution personnel as the dimension, one iteration may use all the data of operator A as the validation set, and use the data of operators B, C, D, E, and other people to train the model. In this way, it can be ensured that the training set and the validation set are completely isolated in the selected business dimension, effectively testing the real generalization ability of the model when facing a completely unseen business entity (such as a new operator or a new device), and eliminating the overestimation of the evaluation results caused by internal correlation, obtaining a more accurate and more realistic application scenario evaluation of the model performance.
[0031] The hierarchical blocked cross-validation based on business dimensions fundamentally solves the problem of internal correlation in production data, eliminates the overestimation of model evaluation results caused by information leakage in standard random validation, and provides more reliable and credible performance guarantees for the actual deployment of the model.
[0032] When the quality risk index of the data stream is generated, if its value exceeds the preset threshold, the business compliance decision-making process will be automatically triggered, and the data correction stage will be entered. The core of the decision-making process is a generative adversarial network as a quality decision-making model. In this stage, the previously calculated quality risk index is not directly used as input to the generative network, but is used as a key business strategy variable. Specifically, the risk index is first converted into a modulation vector through a small mapping network, which is used to dynamically and finely adjust the range and degree of subsequent data correction. The modulation vector is injected into multiple levels of the generator network of the generative adversarial network. Its role is similar to a continuously adjustable knob. The higher the risk index, the stronger the adjustment effect of the modulation vector, which will guide the generative network to apply a larger range or deeper degree of adjustment when correcting the original data. To generate multiple correction schemes with diversity, n different random noise vector samples are taken at the input end of the generator. Each sampling, under the unified constraint of the modulation vector, will drive the generator to output a decision candidate version that is slightly different in content but meets the repair target under the current risk level. N selectable data versions that have been modified in different ways are obtained. After obtaining n decision candidate versions, the optimal version needs to be selected from them. The core is to calculate the weighted difference between each decision candidate version and the original data to measure the loss of data fidelity. The weight here directly reuses the business importance index calculated by the previous process topology model for each data point, which means that modifying key data points with higher business value will produce much higher difference scores than modifying ordinary data points. The potential impact of each modification on the core business value is quantified, and important data is given priority protection. The final scheme is determined through an optimal decision selection step, which calculates a comprehensive business cost for each candidate version, which is a multi-dimensional evaluation result obtained through a multi-objective optimization algorithm, and specifically includes three core parts: 1) residual quality risk: the data of the candidate version is re-input into the risk prediction model to obtain a corrected residual quality risk index representing the effectiveness of repair; 2) data fidelity cost: the weighted difference degree calculated in the previous step is directly used, representing the deviation degree of the modification from the original information; 3) business execution cost: the calculation resources and time consumption of the data processing operation required to generate the candidate version are evaluated; the three costs are comprehensively considered to calculate the total cost for each candidate version; the candidate version with the lowest cost is finally selected as the optimal decision, and is used to automatically replace the original high-risk data segment, completing the entire closed-loop supervision and disposal process. By integrating the three core dimensions of repair effectiveness, data fidelity and business execution cost into the consideration of comprehensive business cost, one-sided decisions caused by single objective optimization are completely avoided.
[0033] The generative adversarial network can dynamically generate multiple degrees of moderate and controllable repair candidate versions according to the risk index, realizing the flexibility of decision-making; at the same time, the selection mechanism of the optimal version adopts multi-objective optimization to seek the best balance point between repair effect, data fidelity (especially for high business value data) and execution cost, ensuring that the final automatic replacement decision is not only effective, but also a globally optimal solution under comprehensive consideration, greatly improving the automation level and intelligent degree of data quality supervision.
[0034] The present application constructs a full-process, closed-loop intelligent quality supervision system from pre-warning, in-process monitoring to post-automatic correction. First, through macro modeling of the business process, the key control points most prone to quality problems can be actively identified, realizing the source prediction of risk; secondly, focusing on these key points, the real-time data stream is deeply analyzed, and the complex and variable quality state is quantified into an intuitive risk index, realizing accurate and dynamic real-time monitoring; based on the risk index, the optimal correction decision is automatically triggered and executed, completing the seamless connection from problem discovery to problem solving, changing the traditional quality inspection mode relying on manual and lag processing, greatly improving the automation level, processing efficiency and reliability of the final results of surveying and mapping data production.
[0035] Embodiment two Unmanned aerial vehicle aerial photogrammetry is one of the mainstream technologies in the current surveying and mapping field, and its operation process is long and has many influencing factors (such as flight attitude, environmental changes, software algorithms, etc.), resulting in high uncertainty in data quality. The present embodiment provides a surveying and mapping data quality supervision system based on machine learning to supervise the quality of surveying and mapping data.
[0036] ReferenceFigure 3 When a UAV mapping task is started, the data processing module receives the aerial survey data stream returned from the UAV in real time, which includes raw aerial photographs (image data) and high-frequency POS data (position and attitude) in the flight control log; synchronously retrieves and records quality attribute information related to this task, such as the pilot ID performing this flight (e.g., Zhang San), the UAV and sensor model used (e.g., DJI M300 RTK + P1 camera), the environmental parameters at the time of the task area (e.g., wind speed 5 m / s, good lighting conditions), and the ground station software version number used for data solving (e.g., DJI Terra v3.9); After the key identification module receives the above information, it immediately constructs a process topology model for this task based on historical similar project data, which includes nodes such as flight planning, data acquisition, POS solving, aerial triangulation, and dense matching. The analysis engine finds that under the condition of "wind speed 5 m / s", the failure rate or reprocessing rate of the aerial triangulation node has significantly increased in history, and combined with sensitivity analysis, it confirms that once this node has a problem, it will directly lead to the precision of the entire project not meeting the standard. Therefore, the system automatically identifies and marks "aerial triangulation" as the quality key control point of this task; The quality risk module uses a time series analysis model to monitor the data stream input into this link and determines its dynamic quality attributes in real time. For example, it finds that the precision of the POS data solving after this task (dynamic quality attribute) has a slight drift, and the time for processing image connection point matching (dynamic quality attribute) is 15% longer than normal. The quality risk prediction model judges that there is a high risk of model solving failure based on the combination of these dynamic attributes, and generates a quality risk index of 0.82 for the data stream being processed; The risk index of 0.82 immediately triggers the quality correction module, and the quality decision model uses 0.82 as a strong strategy variable to control the repair degree, generating three candidate versions: version A is to remove some images with the lowest connection point matching degree and re-solve; version B is to keep all images but automatically optimize the camera parameters and increase the number of iterations; version C is to introduce historical images from adjacent areas as auxiliary constraints for joint adjustment; then, the system evaluates the three versions: version A has the lowest residual risk but high data fidelity cost; version B has balanced costs in all aspects; version C has the best effect but the highest business execution cost (fetching historical data, increasing calculation); finally, the multi-objective optimization algorithm calculates that the comprehensive business cost of version B is the lowest, so the system selects the aerial triangulation result of version B to automatically replace the original high-risk solving result.
[0037] While embodiments of the application have been shown and described, it is to be understood that the embodiments described are merely exemplary of the principles and application of the present application. Numerous modifications and adaptions can be effected without departing from the spirit and scope of the present application, which is not limited to the exact construction and arrangement described. It is intended, therefore, to cover all modifications and adaptions that fall within the scope of the claims and their equivalents.
Claims
1. A machine learning-based method for monitoring the quality of surveying and mapping data, characterized in that, include: The system receives surveying and mapping data streams in real time, synchronously acquires quality attribute information representing the execution of business processes, constructs a process topology model based on the quality attribute information, and uses an analysis engine to quantify the frequency of compliance deviations and identify key quality control points. Based on the quality critical control points and data lineage tracing of the lifecycle of the surveying and mapping data stream in the business process, the dynamic quality attributes of the surveying and mapping data stream are automatically determined using time series analysis methods. The dynamic quality attributes are defined by business rules, and a quality risk prediction model is established. A quality risk index is generated for the surveying and mapping data stream based on the dynamic quality attributes. Based on the quality risk index, a business compliance decision-making process is triggered. The quality risk index is used as a business strategy variable to input into the quality decision-making model to adjust the decision scope and degree of the surveying and mapping data stream, generating n decision candidate versions. The difference between each decision candidate version and the surveying and mapping data stream is calculated, and the optimal candidate version is selected for automatic replacement.
2. The method for monitoring the quality of surveying and mapping data based on machine learning according to claim 1, characterized in that, The process topology model is constructed based on the quality attribute information: The quality attribute information representing the execution of the business process includes the executor, the equipment used, the operating environment parameters, and the processing software version. Each processing step of the quality attribute information is represented as a node in the graph, and the flow relationship of the surveying data stream in the execution of the business process is represented as a directed edge in the graph. Based on historical logs, risk weights are assigned to the node attributes through machine learning, and a process topology graph integrating multi-dimensional attributes and associated risks is constructed. A process topology model is constructed based on a process topology graph training analysis engine. The analysis engine consists of a graph neural inference network, which is used to learn the nonlinear propagation and accumulation pattern of the risk weights in the process topology graph structure, quantify the frequency of compliance deviations, and identify key quality control points.
3. The method for monitoring the quality of surveying and mapping data based on machine learning according to claim 2, characterized in that, The process of identifying critical quality control points is as follows: Using the aforementioned process topology model, a sensitivity analysis oriented towards business outcomes is performed on each node in the process topology diagram. The sensitivity analysis is the degree of impact on the business value of the final deliverable of the entire process when the business output of the node experiences quality fluctuations that conform to historical patterns during model simulation, and the degree of impact is quantified into a business importance indicator. The frequency of historical compliance deviations associated with each node is statistically analyzed based on historical audit logs, serving as a management indicator characterizing the stability of that node. The business importance index and the stability management index are weighted and combined to calculate a comprehensive management priority score for each node, and the node with the highest comprehensive management priority score is determined as the quality critical control point.
4. The method for monitoring the quality of surveying and mapping data based on machine learning according to claim 1, characterized in that, The specific process of automatically determining the dynamic quality attributes of the mapping data stream using time series analysis is as follows: For the identified critical quality control points, dynamic quality attributes are extracted from the surveying and mapping data stream in a targeted manner according to the business rules using time series analysis methods. The time series analysis method adopts a multi-scale adaptive time series modeling framework, which combines graph convolutional networks with spatiotemporal co-modeling to capture the long-term and short-term dependencies of surveying and mapping data streams. The dynamic quality attributes include data integrity, data accuracy, data consistency, data transmission latency, data processing latency, and surveying equipment stability.
5. The method for monitoring the quality of surveying and mapping data based on machine learning according to claim 1, characterized in that, The specific process of establishing a quality risk prediction model and generating a quality risk index for the mapping data stream based on the dynamic quality attributes is as follows: The quality risk prediction model uses a gradient boosting tree to generate a quality risk index for the mapping data stream based on the dynamic quality attributes. The gradient boosting tree constructs a series of decision trees by progressively weighting weak learners and adaptively adjusts the weights of the leaf nodes of each decision tree to reflect the quality risk index. Cross-validation is used for training and validation, and reinforcement learning is used to continuously provide feedback on the deviation between the model's prediction results and the actual quality fluctuations, thereby optimizing the prediction results.
6. The method for monitoring the quality of surveying and mapping data based on machine learning according to claim 5, characterized in that, The specific process of using cross-validation for training and validation is as follows: A hierarchical blocking cross-validation based on business dimensions is adopted, which groups historical logs and corresponding real quality labels according to the quality attribute information that characterizes the execution of business processes. The quality attribute information includes the parameters of the executor, the execution equipment, and the working environment, and divides them into data blocks. In each iteration of cross-validation, the complete data block is used as the validation set, and the remaining blocks are used as the training set.
7. The method for monitoring the quality of surveying and mapping data based on machine learning according to claim 1, characterized in that, The specific process for triggering business compliance decisions based on the quality risk index is as follows: Based on the quality risk index, a business compliance decision-making process is triggered. The quality risk index is used as a business strategy variable and input into the quality decision-making model to adjust the decision range and degree of the mapping data, generating n decision candidate versions. The quality decision-making model is a generative adversarial network, whose generator maps the quality risk index into a modulation vector and injects it into the network layer to control the scope and degree of repair. By sampling the input random noise vector n times, n decision candidate versions are generated.
8. The method for monitoring the quality of surveying and mapping data based on machine learning according to claim 1, characterized in that, The calculation of the difference between each decision candidate version and the original data, and the selection of the optimal candidate version for automatic replacement, specifically involves: This is accomplished collaboratively by the business value impact assessment module and the optimal decision selection module. The business value impact assessment module calculates the weighted difference between each decision candidate version and the original data, where the weight is determined by the business importance index calculated by the process topology model for each data point, and a higher difference score will be generated when it is modified. The optimal decision selection module executes a multi-objective optimization algorithm to calculate the comprehensive business cost for each candidate version and selects the version with the lowest cost as the optimal version. The comprehensive business cost includes: the residual quality risk index of the candidate version after correction, the weighted difference calculated by the business value impact assessment module as the data fidelity cost, and the business execution cost corresponding to the data processing operations required to generate the version.
9. A surveying and mapping data quality supervision system based on machine learning, characterized in that: Data processing module: Receives surveying and mapping data streams in real time and synchronously acquires quality attribute information representing the execution of business processes; Key identification module: Based on the quality attribute information, a process topology model is constructed. The process topology model uses an analysis engine to quantify the frequency of compliance deviations and identify key quality control points. Quality Risk Module: Based on key quality control points and data lineage tracing of the lifecycle of surveying and mapping data in the business process, the module automatically determines the dynamic quality attributes of the surveying and mapping data stream using time series analysis methods. The dynamic quality attributes are defined by business rules, and a quality risk prediction model is established. Based on the dynamic quality attributes, a quality risk index is generated for the surveying and mapping data stream. Quality Correction Module: Based on the quality risk index, the business compliance decision-making process is triggered. The quality risk index is used as a business strategy variable to input into the quality decision-making model to adjust the decision range and degree of the surveying and mapping data, generate n decision candidate versions, calculate the difference between each decision candidate version and the original data, and select the optimal candidate version for automatic replacement.
Citation Information
Patent Citations
Diversified community nursing service tracing and service quality improving method
CN118333642A
Informatization processing method based on big data
CN119003495A
Gas station equipment monitoring system and method based on Internet of Things
CN120123951A
Management method of financial service platform based on Internet
CN120389852A
Urban construction and traffic engineering investigation and design enterprise product quality management system and method
CN120450504A