A management method and system for multi-terminal engineering internal data
By conducting quality evaluation and data fusion on multi-end engineering internal data, generating collaborative fusion engineering information and training engineering value extraction models, the problems of poor accuracy and low adaptability of value information extraction in the existing technology are solved, and the reliability and effectiveness of data management are improved.
Patent Information
- Application Number
- CN202411974936.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-31
AI Technical Summary
In the prior art, the value information extraction of multi-end engineering internal data has poor accuracy and low adaptability, resulting in unreliable data management.
By obtaining multi-end engineering internal data, classification, redundancy and integrity analysis are carried out, data quality is determined, and data fusion is carried out through spatiotemporal and correlation associations, collaborative fusion engineering information is generated, and engineering value extraction models are trained.
It improves the accuracy and fitness of value information extraction, and enhances the reliability and effectiveness of data management.
Smart Images

Figure CN119377582B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of engineering internal data processing, and in particular to a management method and system for multi-terminal engineering internal data. Background Art
[0002] The multi-terminal engineering internal data management solution aims to solve the complexity of collaborative management of engineering internal data in the current cross-platform and cross-device environment. With the continuous expansion of engineering project scale and the rapid development of technology, engineering internal data has become diversified and complex, involving multiple stages such as design, construction, operation and maintenance, and requires efficient transmission and sharing between different terminal devices (such as PCs, tablets, mobile phones, etc.).
[0003] In the existing technology, the quality of multi-terminal engineering internal data is uneven, and the temporal, spatial and correlation relationships between multiple types of data are relatively complex, resulting in poor accuracy and low adaptability in the extraction of valuable information, which is not conducive to subsequent data management and reduces the reliability of data management.
[0004] Therefore, how to improve the accuracy and adaptability of valuable information extraction is a technical problem that needs to be solved. Summary of the invention
[0005] The purpose of the present invention is to solve the problems of poor accuracy and low adaptability in the prior art in extracting valuable information, and to propose a management method for multi-terminal engineering internal data, which includes:
[0006] Obtain multi-terminal engineering internal data, classify the multi-terminal engineering internal data, analyze the redundancy and integrity of each type of multi-terminal engineering internal data, and determine the quality of each type of multi-terminal engineering internal data;
[0007] Screen and preprocess the multi-terminal engineering internal data by the quality of each type of multi-terminal engineering internal data, identify the spatiotemporal associations between the multi-terminal engineering internal data, and fuse each type of multi-terminal engineering internal data to obtain a spatiotemporal fusion data set;
[0008] Identify the correlation between multi-terminal engineering internal data, and fuse each type of multi-terminal engineering internal data to obtain relevant fused data sets;
[0009] Combine the spatiotemporal fusion data set and the related fusion data set to generate collaborative fusion engineering information, train the engineering value extraction model based on the collaborative fusion engineering information, and realize the management of multi-terminal engineering internal data through engineering value.
[0010] In some embodiments of the present application, the redundancy and integrity of each type of multi-terminal engineering internal data are analyzed, including:
[0011] Divide multi-terminal engineering internal data into structured data and unstructured data;
[0012] For structured data, set the corresponding data unit length according to the standard size;
[0013] For unstructured data, analyze the complexity of each type of multi-terminal engineering internal data and set the corresponding data unit length according to the complexity;
[0014] Check the missing values, empty values and NULL values of each data unit in structured and unstructured data, and define the integrity of each type of multi-terminal engineering internal data based on the missing values, empty values and NULL values;
[0015] Determine the duplication categories between data units in structured data and unstructured data. The duplication categories include complete duplication and partial duplication. For complete duplication, delete multiple data units under complete duplication and only retain one data unit.
[0016] For partial duplication, the data units are divided into different clusters through clustering algorithm, and the overlap between each data unit and other data units in the same cluster is calculated to build a cluster overlap list. The redundancy degree of each data unit is calculated based on the cluster overlap list, so as to determine the redundancy degree of each type of multi-terminal engineering internal data; ;
[0017] in, For the The first The degree of redundancy of each data unit, There are two preset weights, For the The first The maximum overlap of data units, For the The number of data units in a cluster, Indicates The number of data units in a cluster after removing the data unit with the largest overlap, Indicates The first The data units are combined with the data units with the maximum overlap after removing the data units with the The overlap between data units;
[0018] Among them, the redundancy degree and integrity degree of multi-terminal engineering internal data are used to describe the redundancy and integrity of multi-terminal engineering internal data respectively.
[0019] In some embodiments of the present application, the quality of each type of multi-terminal engineering internal data is determined, including:
[0020] The quality of each type of multi-terminal engineering internal data is calculated based on the redundancy and integrity of the multi-terminal engineering internal data.
[0021] In some embodiments of the present application, the multi-terminal engineering internal data is screened and preprocessed according to the quality of each type of multi-terminal engineering internal data, including:
[0022] The multi-terminal engineering internal data with quality lower than the quality threshold will be deleted, and the multi-terminal engineering internal data with quality not lower than the quality threshold will be pre-processed.
[0023] In some embodiments of the present application, the spatiotemporal association between multi-terminal engineering internal data is identified, and each type of multi-terminal engineering internal data is fused to obtain a spatiotemporal fusion data set, including:
[0024] Use time analysis tools to analyze the changing trends of multi-terminal engineering internal data over time to obtain time relationships, use GIS software to project multi-terminal engineering internal data onto geographic space, and analyze the spatial relationships between multi-terminal engineering internal data on geographic space;
[0025] The multi-terminal engineering internal data are aligned in time and space, and the time features and spatial features are extracted based on the time and spatial relationships of the multi-terminal engineering internal data. The time features and spatial features are fused to obtain a time-space fusion data set.
[0026] In some embodiments of the present application, the correlation between the multi-terminal engineering internal data is identified, and each type of multi-terminal engineering internal data is fused to obtain a related fused data set, including:
[0027] Extract data features from multi-terminal engineering internal data, perform data standardization on the data features, draw a scatter plot between the data features of the multi-terminal engineering internal data, and analyze the scatter plot to determine the correlation association type between the data features of the multi-terminal engineering internal data, the correlation association type includes linear correlation and nonlinear correlation;
[0028] Calculate the correlation coefficient of the data features under the linear correlation relationship, assign the fusion weight of the data features according to the correlation coefficient, fuse the data features with linear correlation, and obtain the linear correlation fusion set;
[0029] The multiple data features under nonlinear correlation are fused by a preset feature fusion model to obtain a nonlinear correlation fusion set;
[0030] The correlation fusion data set includes two parts: linear correlation fusion set and nonlinear correlation fusion set.
[0031] In some embodiments of the present application, the spatiotemporal fusion data set and the related fusion data set are combined to generate collaborative fusion engineering information, including:
[0032] Integrate the spatiotemporal fusion data set and the related fusion data set to generate collaborative fusion engineering information, and evaluate the comprehensive information index of the data of the spatiotemporal fusion data set. According to the linear correlation fusion set and the nonlinear correlation fusion set in the related fusion data set, establish a linear relationship network and a nonlinear relationship network respectively, and analyze the linear relationship network and the nonlinear relationship network to determine the comprehensive information index of the data relationship of the linear correlation fusion set and the nonlinear correlation fusion set respectively.
[0033] The comprehensive information indicators of data and the comprehensive information indicators of data relationships are added to the collaborative fusion engineering information.
[0034] In some embodiments of the present application, the engineering value extraction model is trained based on the collaborative fusion engineering information, including:
[0035] Determine the multi-dimensional information complexity level of collaborative fusion engineering information through the comprehensive information index of data and the comprehensive information index of data relationship in collaborative fusion engineering information; ;
[0036] in, To collaboratively integrate the multi-dimensional information complexity level of engineering information, To coordinate and integrate the comprehensive information indicators of data in engineering information, are the combined weights corresponding to the linear correlation and nonlinear correlation, respectively. are the comprehensive information indicators of the data relationships corresponding to the linear correlation and nonlinear correlation, All are preset constants;
[0037] The model parameters are determined by collaboratively integrating the multi-dimensional information complexity level of the engineering information, and the engineering value extraction model is trained according to the model parameters and the collaboratively integrated engineering information.
[0038] Correspondingly, the present application also provides a management system for multi-terminal engineering internal data, including:
[0039] An evaluation module is used to obtain multi-terminal engineering internal data, classify the multi-terminal engineering internal data, analyze the redundancy and integrity of each type of multi-terminal engineering internal data, and determine the quality of each type of multi-terminal engineering internal data;
[0040] The first fusion module is used to screen and preprocess the multi-terminal engineering internal data according to the quality of each type of multi-terminal engineering internal data, identify the spatiotemporal association between the multi-terminal engineering internal data, and fuse each type of multi-terminal engineering internal data to obtain a spatiotemporal fusion data set;
[0041] The second fusion module is used to identify the correlation between the multi-terminal engineering internal data and fuse each type of multi-terminal engineering internal data to obtain a related fusion data set;
[0042] A determination module is used to combine the spatiotemporal fusion data set and the related fusion data set to generate collaborative fusion engineering information, train the engineering value extraction model based on the collaborative fusion engineering information, and realize the management of multi-terminal engineering internal data through engineering value.
[0043] Compared with the prior art, the present invention has the following beneficial effects:
[0044] 1. Analyze the redundancy and integrity of each type of multi-terminal engineering internal data, determine the quality of each type of multi-terminal engineering internal data, and screen and pre-process the multi-terminal engineering internal data based on the quality of each type of multi-terminal engineering internal data. Divide the data into data units, analyze the degree of integrity and redundancy, and define the data quality by combining the two, thereby improving the control of data quality and providing a reliable foundation for subsequent data fusion and model training.
[0045] 2. Identify the spatiotemporal associations and correlation associations between multi-terminal engineering internal data to construct spatiotemporal fusion data sets and related fusion data sets, and then generate collaborative fusion engineering information. The collaborative fusion engineering information not only contains the spatiotemporal characteristics and correlation characteristics of multi-terminal engineering internal data, but also reveals the inherent connection and mutual influence between the data, providing rich basic data for subsequent engineering value extraction. In addition, the collaborative fusion engineering information also evaluates the spatiotemporal fusion data set and the related fusion data set in complex situations, that is, the comprehensive information index. It helps to confirm the model parameters, so as to accurately train the engineering value extraction model. It improves the accuracy and adaptability of value information extraction, indirectly assists the management of internal data, and improves the effectiveness of data management. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 A flowchart of a method for managing multi-terminal engineering internal data proposed by the present invention;
[0047] Figure 2 A structural diagram of a management system for multi-terminal engineering internal data proposed by the present invention. DETAILED DESCRIPTION
[0048] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.
[0049] Reference Figure 1A method for managing multi-terminal engineering internal data includes the following steps:
[0050] Step S101, obtaining multi-terminal engineering internal data, classifying the multi-terminal engineering internal data, analyzing the redundancy and integrity of each type of multi-terminal engineering internal data, and determining the quality of each type of multi-terminal engineering internal data.
[0051] In this embodiment, multi-terminal engineering internal data refers to various data and information related to internal work in multi-terminal projects (engineering projects involving multiple terminals or platforms). These data are usually generated during the construction of engineering projects and are used to support the decision-making, execution and monitoring of engineering projects. Multi-terminal engineering internal data mainly comes from multiple channels such as project management software, on-site monitoring equipment, and manual records. In order to ensure the comprehensiveness and accuracy of the data, it is necessary to establish unified data collection standards and processes to ensure that the data source is reliable and the format is unified.
[0052] Project management software: Obtain data on project progress, cost, quality, etc. through API interface or data export function.
[0053] On-site monitoring equipment: Through the Internet of Things technology, real-time data such as environmental parameters and equipment operating status of the construction site are collected.
[0054] Manual records: including engineering logs, meeting minutes, quality inspection reports, etc., require the establishment of electronic archiving and filing mechanisms.
[0055] According to the source and nature of multi-terminal engineering office data, it can be divided into the following categories:
[0056] Progress data: including project schedule, actual progress records, milestone completion status, etc.
[0057] Cost data: including budget allocation, actual cost expenditure, cost overruns, etc.
[0058] Quality data: including quality inspection reports, quality accident records, quality improvement measures, etc.
[0059] Resource data: including human resource allocation, material supply, equipment usage records, etc.
[0060] Environmental data: including construction site environmental parameters, meteorological conditions, safety monitoring data, etc.
[0061] In this embodiment, redundancy analysis aims to identify and eliminate duplicate and useless information in the data to improve data processing efficiency and storage utilization. Redundancy analysis can be performed using data analysis tools such as FineBI. Integrity analysis aims to ensure that each type of multi-terminal engineering internal data contains all necessary information and that the information is accurate.
[0062] In some embodiments of the present application, the redundancy and integrity of each type of multi-terminal engineering internal data are analyzed, including:
[0063] Divide multi-terminal engineering internal data into structured data and unstructured data;
[0064] For structured data, set the corresponding data unit length according to the standard size;
[0065] For unstructured data, analyze the complexity of each type of multi-terminal engineering internal data and set the corresponding data unit length according to the complexity;
[0066] Check the missing values, empty values and NULL values of each data unit in structured and unstructured data, and define the integrity of each type of multi-terminal engineering internal data based on the missing values, empty values and NULL values;
[0067] Determine the duplication categories between data units in structured data and unstructured data. The duplication categories include complete duplication and partial duplication. For complete duplication, delete multiple data units under complete duplication and only retain one data unit.
[0068] For partial duplication, the data units are divided into different clusters through clustering algorithm, and the overlap between each data unit and other data units in the same cluster is calculated to build a cluster overlap list. The redundancy degree of each data unit is calculated based on the cluster overlap list, so as to determine the redundancy degree of each type of multi-terminal engineering internal data; ;
[0069] in, For the The first The degree of redundancy of each data unit, There are two preset weights, For the The first The maximum overlap of data units, For the The number of data units in a cluster, Indicates The number of data units in a cluster after removing the data unit with the largest overlap, Indicates The first The data units are combined with the data units with the maximum overlap after removing the data units with the The overlap between data units;
[0070] Among them, the redundancy degree and integrity degree of multi-terminal engineering internal data are used to describe the redundancy and integrity of multi-terminal engineering internal data respectively.
[0071] In this embodiment, structured data generally refers to tabular data in a relational database, which has clear fields and records. The standard size of such data is one row or one field, which is the data unit length. Unstructured data includes text files, JSON objects, images, audio, etc., and their formats and contents are more flexible. Different data categories may correspond to different complex situations (the complexity is quantified through historical data), and the corresponding data unit length is set.
[0072] In this embodiment, for each data unit, the missing value, empty value and NULL value are analyzed to define the integrity of each type of multi-terminal engineering internal data. Complete duplication means that two or more data units are exactly the same, and partial duplication means that two or more data units are partially the same. The data units are divided into different clusters through clustering algorithms, and the data units with mutual duplication relationships are divided into the same cluster. The overlap list of clusters is a list of overlaps listed for each data unit and other data units. The overlap can be calculated according to a similarity algorithm (such as edit distance, cosine similarity, etc.).
[0073] In this embodiment, the redundancy degree of a data unit is determined by combining the maximum overlap value in the overlap list and the class average values of other overlap values. Represents the class average of other overlaps (a bit larger than the general average).
[0074] In some embodiments of the present application, the quality of each type of multi-terminal engineering internal data is determined, including:
[0075] The quality of each type of multi-terminal engineering internal data is calculated based on the redundancy and integrity of the multi-terminal engineering internal data.
[0076] In this embodiment, the quality is calculated by comprehensively considering the redundancy and integrity of multi-terminal engineering internal data. It can be calculated by weighted summation or weighted average, or some other methods that comprehensively consider the redundancy and integrity of multi-terminal engineering internal data. All of these methods fall within the scope of protection of this application and are in line with the calculation concept of the quality of multi-terminal engineering internal data.
[0077] Step S102, screening and preprocessing the multi-terminal engineering internal data by the quality of each type of multi-terminal engineering internal data, identifying the spatiotemporal correlation between the multi-terminal engineering internal data, and fusing each type of multi-terminal engineering internal data to obtain a spatiotemporal fusion data set.
[0078] In this embodiment, the quality of the engineering internal data at multiple terminals is controlled, and the data with better quality is preprocessed, including cleaning, filling, etc. The data with very poor quality is eliminated.
[0079] In some embodiments of the present application, the multi-terminal engineering internal data is screened and preprocessed according to the quality of each type of multi-terminal engineering internal data, including:
[0080] The multi-terminal engineering internal data with quality lower than the quality threshold will be deleted, and the multi-terminal engineering internal data with quality not lower than the quality threshold will be pre-processed.
[0081] In some embodiments of the present application, the spatiotemporal association between multi-terminal engineering internal data is identified, and each type of multi-terminal engineering internal data is fused to obtain a spatiotemporal fusion data set, including:
[0082] Use time analysis tools to analyze the changing trends of multi-terminal engineering internal data over time to obtain time relationships, use GIS software to project multi-terminal engineering internal data onto geographic space, and analyze the spatial relationships between multi-terminal engineering internal data on geographic space;
[0083] The multi-terminal engineering internal data are aligned in time and space, and the time features and spatial features are extracted based on the time and spatial relationships of the multi-terminal engineering internal data. The time features and spatial features are fused to obtain a time-space fusion data set.
[0084] In this embodiment, time analysis tools such as time series graphs and moving averages are used to analyze the changing trend of data over time, and to identify long-term trends, seasonal changes, and periodic fluctuations in the data. GIS technology is used to map data to geographic space, and to analyze the distribution patterns and relationships of data in space. This includes the aggregation, dispersion, and spatial autocorrelation of data. Spatial relationships between data are identified, such as adjacent relationships, inclusion relationships, and distance relationships. These relationships are essential for understanding the spatial dependence and interaction of data. Features that can reflect their spatiotemporal characteristics are extracted from each type of data, such as time features (timestamps, cycles, etc.) and spatial features (position, shape, distance, etc.). Appropriate fusion strategies are selected, such as Kalman filtering, Bayesian fusion, etc. These strategies should be able to fully consider the spatiotemporal correlation and uncertainty of data. The extracted features are fused according to the fusion strategy to form a spatiotemporal fusion data set. This includes combining data at different time points and spatial locations, and handling conflicts and redundancies between data.
[0085] Step S103, identifying the correlation between the multi-terminal engineering internal data, and fusing each type of multi-terminal engineering internal data to obtain a related fused data set.
[0086] In this embodiment, the correlation between the internal engineering data at multiple terminals is analyzed, and data fusion is performed accordingly.
[0087] In some embodiments of the present application, the correlation between the multi-terminal engineering internal data is identified, and each type of multi-terminal engineering internal data is fused to obtain a related fused data set, including:
[0088] Extract data features from multi-terminal engineering internal data, perform data standardization on the data features, draw a scatter plot between the data features of the multi-terminal engineering internal data, and analyze the scatter plot to determine the correlation association type between the data features of the multi-terminal engineering internal data, the correlation association type includes linear correlation and nonlinear correlation;
[0089] Calculate the correlation coefficient of the data features under the linear correlation relationship, assign the fusion weight of the data features according to the correlation coefficient, fuse the data features with linear correlation, and obtain the linear correlation fusion set;
[0090] The multiple data features under nonlinear correlation are fused by a preset feature fusion model to obtain a nonlinear correlation fusion set;
[0091] The correlation fusion data set includes two parts: linear correlation fusion set and nonlinear correlation fusion set.
[0092] In this embodiment, a scatter plot of two variables is drawn to observe whether the data points are roughly distributed near a straight line, which is accomplished by setting a similarity threshold for fitting. If so, a linear relationship may exist. Otherwise, it is considered a nonlinear relationship. The correlation coefficient of the data features under the linear correlation relationship, such as the Pearson correlation coefficient, etc. The linear correlation relationship is relatively simple, and the fusion weight of each type of data feature can be directly assigned according to the correlation coefficient. Fusion is performed in a weighted summation manner.
[0093] In this embodiment, nonlinear relationship means that the change between variables is not in a fixed proportion, and may be a curve, surface or other complex form. Therefore, the traditional linear correlation coefficient (such as Pearson correlation coefficient) may not accurately reflect this relationship. Machine learning algorithms, especially neural networks, can automatically learn complex relationships in data, including nonlinear relationships. By training a neural network model, we can extract the implicit relationship between features and calculate the correlation based on it.
[0094] Step S104, combining the spatiotemporal fusion data set and the related fusion data set to generate collaborative fusion engineering information, training the engineering value extraction model based on the collaborative fusion engineering information, and realizing the management of multi-terminal engineering internal data through engineering value.
[0095] In this embodiment, the collaborative fusion engineering information includes and integrates the spatiotemporal fusion data set and the related fusion data set, and the complex situation dimensions (comprehensive information indicators) of the spatiotemporal fusion data set and the related fusion data set. Select appropriate machine learning algorithms (such as regression models, classification models, etc.) to build an engineering value extraction model. The engineering value includes direct value information in multiple dimensions such as the progress, quality, cost, and safety of the project. By referring to these direct value information, the management of multi-terminal engineering internal data is realized.
[0096] In some embodiments of the present application, the spatiotemporal fusion data set and the related fusion data set are combined to generate collaborative fusion engineering information, including:
[0097] Integrate the spatiotemporal fusion data set and the related fusion data set to generate collaborative fusion engineering information, and evaluate the comprehensive information index of the data of the spatiotemporal fusion data set. According to the linear correlation fusion set and the nonlinear correlation fusion set in the related fusion data set, establish a linear relationship network and a nonlinear relationship network respectively, and analyze the linear relationship network and the nonlinear relationship network to determine the comprehensive information index of the data relationship of the linear correlation fusion set and the nonlinear correlation fusion set respectively.
[0098] The comprehensive information indicators of data and the comprehensive information indicators of data relationships are added to the collaborative fusion engineering information.
[0099] In this embodiment, the comprehensive information index of the data of the spatiotemporal fusion data set is evaluated, and the data volume, number of data types, distribution, outliers, etc. of the spatiotemporal fusion data set are comprehensively considered to determine the comprehensive information index of a data, which describes the complexity and difficulty of the data. The data is regarded as nodes, and the relationship between the data is regarded as edges to construct a data network. For linear and nonlinear relationships, network analysis indicators (such as degree, clustering coefficient, path length, etc.) are used to evaluate the complexity and structural characteristics of the data relationship, and the various parameters on the network diagram are comprehensively considered to determine the comprehensive information index of the data relationship (linear relationship and nonlinear relationship), which describes the complexity and difficulty of the data relationship.
[0100] In some embodiments of the present application, the engineering value extraction model is trained based on the collaborative fusion engineering information, including:
[0101] Determine the multi-dimensional information complexity level of collaborative fusion engineering information through the comprehensive information index of data and the comprehensive information index of data relationship in collaborative fusion engineering information; ;
[0102] in, To collaboratively integrate the multi-dimensional information complexity level of engineering information, To coordinate and integrate the comprehensive information indicators of data in engineering information, are the combined weights corresponding to the linear correlation and nonlinear correlation, respectively. are the comprehensive information indicators of the data relationships corresponding to the linear correlation and nonlinear correlation, All are preset constants;
[0103] The model parameters are determined by collaboratively integrating the multi-dimensional information complexity level of the engineering information, and the engineering value extraction model is trained according to the model parameters and the collaboratively integrated engineering information.
[0104] In this embodiment, The complexity of the data relationship is expressed as a correction of the complexity of the data itself, so as to determine a comprehensive complexity impact. Different levels of multidimensional information complexity correspond to different combinations of model parameters, including learning rate, regularization coefficient, number of hidden layer units, etc., so as to construct a targeted adaptive model parameter, which is used in the training process to best fit the parameter combination of the current data complexity.
[0105] Compared with the prior art, the present invention has the following beneficial effects:
[0106] 1. Analyze the redundancy and integrity of each type of multi-terminal engineering internal data, determine the quality of each type of multi-terminal engineering internal data, and screen and pre-process the multi-terminal engineering internal data based on the quality of each type of multi-terminal engineering internal data. Divide the data into data units, analyze the degree of integrity and redundancy, and define the data quality by combining the two, thereby improving the control of data quality and providing a reliable foundation for subsequent data fusion and model training.
[0107] 2. Identify the spatiotemporal associations and correlation associations between multi-terminal engineering internal data to construct spatiotemporal fusion data sets and related fusion data sets, and then generate collaborative fusion engineering information. The collaborative fusion engineering information not only contains the spatiotemporal characteristics and correlation characteristics of multi-terminal engineering internal data, but also reveals the inherent connection and mutual influence between the data, providing rich basic data for subsequent engineering value extraction. In addition, the collaborative fusion engineering information also evaluates the spatiotemporal fusion data set and the related fusion data set in complex situations, that is, the comprehensive information index. It helps to confirm the model parameters, so as to accurately train the engineering value extraction model. It improves the accuracy and adaptability of value information extraction, indirectly assists the management of internal data, and improves the effectiveness of data management.
[0108] Correspondingly, the present application also provides a management system for multi-terminal engineering internal data, such as Figure 2 As shown, including,
[0109] An evaluation module is used to obtain multi-terminal engineering internal data, classify the multi-terminal engineering internal data, analyze the redundancy and integrity of each type of multi-terminal engineering internal data, and determine the quality of each type of multi-terminal engineering internal data;
[0110] The first fusion module is used to screen and preprocess the multi-terminal engineering internal data according to the quality of each type of multi-terminal engineering internal data, identify the spatiotemporal association between the multi-terminal engineering internal data, and fuse each type of multi-terminal engineering internal data to obtain a spatiotemporal fusion data set;
[0111] The second fusion module is used to identify the correlation between the multi-terminal engineering internal data and fuse each type of multi-terminal engineering internal data to obtain a related fusion data set;
[0112] A determination module is used to combine the spatiotemporal fusion data set and the related fusion data set to generate collaborative fusion engineering information, train the engineering value extraction model based on the collaborative fusion engineering information, and realize the management of multi-terminal engineering internal data through engineering value.
[0113] Through the description of the above implementation methods, those skilled in the art can clearly understand that the present invention can be implemented by hardware, or by software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present invention can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each implementation scenario of the present invention.
[0114] Those skilled in the art will appreciate that the accompanying drawings are merely schematic diagrams of a preferred implementation scenario, and the modules or processes in the accompanying drawings are not necessarily required for implementing the present invention.
[0115] Those skilled in the art will appreciate that the modules in the system in the implementation scenario can be distributed in the system of the implementation scenario according to the implementation scenario description, or can be changed accordingly and located in one or more systems different from the implementation scenario. The modules in the above implementation scenario can be combined into one module, or can be further split into multiple submodules.
[0116] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.
Claims
1. A method for managing multi-terminal engineering internal data, characterized in that: include, Obtain multi-terminal engineering internal data, classify the multi-terminal engineering internal data, analyze the redundancy and integrity of each type of multi-terminal engineering internal data, and determine the quality of each type of multi-terminal engineering internal data; Screen and preprocess the multi-terminal engineering internal data by the quality of each type of multi-terminal engineering internal data, identify the spatiotemporal associations between the multi-terminal engineering internal data, and fuse each type of multi-terminal engineering internal data to obtain a spatiotemporal fusion data set; Identify the correlation between multi-terminal engineering internal data, and fuse each type of multi-terminal engineering internal data to obtain relevant fused data sets; Combine the spatiotemporal fusion data set and the related fusion data set to generate collaborative fusion engineering information, train the engineering value extraction model based on the collaborative fusion engineering information, and manage the multi-terminal engineering internal data through engineering value; in, Analyze the redundancy and integrity of each type of multi-terminal engineering office data, including: Divide multi-terminal engineering internal data into structured data and unstructured data; For structured data, set the corresponding data unit length according to the standard size; For unstructured data, analyze the complexity of each type of multi-terminal engineering internal data and set the corresponding data unit length according to the complexity; Check the missing values, empty values and NULL values of each data unit in structured and unstructured data, and define the integrity of each type of multi-terminal engineering internal data based on the missing values, empty values and NULL values; Determine the duplication categories between data units in structured data and unstructured data. The duplication categories include complete duplication and partial duplication. For complete duplication, delete multiple data units under complete duplication and only retain one data unit. For partial duplication, the data units are divided into different clusters through clustering algorithm, and the overlap between each data unit and other data units in the same cluster is calculated to build a cluster overlap list. The redundancy degree of each data unit is calculated based on the cluster overlap list, so as to determine the redundancy degree of each type of multi-terminal engineering internal data; ; in, For the The first The degree of redundancy of each data unit, There are two preset weights, For the The first The maximum overlap of data units, For the The number of data units in a cluster, Indicates The number of data units in a cluster after removing the data unit with the largest overlap, Indicates The first The data units are combined with the data units with the maximum overlap after removing the data units with the The overlap between data units; Among them, the redundancy degree and integrity degree of multi-terminal engineering internal data are used to describe the redundancy and integrity of multi-terminal engineering internal data respectively; And determine the quality of each type of multi-terminal engineering office data, including, The quality of each type of multi-terminal engineering internal data is calculated by comprehensively considering the redundancy and integrity of the multi-terminal engineering internal data; Combine spatiotemporal fusion datasets and related fusion datasets to generate collaborative fusion engineering information, including, Integrate the spatiotemporal fusion data set and the related fusion data set to generate collaborative fusion engineering information, and evaluate the comprehensive information index of the data of the spatiotemporal fusion data set. According to the linear correlation fusion set and the nonlinear correlation fusion set in the related fusion data set, establish a linear relationship network and a nonlinear relationship network respectively, and analyze the linear relationship network and the nonlinear relationship network to determine the comprehensive information index of the data relationship of the linear correlation fusion set and the nonlinear correlation fusion set respectively. Adding comprehensive information indicators of data and comprehensive information indicators of data relationships to collaborative fusion engineering information; Training engineering value extraction models based on collaborative fusion engineering information includes: Determine the multi-dimensional information complexity level of collaborative fusion engineering information through the comprehensive information index of data and the comprehensive information index of data relationship in collaborative fusion engineering information; ; in, To collaboratively integrate the multi-dimensional information complexity level of engineering information, To coordinate and integrate the comprehensive information indicators of data in engineering information, are the combined weights corresponding to the linear correlation and nonlinear correlation, respectively. are the comprehensive information indicators of the data relationships corresponding to the linear correlation and nonlinear correlation, All are preset constants; The model parameters are determined by collaboratively integrating the multi-dimensional information complexity level of the engineering information, and the engineering value extraction model is trained according to the model parameters and the collaboratively integrated engineering information.
2. The method for managing multi-terminal engineering internal data according to claim 1 is characterized in that: Screen and pre-process the multi-terminal engineering internal data based on the quality of each type of multi-terminal engineering internal data. include, The multi-terminal engineering internal data with quality lower than the quality threshold will be deleted, and the multi-terminal engineering internal data with quality not lower than the quality threshold will be pre-processed.
3. The method for managing multi-terminal engineering internal data according to claim 1 is characterized in that: Identify the spatiotemporal correlation between multi-terminal engineering internal data, and fuse each type of multi-terminal engineering internal data to obtain a spatiotemporal fusion data set, including: Use time analysis tools to analyze the changing trends of multi-terminal engineering internal data over time to obtain time relationships, use GIS software to project multi-terminal engineering internal data onto geographic space, and analyze the spatial relationships between multi-terminal engineering internal data on geographic space; The multi-terminal engineering internal data are aligned in time and space, and the time features and spatial features are extracted based on the time and spatial relationships of the multi-terminal engineering internal data. The time features and spatial features are fused to obtain a time-space fusion data set.
4. The method for managing multi-terminal engineering internal data according to claim 1, characterized in that: Identify the correlation between multi-terminal engineering internal data, and fuse each type of multi-terminal engineering internal data to obtain relevant fused data sets. include, Extract data features from multi-terminal engineering internal data, perform data standardization on the data features, draw a scatter plot between the data features of the multi-terminal engineering internal data, and analyze the scatter plot to determine the correlation association type between the data features of the multi-terminal engineering internal data, the correlation association type includes linear correlation and nonlinear correlation; Calculate the correlation coefficient of the data features under the linear correlation relationship, assign the fusion weight of the data features according to the correlation coefficient, fuse the data features with linear correlation, and obtain the linear correlation fusion set; The multiple data features under nonlinear correlation are fused by a preset feature fusion model to obtain a nonlinear correlation fusion set; The correlation fusion data set includes two parts: linear correlation fusion set and nonlinear correlation fusion set.
5. A management system for multi-terminal engineering internal data, characterized in that: Used to implement the management method for multi-terminal engineering internal data as described in any one of claims 1 to 4, the system includes: An evaluation module is used to obtain multi-terminal engineering internal data, classify the multi-terminal engineering internal data, analyze the redundancy and integrity of each type of multi-terminal engineering internal data, and determine the quality of each type of multi-terminal engineering internal data; The first fusion module is used to screen and preprocess the multi-terminal engineering internal data according to the quality of each type of multi-terminal engineering internal data, identify the spatiotemporal association between the multi-terminal engineering internal data, and fuse each type of multi-terminal engineering internal data to obtain a spatiotemporal fusion data set; The second fusion module is used to identify the correlation between the multi-terminal engineering internal data and fuse each type of multi-terminal engineering internal data to obtain a related fusion data set; A determination module is used to combine the spatiotemporal fusion data set and the related fusion data set to generate collaborative fusion engineering information, train the engineering value extraction model based on the collaborative fusion engineering information, and realize the management of multi-terminal engineering internal data through engineering value.
Citation Information
Patent Citations
Airborne LiDAR point cloud data vulnerability rapid detection method based on density histogram
CN110008207A
System and method for automated data harmonization
US20240020292A1