Data-driven digital twin model construction processing method

By going through data collection, preprocessing, analysis, and fusion steps, the problems of data duplication and ambiguity in digital twin models have been solved, achieving accuracy and uniqueness in model construction.

CN120873458APending Publication Date: 2025-10-31GUIZHOU INST OF GEOLOGY & MINERAL SURVEYING & MAPPING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510963142.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-14
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

In the existing construction of pipeline network digital twin models driven by multi-source heterogeneous data, data duplication and ambiguity are prone to occur, leading to inaccurate model construction.

Method used

Data accuracy is ensured through steps including data collection, preprocessing, preliminary fusion, analysis to remove duplicate data, judgment and processing of fuzzy data, and deep fusion.

Benefits of technology

It improves the accuracy of digital twin model construction by manually judging fuzzy data to ensure that each entity in the model has a unique data representation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120873458A_ABST
    Figure CN120873458A_ABST
Patent Text Reader

Abstract

The invention relates to a data-driven digital twinborn model construction processing method. The method comprises the following steps: collecting data for constructing a digital twinborn model; the collected digital twinborn model data are preprocessed, and preliminary fusion of the data is carried out; analyzing the data based on the fusion process, deleting repeated data, and performing secondary analysis; extracting fuzzy data existing in the analysis process, and performing judgment processing on the fuzzy data; and traversing the digital twin model data, and after checking and confirming that no repeated data and fuzzy data exist, carrying out deep data fusion. According to the method, the repeated data and the fuzzy data can be analyzed and extracted, the repeated data can be deleted, and the entity is ensured to have unique model data correspondingly, so that the accuracy of model construction is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of digital twin model technology, specifically, it relates to a data-driven digital twin model construction and processing method. Background Technology

[0002] Digital twins refer to the digital modeling and simulation of real-world entities, systems, or processes to better understand, analyze, and optimize their behavior. Multi-source heterogeneous data fusion integrates data from different sources and of different types to provide more comprehensive and accurate information to support the creation and updating of digital twin models. In this context, modern society generates a vast amount of data, which may include structured, semi-structured, and unstructured data. Multi-source heterogeneous digital twin technology can be applied to various fields, such as urban construction management, healthcare, and e-commerce. To build accurate and comprehensive digital twin models, it is necessary to integrate this multi-source heterogeneous data, which helps eliminate information silos and gain deeper insights.

[0003] Furthermore, existing multi-source heterogeneous data-driven digital twin model construction for pipeline networks often suffers from duplication and ambiguity due to the excessive amount of data stored. Specific instances of duplication may include: different collected data information; different sampling methods used for the same model unit, resulting in the same model unit existing in different data sources; and different model units having essentially the same structure but existing simultaneously in reality. This leads to data duplication and ambiguity, which existing processing systems cannot adequately handle, resulting in inaccuracies in the construction of digital twin models.

[0004] In view of this, the present invention is proposed. Summary of the Invention

[0005] The purpose of this invention is to provide a data-driven digital twin model construction and processing method that can process repetitive and ambiguous data in the digital twin model and improve the accuracy of model construction.

[0006] The basic concept of the technical solution adopted in this invention is:

[0007] A data-driven digital twin model construction method includes the following steps:

[0008] S1. Collect data for constructing the digital twin model;

[0009] S2. Preprocess the collected digital twin model data and perform initial data fusion;

[0010] S3. Analyze the data based on the fusion process, delete duplicate data, and perform secondary analysis;

[0011] S4. Extract fuzzy data that exists during the analysis process, and judge and process the fuzzy data;

[0012] S5. After traversing the digital twin model data and checking to confirm that there is no duplicate or ambiguous data, perform deep data fusion.

[0013] According to the data-driven digital twin model construction and processing method described above, in step S1, the step of collecting data for constructing the digital twin model includes:

[0014] Determine the content of the digital twin model that needs to be built;

[0015] Obtain the relevant data for the constructed digital twin model.

[0016] According to the data-driven digital twin model construction and processing method described above, in S2, the step of preprocessing the collected digital twin model data includes: data cleaning, data denoising, and missing data filling of the acquired data.

[0017] According to the data-driven digital twin model construction and processing method described above, in S2, the step of preliminary data fusion includes:

[0018] The data required for building a digital twin model is categorized, and sub-units are constructed:

[0019] Classification: Data A, Data B, Data C... Data N;

[0020] Sub-unit: Data A: A1, A2, A3...An;

[0021] Data B: B1, B2, B3...Bn;

[0022] Data C: C1, C2, C3...Cn; ......;

[0024] Data N: N1, N2, N3, ..., Nn

[0025] The preprocessed data is matched one by one to its category and its sub-unit, and then the data is filled into the sub-unit.

[0026] According to the data-driven digital twin model construction and processing method described above, in S3, the steps of analyzing the data based on the fusion process, deleting duplicate data, and performing secondary analysis include:

[0027] When new data is entered into a sub-cell, the duplication rate between the new data and the already entered data is calculated.

[0028] The first analysis checks whether the duplication rate is greater than the first threshold. If so, the newly entered data is recorded as duplicate data and then deleted.

[0029] The second analysis checks whether the repetition rate is greater than the second threshold and not greater than the first threshold. If so, the newly entered data is recorded as fuzzy data.

[0030] According to the data-driven digital twin model construction and processing method described above, in S4, the step of judging and processing fuzzy data includes: parsing the fuzzy data, marking the coordinates of data points that are duplicated, and feeding the fuzzy data back to the staff after marking. The staff judges the fuzzy data. If it is judged to be duplicated data, the staff deletes it directly. If it is judged not to be duplicated data, it is returned directly.

[0031] According to the data-driven digital twin model construction and processing method described above, in S5, the step of deep data fusion includes: based on existing non-repeating data, extracting the coordinates of data points in each sub-unit data, merging and aligning the models corresponding to the sub-unit data, thereby constructing a set model.

[0032] Compared with the prior art, the present invention has the following advantages:

[0033] This invention, through the construction of a digital model, can extract ambiguous data when there is ambiguous information in the fusion process, and feed the comparison content back to the staff. The staff judges the ambiguous data. If it is judged to be duplicate data, the staff deletes it directly. If it is judged to be non-duplicate data, it can be returned directly. Therefore, the accuracy of judging ambiguous information is guaranteed to a certain extent.

[0034] This invention can analyze and extract duplicate and ambiguous data, and delete duplicate data to ensure that each entity has unique model data, thereby improving the accuracy of model construction.

[0035] The specific embodiments of the present invention will now be described in further detail with reference to the accompanying drawings. Attached Figure Description

[0036] In the attached diagram:

[0037] Figure 1 A schematic diagram illustrating the steps involved in constructing a data-driven digital twin model. Detailed Implementation

[0038] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments will be clearly and completely described below with reference to the accompanying drawings. The following embodiments are used to illustrate the present invention.

[0039] like Figure 1 As shown, a data-driven digital twin model construction method includes the following steps:

[0040] S1. Collect data for constructing the digital twin model;

[0041] S2. Preprocess the collected digital twin model data and perform initial data fusion;

[0042] S3. Analyze the data based on the fusion process, delete duplicate data, and perform secondary analysis;

[0043] S4. Extract fuzzy data that exists during the analysis process, and judge and process the fuzzy data;

[0044] S5. After traversing the digital twin model data and checking to confirm that there is no duplicate or ambiguous data, perform deep data fusion.

[0045] The constructed digital model can extract ambiguous information during the data fusion process. The extracted data, which is unclear whether it is duplicated, is then compared with the data. This comparison is fed back to the staff, who then judge the ambiguous data. If it is determined to be duplicated, it is deleted directly; if it is determined not to be duplicated, it can be returned directly. Therefore, the accuracy of judging ambiguous information is guaranteed to a certain extent.

[0046] In a specific implementation, step S1, which involves collecting data for constructing the digital twin model, includes:

[0047] The staff determines the content of the digital twin model to be built (i.e., what the model is, its components, the source of data collection, the accuracy requirements of the model, etc.);

[0048] Staff members obtain relevant data for the constructed digital twin model; the relevant data includes information such as attributes and data point coordinates.

[0049] Furthermore, the method steps for preprocessing the collected digital twin model data based on the above S2 include: data cleaning, data denoising, and missing data imputation of the acquired data.

[0050] Furthermore, in S2, the preprocessing steps for the collected digital twin model data include:

[0051] Sub-units are constructed based on the different data acquired;

[0052] Classification: Data A, Data B, Data C... Data N;

[0053] Sub-unit: Data A: A1, A2, A3...An;

[0054] Data B: B1, B2, B3...Bn;

[0055] Data C: C1, C2, C3...Cn; ......;

[0057] Data N: N1, N2, N3, ..., Nn;

[0058] The preprocessed data is matched one by one to its category and its sub-unit, and then the data is filled into the sub-unit.

[0059] Specifically, the first preprocessed data is selected, its category is matched, and then its sub-unit is matched within the matched category. After matching, the data is filled into the sub-unit; the same operation is performed on the other data in sequence.

[0060] Furthermore, in S3, the steps of analyzing data based on the fusion process, deleting duplicate data, and performing secondary analysis include:

[0061] During the data entry process, when new data is entered into a sub-cell, the duplication rate between the newly entered data and the already entered data is calculated.

[0062] The first analysis checks if the duplication rate exceeds the first threshold; if so, the newly entered data is recorded as duplicate data and deleted; if not, a second analysis is performed.

[0063] The second analysis checks whether the repetition rate is greater than the second threshold and not greater than the first threshold; if so, the newly entered data is recorded as fuzzy data; if not, the new data is entered directly.

[0064] Among them, 0.85≤first threshold≤0.95, 0.6≤second threshold≤0.8; in actual operation, the first threshold and the second threshold are selected from this range and then calculated and analyzed.

[0065] Furthermore, in S4, the steps for judging and processing fuzzy data include: when there is fuzzy data in the data fusion, parsing the fuzzy data, marking the coordinates of duplicate data points, and feeding back the fuzzy data to the staff. The staff judges the fuzzy data. If it is judged to be duplicate data, the staff deletes it directly. If it is judged not to be duplicate data, it can be returned directly.

[0066] In practice, the fuzzy data returned is usually a dataset containing numerous data point coordinates. Real-world entities often share similar structural forms and spatial coordinates. Due to sampling accuracy limitations, two scenarios arise when constructing digital twin models: First, the model datasets corresponding to two entities contain many identical data point coordinates, making it difficult for the system to distinguish them as two entity structures. Second, different sampling devices sample the same entity structure, but the resulting datasets contain many non-overlapping data point coordinates, making it difficult for the system to determine whether it represents the same entity structure or two or more entities. These scenarios result in fuzzy data, requiring feedback to staff who, in conjunction with other data (such as field survey data), manually analyze whether it corresponds to one or multiple models. If it's determined to be multiple models, the data is retained and returned; if it's determined to be one model, duplicate data is deleted.

[0067] Furthermore, in S5, the steps for deep data fusion include: extracting the coordinates of data points from each sub-unit data based on existing non-repeating data, merging and aligning the models corresponding to the sub-unit data, thereby constructing a ensemble model.

[0068] Specifically, each sub-unit corresponds to a sub-model. After extracting the coordinates of the data points of the sub-unit, the spatial shape and spatial position of the edge of the sub-model can be known. When merging and aligning the models corresponding to the sub-unit data, the edges of adjacent sub-models are overlapped to connect the adjacent sub-models, thereby constructing a set model.

[0069] The embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of the present invention.

Claims

1. A data-driven digital twin model construction and processing method, characterized in that, Includes the following steps: S1. Collect data for constructing the digital twin model; S2. Preprocess the collected digital twin model data and perform initial data fusion; S3. Analyze the data based on the fusion process, delete duplicate data, and perform secondary analysis; S4. Extract fuzzy data that exists during the analysis process, and judge and process the fuzzy data; S5. After traversing the digital twin model data and checking to confirm that there is no duplicate or ambiguous data, perform deep data fusion.

2. The data-driven digital twin model construction and processing method according to claim 1, characterized in that, In S1, the steps for collecting data to construct the digital twin model include: Determine the content of the digital twin model that needs to be built; Obtain the relevant data for the constructed digital twin model.

3. The data-driven digital twin model construction and processing method according to claim 2, characterized in that, In S2, the steps for preprocessing the collected digital twin model data include: data cleaning, data denoising, and missing data imputation.

4. The data-driven digital twin model construction and processing method according to claim 3, characterized in that, In S2, the steps for preliminary data fusion include: The data required for building a digital twin model is categorized, and sub-units are constructed: Classification: Data A, Data B, Data C... Data N; Sub-unit: Data A: A1, A2, A3...An; Data B: B1, B2, B3...Bn; Data C: C1, C2, C3...Cn; ......; Data N: N1, N2, N3, ..., Nn The preprocessed data is matched one by one to its category and its sub-unit, and then the data is filled into the sub-unit.

5. The data-driven digital twin model construction and processing method according to claim 4, characterized in that, In S3, the steps for analyzing data based on the fusion process, deleting duplicate data, and performing secondary analysis include: When new data is entered into a sub-cell, the duplication rate between the new data and the already entered data is calculated. The first analysis checks whether the duplication rate is greater than the first threshold. If so, the newly entered data is recorded as duplicate data and then deleted. The second analysis checks whether the repetition rate is greater than the second threshold and not greater than the first threshold. If so, the newly entered data is recorded as fuzzy data.

6. The data-driven digital twin model construction and processing method according to claim 1, characterized in that, In S4, the steps for judging and processing fuzzy data include: parsing the fuzzy data, marking the coordinates of data points that are duplicated, and then feeding the fuzzy data back to the staff. The staff judges the fuzzy data. If it is judged to be duplicated, the staff deletes it directly. If it is judged not to be duplicated, it is returned directly.

7. The data-driven digital twin model construction and processing method according to claim 1, characterized in that, In S5, the steps for deep data fusion include: extracting the coordinates of data points from each sub-unit based on existing non-repeating data, merging and aligning the models corresponding to the sub-unit data, thereby constructing a ensemble model.