Twin multi-data processing method and system for digital twin model
By designing the multi-protocol support, data preprocessing, data processing core and data storage and management module of the digital twin model system, the problems of poor data transmission and missing data in the existing system when processing large amounts of data are solved, and efficient data processing and high-precision digital model construction are realized.
Patent Information
- Application Number
- CN202510159178.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-05-16
AI Technical Summary
When processing large amounts of data, existing digital twin model systems have problems such as poor data transmission, missing data, long model training time and overfitting, making it difficult to effectively process new data.
A twin multi-data processing method and system for digital twin models is designed, including multi-protocol support module, data preprocessing module, data processing core module and data storage and management module. The system uses multi-protocol support modules to process data from different protocols, the data preprocessing module performs data cleaning and format conversion, the data processing core module uses distributed computing and parallel processing algorithms to integrate data, and the data storage and management module performs distributed storage and data indexing.
Through distributed computing and parallel processing algorithms, data processing efficiency is significantly improved, data redundant processing time is reduced, high-quality data resources are integrated, potential correlations of multi-source data are mined, data credibility and accuracy are improved, and high-precision digital model is built.
Smart Images

Figure CN120011440A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of multi-data processing, and specifically relates to a twin multi-data processing method and system for a digital twin model. Background Art
[0002] Digital twins make full use of data such as physical models, sensor updates, and operation history, integrate multi-disciplinary, multi-physical quantity, multi-scale, and multi-probability simulation processes, and complete mapping in virtual space, thereby reflecting the entire life cycle of the corresponding physical equipment. In the process of building a digital twin model, a lot of data needs to be collected through many sensors. However, in a digital twin system, it may contain multiple subsystems, such as a data acquisition subsystem, a data storage subsystem, and a data processing subsystem. If there is a problem with the interface integration between these subsystems, it may lead to poor data transmission, and the collected data cannot be correctly transmitted to the storage device, which is easy to cause data loss.
[0003] At the same time, existing data processing algorithms are designed based on small-scale data. When faced with large amounts of data, these algorithms may fail, and it may be difficult for the system to obtain a large amount of high-quality labeled data. Moreover, as the amount of data continues to increase, the time for model training will increase significantly, and problems such as overfitting may occur, making the model unable to effectively process new data. Summary of the invention
[0004] The technical problem to be solved by the present invention is to overcome the shortcomings of the above-mentioned prior art and provide a twin multi-data processing method and system for a digital twin model.
[0005] The technical solution adopted to solve the above technical problems is: a twin multi-data processing method and system for a digital twin model, including four subsystems: a multi-protocol support module, a data preprocessing module, a data processing core module, and a data storage and management module. The multi-protocol support module is responsible for communicating with various data sources and external systems to ensure that data conforming to different protocols can be received and sent; the data preprocessing module is responsible for converting data in different formats into a unified format; the data processing core module is responsible for fusing the data transmitted by the data preprocessing module; the data storage and management module is responsible for storing the status data and historical data of the digital twin model.
[0006] The multi-protocol support module scans the incoming data connection request. The scanning process checks the specific fields in the data packet header. The scanned protocol feature information will be matched with the protocol registration library inside the module. The multi-protocol support module will start the corresponding protocol adaptation layer, and the adaptation layer will establish a connection with the data source according to the requirements of the protocol while converting the data format. After the connection is successfully established, the multi-protocol support module starts to receive data from the data source, and the received data will be converted into a unified data format inside the system and passed to the next module or data buffer unit.
[0007] The multi-protocol support module includes a data buffer unit for temporarily storing a large amount of incoming data. When the amount of data increases sharply, the data buffer unit first stores the data and then gradually passes the data to the subsequent data preprocessing module according to the system's processing capacity to avoid data loss and system crash.
[0008] The data preprocessing module includes a data cleaning unit, a data standardization and normalization unit, and a data integration and conversion unit. The data cleaning unit determines whether there is missing data by checking the value of each column in the data set, and fills the missing values with the mean and median according to the distribution of the data.
[0009] The data standardization and normalization unit calculates the mean μ and standard deviation σ of each column that needs to be standardized using the formula:
[0010]
[0011] Convert each data point X to a standardized Z value, where the z value indicates how many standard deviations the data point is from the mean;
[0012] Find the minimum value min and maximum value max of each column to be normalized, using the formula
[0013]
[0014] Convert each data point X to a normalized value X norm , which ranges from 0 to 1;
[0015] The data integration and conversion unit matches and merges the attributes of the same entity in different data sources, and converts data of different data types into a type suitable for subsequent analysis.
[0016] The data processing core module includes a distributed computing framework unit, a parallel processing algorithm unit and a data fusion algorithm unit. The distributed computing framework unit adopts a distributed computing framework to decompose and distribute data processing tasks to multiple computing node clusters. Each cluster is responsible for processing data in a part of the area and then summarizing the results.
[0017] The parallel processing algorithm unit can utilize multiple processors or computing cores to process data simultaneously. The parallel processing algorithm unit divides a large amount of data into multiple small blocks and processes them simultaneously.
[0018] The data fusion algorithm unit calculates the correlation between data from different data sources through the Pearson correlation coefficient, using the formula:
[0019]
[0020] γ is the Pearson correlation coefficient, where n is the number of data points, and are the means of X and Y respectively. The correlation threshold is set according to the application requirements. When the correlation coefficient is higher than the threshold, it is considered that the data has a strong correlation and further fusion is performed;
[0021] The weight of each data source in the fusion is determined according to the quality, reliability and importance of the data. For the data x1, x2, ..., x n , whose weights are ω1, ω2, …, ω n (and ), the fused result y is calculated by the formula y=ω1x1+ω2x2+…+ω n x n calculate;
[0022] When performing data fusion, the state of the system is estimated and predicted based on the dynamic model of the system. Assume that the state equation of the system is
[0023] x k =Ax k-1 +Bμ k-1 +ω k-1
[0024] where x k is the system state vector at time k, A is the state transfer matrix, B is the control input matrix, μ k-1 is the control input vector, ω k-1 is the process noise vector, which predicts the state at the current moment through the state estimation at the previous moment and the known system model;
[0025] At the same time, according to the observation equation
[0026] z k =Hx k +v k
[0027] where z k is the observation vector at time k, H is the observation matrix, v kis the observation noise vector, combining the predicted state and the actual observation value, using the Kalman gain K k To update the state estimate, the Kalman gain is calculated as
[0028] K k =P k / k-1 H T (HP k / k-1 H T +R) -1
[0029] Where P k / k-1 is the prediction covariance matrix, R is the observation noise covariance matrix, and the updated state estimation formula is
[0030] x k / k =x k / k-1 +K k (z k -Hx k / k-1 )
[0031] Through continuous prediction and observation updates, dynamic data fusion and state estimation are achieved.
[0032] The data storage and management module includes a distributed storage unit and a data index and metadata management unit. The distributed storage unit stores data in a dispersed manner on multiple storage nodes to improve storage capacity and data reliability.
[0033] The data index and metadata management unit establishes an effective index to quickly locate required data, while recording information on the source, collection time and processing history of the data.
[0034] The specific steps include:
[0035] Step 1: For physical entities, deploy a variety of sensors in key locations and key processes, collect relevant data from IoT devices and network platforms, and mine and integrate historical data related to physical entities, including past operation records, maintenance data, and fault logs. These historical data provide a basis for analyzing long-term trends and behavior patterns of physical entities.
[0036] Step 2: Identify and process outliers in the data through statistical analysis methods, delete outliers, replace them with adjacent values, and perform corrections based on model prediction. For data with missing values, use the mean and median to fill in the missing values, extract the features of data from different data sources, and then fuse these features. Then fuse the original data from different data sources, arrange them in a certain order to form a new data set, and finally use weighted average to synthesize the results to obtain the final data fusion result;
[0037] Step 3: For large-scale multi-source data, use distributed storage technology to store the data in multiple nodes. For structured data, use the relational database MySQL, and for semi-structured and unstructured data, use the NoSQL database.
[0038] The beneficial effects of the present invention are as follows: the present invention significantly improves the processing efficiency of data information through distributed computing, parallel processing algorithms and fusion processing of data, enhances the ability to mine data value, and reduces the time for redundant data processing. At the same time, fusion processing also integrates high-quality data resources and mines potential correlations between multi-source data. The system can process multiple data through the data processing core module, and thus can build a more complete device status portrait. At the same time, the system uses the inherent logical relationship between different data to verify each other, improves data credibility and accuracy, and uses data analysis algorithms to deeply mine massive data, thereby building a high-precision digital model. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 It is a schematic diagram of the module flow of the present invention;
[0040] Figure 2 It is a schematic diagram of a system page of the present invention;
[0041] Figure 3 It is a schematic diagram of the system operation of the present invention. DETAILED DESCRIPTION
[0042] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0043] like Figure 1 As shown, a twin multi-data processing method and system of a digital twin model of the present embodiment includes four subsystems: a multi-protocol support module, a data preprocessing module, a data processing core module, and a data storage and management module. The multi-protocol support module is responsible for communicating with various data sources and external systems to ensure that data conforming to different protocols can be received and sent; the data preprocessing module is responsible for converting data in different formats into a unified format; the data processing core module is responsible for fusing the data transmitted by the data preprocessing module; the data storage and management module is responsible for storing the status data and historical data of the digital twin model.
[0044] The multi-protocol support module scans the incoming data connection requests. The scanning process checks the specific fields in the data packet header. The scanned protocol feature information will be matched with the protocol registration library inside the module. The multi-protocol support module will start the corresponding protocol adaptation layer, and the adaptation layer will establish a connection with the data source according to the requirements of the protocol while converting the data format. After the connection is successfully established, the multi-protocol support module starts to receive data from the data source. The received data will be converted into a unified data format within the system and passed to the next module or data buffer unit.
[0045] The multi-protocol support module includes a data buffer unit for temporarily storing a large amount of incoming data. When the amount of data increases sharply, the data buffer unit first stores the data and then gradually passes the data to the subsequent data preprocessing module according to the system's processing capacity to avoid data loss and system crash.
[0046] The data preprocessing module includes a data cleaning unit, a data standardization and normalization unit, and a data integration and conversion unit. The data cleaning unit determines whether there is missing data by checking the value of each column in the data set. According to the distribution of the data, the mean and median are used to fill the missing values. Various sensors are widely deployed on physical entities related to fire protection. Temperature sensors, smoke sensors and gas sensors are installed in buildings to monitor environmental parameters in real time. Pressure sensors and flow sensors are installed on fire protection equipment to obtain the operating status information of the equipment. Fire fighters are equipped with positioning sensors and physiological status sensors to understand their location and physical condition. These sensors can be used to collect raw data.
[0047] Data Standardization and Normalization Unit For each column that needs to be standardized, calculate its mean μ and standard deviation σ using the formula:
[0048]
[0049] Convert each data point X to a standardized Z value, where the z value indicates how many standard deviations the data point is from the mean;
[0050] Find the minimum value min and maximum value max of each column to be normalized, using the formula
[0051]
[0052] Convert each data point X to a normalized value X norm , which ranges from 0 to 1;
[0053] The data integration and conversion unit matches and merges the attributes of the same entity in different data sources, and converts data of different data types into a type suitable for subsequent analysis.
[0054] The data processing core module includes a distributed computing framework unit, a parallel processing algorithm unit, and a data fusion algorithm unit. The distributed computing framework unit uses a distributed computing framework to decompose and distribute data processing tasks to multiple computing node clusters. Each cluster is responsible for processing data in a certain area and then summarizing the results.
[0055] The parallel processing algorithm unit can use multiple processors or computing cores to process data at the same time. The parallel processing algorithm unit divides a large amount of data into multiple small blocks and processes them simultaneously.
[0056] The data fusion algorithm unit calculates the correlation between data from different data sources through the Pearson correlation coefficient, using the formula:
[0057]
[0058] γ is the Pearson correlation coefficient, where n is the number of data points, and are the means of X and Y respectively. The correlation threshold is set according to the application requirements. When the correlation coefficient is higher than the threshold, it is considered that the data has a strong correlation and further fusion is performed;
[0059] The weight of each data source in the fusion is determined according to the quality, reliability and importance of the data. For the data x1, x2, ..., x n , whose weights are ω1, ω2, …, ω n (and ), the fused result y is calculated by the formula y=ω1x1+ω2x2+…+ω n x n calculate;
[0060] When performing data fusion, the state of the system is estimated and predicted based on the dynamic model of the system. Assume that the state equation of the system is
[0061] x k =Ax k-1 +Bμ k-1 +ω k-1
[0062] where x k is the system state vector at time k, A is the state transfer matrix, B is the control input matrix, μ k-1 is the control input vector, ω k-1 is the process noise vector, which predicts the state at the current moment through the state estimation at the previous moment and the known system model;
[0063] At the same time, according to the observation equation
[0064] zk =Hx k +v k
[0065] where z k is the observation vector at time k, H is the observation matrix, v k is the observation noise vector, combining the predicted state and the actual observation value, using the Kalman gain K k To update the state estimate, the Kalman gain is calculated as
[0066] K k =P k / k-1 H T (HP k / k-1 H T +R) -1
[0067] Where P k / k-1 is the prediction covariance matrix, R is the observation noise covariance matrix, and the updated state estimation formula is
[0068] x k / k =x k / k-1 +K k (z k -Hx k / k-1 )
[0069] Through continuous prediction and observation updates, dynamic data fusion and state estimation are achieved;
[0070] The data processing core module processes and analyzes real-time data quickly, can monitor the operating status of each part of the fire protection system in real time, determine whether any abnormal situation occurs, analyze smoke and temperature data in real time, determine whether the fire warning threshold is reached, and predict the development trend of the fire based on historical data and real-time data, including the direction and speed of fire spread, and possible affected areas. The data processing core module uses multiple algorithms to fuse and analyze the operating data of firefighting equipment, the distribution and status data of firefighters, etc., to evaluate whether the allocation of firefighting resources is reasonable and whether it can meet current firefighting needs, and provide a basis for optimizing the allocation of firefighting resources. It can process a large amount of information from all aspects at the same time in a short period of time, and then obtain the final data fusion result.
[0071] The data storage and management module includes a distributed storage unit and a data index and metadata management unit. The distributed storage unit stores data in multiple storage nodes to improve storage capacity and data reliability.
[0072] The data index and metadata management unit establishes an effective index to quickly locate the required data, while recording information on the data's source, collection time, and processing history.
[0073] The specific steps include:
[0074] Step 1: For physical entities, deploy a variety of sensors in key locations and key processes, collect relevant data from IoT devices and network platforms, and mine and integrate historical data related to physical entities, including past operation records, maintenance data, and fault logs. These historical data provide a basis for analyzing long-term trends and behavior patterns of physical entities.
[0075] Step 2: Identify and process outliers in the data through statistical analysis methods, delete outliers, replace them with adjacent values, and perform corrections based on model prediction. For data with missing values, use the mean and median to fill in the missing values, extract the features of data from different data sources, and then fuse these features. Then fuse the original data from different data sources, arrange them in a certain order to form a new data set, and finally use weighted average to synthesize the results to obtain the final data fusion result;
[0076] Step 3: For large-scale multi-source data, use distributed storage technology to store the data in multiple nodes. For structured data, use the relational database MySQL, and for semi-structured and unstructured data, use the NoSQL database.
[0077] The above are only preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention.
Claims
1. A twin multi-data processing system of a digital twin model, characterized by: It includes four subsystems: multi-protocol support module, data pre-processing module, data processing core module, and data storage and management module. The multi-protocol support module is responsible for communicating with various data sources and external systems to ensure that data that conforms to different protocols can be received and sent; The data preprocessing module is responsible for converting data in different formats into a unified format; the data processing core module is responsible for integrating the data transmitted by the data preprocessing module; The data storage and management module is responsible for storing the status data and historical data of the digital twin model.
2. The twin multi-data processing system of a digital twin model according to claim 1, characterized in that: The multi-protocol support module scans the incoming data connection request. The scanning process checks the specific fields in the data packet header. The scanned protocol feature information will be matched with the protocol registration library inside the module. The multi-protocol support module will start the corresponding protocol adaptation layer, and the adaptation layer will establish a connection with the data source according to the requirements of the protocol while converting the data format. After the connection is successfully established, the multi-protocol support module starts to receive data from the data source, and the received data will be converted into a unified data format inside the system and passed to the next module or data buffer unit.
3. The twin multi-data processing system of a digital twin model according to claim 2, characterized in that: The multi-protocol support module includes a data buffer unit for temporarily storing a large amount of incoming data. When the amount of data increases sharply, the data buffer unit first stores the data and then gradually transfers the data to the subsequent data pre-processing module according to the processing capacity of the system.
4. The twin multi-data processing system of a digital twin model according to claim 1, characterized in that: The data preprocessing module includes a data cleaning unit, a data standardization and normalization unit, and a data integration and conversion unit. The data cleaning unit determines whether there is missing data by checking the value of each column in the data set, and fills the missing values with the mean and median according to the distribution of the data.
5. The twin multi-data processing system of a digital twin model according to claim 4, characterized in that: The data standardization and normalization unit calculates the mean μ and standard deviation σ of each column that needs to be standardized using the formula: Convert each data point X into a standardized Z value, where the Z value indicates how many standard deviations the data point is from the mean; Find the minimum value min and maximum value max of each column to be normalized, using the formula, Convert each data point X to a normalized value X norm , which ranges from 0 to 1; The data integration and conversion unit matches and merges the attributes of the same entity in different data sources, and converts data of different data types into a type suitable for subsequent analysis.
6. The twin multi-data processing system of a digital twin model according to claim 5, characterized in that: The data processing core module includes a distributed computing framework unit, a parallel processing algorithm unit and a data fusion algorithm unit. The distributed computing framework unit adopts a distributed computing framework to decompose and distribute data processing tasks to multiple computing node clusters. Each cluster is responsible for processing data in a part of the area and then summarizing the results. The parallel processing algorithm unit can utilize multiple processors or computing cores to process data simultaneously. The parallel processing algorithm unit divides a large amount of data into multiple small blocks and processes them simultaneously.
7. The twin multi-data processing system of a digital twin model according to claim 6, characterized in that: The data fusion algorithm unit calculates the correlation between data from different data sources through the Pearson correlation coefficient, using the formula: γ is the Pearson correlation coefficient, where n is the number of data points, and are the means of X and Y respectively. The correlation threshold is set according to the application requirements. When the correlation coefficient is higher than the threshold, it is considered that the data has a strong correlation and further fusion is performed; The weight of each data source in the fusion is determined according to the quality, reliability and importance of the data. For the data x1, x2, ..., x n , whose weights are ω1, ω2, …, ω n (and ), the fused result y is calculated by the formula y=ω1x1+ω2x2+…+ω n x n calculate; When performing data fusion, the state of the system is estimated and predicted based on the dynamic model of the system. Assuming that the state equation of the system is: x k =Ax k-1 +Bμ k-1 +oh k-1 where x k is the system state vector at time k, A is the state transfer matrix, B is the control input matrix, μ k-1 is the control input vector, ω k-1 is the process noise vector, which predicts the state at the current moment through the state estimation at the previous moment and the known system model; At the same time, according to the observation equation, z k =Hx k +v k where z k is the observation vector at time k, H is the observation matrix, v k is the observation noise vector, combining the predicted state and the actual observation value, using the Kalman gain K k To update the state estimate, the Kalman gain is calculated as follows: K k =P k / k-1 H T (HP k / k-1 H T +R) -1 Where P k / k-1 is the prediction covariance matrix, R is the observation noise covariance matrix, and the updated state estimation formula is, x k / k =x k / k-1 +K k (z k -Hx k / k-1 ) Through continuous prediction and observation updates, dynamic data fusion and state estimation are achieved.
8. The twin multi-data processing system of a digital twin model according to claim 1, characterized in that: The data storage and management module includes a distributed storage unit and a data index and metadata management unit, and the distributed storage unit stores data in a dispersed manner on multiple storage nodes; The data index and metadata management unit establishes an effective index to quickly locate required data, while recording information on the source, collection time and processing history of the data.
9. A method for processing multiple data of a twin of a digital twin model according to any one of claims 1 to 8, characterized in that: The specific steps include: Step 1: For physical entities, deploy a variety of sensors in key locations and key processes, collect relevant data from IoT devices and network platforms, and mine and integrate historical data related to physical entities, including past operation records, maintenance data, and fault logs. These historical data provide a basis for analyzing long-term trends and behavior patterns of physical entities. Step 2: Identify and process outliers in the data through statistical analysis methods, delete outliers, replace them with adjacent values, and perform corrections based on model prediction. For data with missing values, use the mean and median to fill in the missing values, extract the features of data from different data sources, and then fuse these features. Then fuse the original data from different data sources, arrange them in a certain order to form a new data set, and finally use weighted average to synthesize the results to obtain the final data fusion result; Step 3: For large-scale multi-source data, use distributed storage technology to store the data in multiple nodes. For structured data, use the relational database MySQL, and for semi-structured and unstructured data, use the NoSQL database.