Data-in-channel prediction method based on digital twin system

By building a data middle platform based on a digital twin system, combining XGBoost-LSTM hybrid model and convolutional neural network, the compatibility and real-time problems of digital twin and data middle platform in heterogeneous system are solved, efficient data management and real-time decision-making are achieved, and equipment prediction accuracy and production optimization are improved.

CN120508586APending Publication Date: 2025-08-19YANGZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510416554.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

The existing digital twin and data middle platform technologies face problems such as poor compatibility of heterogeneous systems, imbalance in computing power and storage, contradictions in model accuracy and real-time performance in actual implementation, resulting in high complexity of data development and difficulty in achieving efficient data management and real-time decision-making.

Method used

A data middle platform based on a digital twin system is built, and XGBoost-LSTM hybrid model is combined with a convolutional neural network to process multivariate heterogeneous data, unified data access through edge computing nodes, combined with data storage and computing layer to realize real-time data processing and prediction.

Benefits of technology

Real-time decision-making with second-level response is realized, improving the accuracy of equipment data prediction, reducing equipment unplanned downtime, optimizing production plans, reducing maintenance costs, supporting multi-protocol data access, reducing heterogeneous system integration costs, and improving data transmission efficiency and credibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120508586A_ABST
    Figure CN120508586A_ABST
Patent Text Reader

Abstract

The invention discloses a data-in-channel prediction method based on a digital twin system, and the method comprises the steps: inputting to-be-predicted data into an XGBoost-LSTM hybrid model, obtaining static features and dynamic features of the to-be-predicted data, splicing the static features and dynamic features, inputting the spliced static features and dynamic features into a full-connection layer of a convolutional neural network, and obtaining a prediction result; the step of completing training of the XGBoost-LSTM hybrid model comprises the steps that the constructed XGBoost-LSTM hybrid model comprises an XGBoost branch and an LSTM model branch, and structured data are input into the XGBoost branch to obtain static features; inputting the original data and the time series data into an LSTM model branch to obtain dynamic characteristics; the static features and the dynamic features are spliced and then serve as input of a full-connection layer in a convolutional neural network, the prediction data serve as output of the full-connection layer in the convolutional neural network, and a trained XGBoost-LSTM hybrid model is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a data center prediction method based on a digital twin system, and belongs to the technical field of digital twins. Background Art

[0002] With the rapid development of concepts such as Industry 4.0 and smart manufacturing, the deep integration of the physical and digital worlds has become a core direction for enterprises' digital transformation. It is predicted that by 2026, more than 50% of industrial enterprises will optimize asset performance management through digital twin technology. Digital twins achieve full lifecycle management and simulation optimization by building high-fidelity virtual mappings of physical entities. The data middle platform, on the other hand, aims to capitalize and service data, providing unified data integration, governance, and application support capabilities. The combination of the two offers significant advantages. Despite the progress made in digital twin and data middle platform technologies, their actual implementation still faces problems such as poor compatibility between heterogeneous systems, imbalance between computing power and storage, and the conflict between model accuracy and real-time performance. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to overcome the defects of the existing technology and provide a data middle-end prediction method based on the digital twin system. According to the data characteristics of the digital twin system, a basic platform for data processing is constructed, which allows the configuration and access of various types of databases, integrates multi-dimensional heterogeneous data, reduces the complexity of data development, and provides convenience and support for the visualization and data management of the digital twin system.

[0004] Preferably, the present invention provides a data center prediction method based on a digital twin system, comprising: Input the data to be predicted into the trained XGBoost-LSTM hybrid model to obtain static features and dynamic features; Static features and dynamic features are concatenated and input into the fully connected layer of the convolutional neural network to predict the device status and device fault conditions. Among them, the trained XGBoost-LSTM hybrid model includes: The pre-acquired sensor time series data, raw data and structured data are processed to obtain an equipment health report; an XGBoost-LSTM hybrid model is constructed, which includes an LSTM model branch and an XGBoost model branch. The structured data is input into the XGBoost model branch to extract the static features of the structured data; the raw data and time series data are input into the LSTM model branch to extract the dynamic features; the static features and dynamic features are concatenated and used as the input of the fully connected layer in the convolutional neural network, and the data including the equipment status and equipment failure is used as the output of the fully connected layer in the convolutional neural network. The mapping relationship between the static features of the data, the dynamic features of the data, the equipment status and the equipment failure is constructed to obtain a trained convolutional neural network.

[0005] Preferably, static features include category associations and nonlinear relationships, and dynamic features include trends and cycles.

[0006] Prioritize processing the pre-acquired time series data, raw data, and structured data to obtain a device health report, including: using edge computing nodes to clean, convert protocols, and compress the pre-acquired time series data, raw data, and structured data before sending them to the message queue; the data set in the message queue is processed through a protocol gateway for unified access of heterogeneous device protocols, and the data set is transmitted to the data storage and management layer.

[0007] Prioritize processing the pre-acquired time series data, raw data, and structured data to obtain a device health report, including: The data storage and management layer includes a hierarchical storage module and a metadata management module, which stores time series data in the open source time series database TDengine in the data storage and management layer. Put the original data into Hadoop distributed storage and use Parquet column storage; Put structured data into the data warehouse and model it according to the star schema; The metadata management module records the source, format, lineage and business meaning of the data sets in the hierarchical storage module.

[0008] Prioritize processing the pre-acquired time series data, raw data, and structured data to obtain a device health report, including: The data processing and computing layer is used to perform real-time calculations on window aggregation, complex event monitoring, offline analysis of data sets, and generation of equipment health reports.

[0009] Prioritize processing of pre-timed data, raw data, and structured data, including: Use the Unity 3D platform to build and render geometric models of physical equipment and products in the workshop; Use digital twin simulation software ANSYS Twin Builder to define physical rules including workshop equipment and products.

[0010] Preferably, the prediction results including equipment status and equipment failure conditions are presented in a visual manner including data charts or 3D models.

[0011] Prioritize using edge computing nodes to clean, convert, and compress pre-acquired time series data, raw data, and structured data before sending them to the message queue, including: Using the edge computing service AWS IoT Greengrass, Kalman filtering is performed on the pre-acquired time series data, raw data, and structured data to denoise them. The protocol gateway is used to uniformly convert the industrial protocol into OPC UA and transmit it to the hierarchical storage module through the message queue.

[0012] Preferably, the present invention provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of any one of the methods described in the first aspect when executing the program.

[0013] Preferably, the present invention provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of any one of the methods described in the first aspect when executed by a processor.

[0014] The beneficial effects achieved by the present invention are: 1. This invention uses digital twins to map the state of physical entities in real time. Combined with an AI prediction model consisting of an XGBoost-LSTM hybrid model and a convolutional neural network, it can support sub-second responses and optimize real-time decision-making. By constructing an XGBoost-LSTM hybrid model, this invention integrates structured "static insights" with the "dynamic perception" of time series, leveraging the complementary advantages of these two models in processing multimodal data. This significantly improves model performance and enhances the accuracy of device data predictions, providing an intelligent solution for complex application scenarios.

[0015] 2. The AI prediction model in this invention can proactively identify equipment failure risks, balance equipment loads, dynamically adjust process parameters, and dynamically optimize production plans, thereby reducing unplanned equipment downtime, extending equipment life, lowering maintenance costs, and reducing energy consumption per unit of product. It also enables precise inventory management, reduces raw material backlogs, and improves capital turnover.

[0016] 3. This invention achieves multi-protocol data access by deploying edge computing nodes, unifying multi-source heterogeneous data, and converting data protocols. This reduces the cost of integrating heterogeneous systems, achieves high-throughput, low-latency data transmission, and reduces the time it takes to upload data to the cloud, ensuring real-time data. It also reduces cloud storage and computing pressure and reduces bandwidth consumption.

[0017] 4. The raw data lake in this invention supports low-cost storage of massive amounts of raw data, the time-series database supports real-time status data of storage devices and sensors, and metadata management makes data lineage clearly visible, breaking down data silos, improving query performance, supporting rapid fault location, and ensuring data credibility and compliance.

[0018] 5. This invention encapsulates AI models as APIs to empower business systems. New products can be designed and verified, and new solutions tested within a digital twin environment. Value-added services can be provided based on device twin data, reducing trial-and-error costs and shortening time to market.

[0019] 6. The low-code tools in this invention enable business personnel to quickly build applications. API standardization can shorten third-party system integration cycles. Data analysis provides personalized services, allowing users to view order production progress in real time, improving transparency.

[0020] 7. This invention opens its API to attract third-party developers and foster an industry application ecosystem. The multi-twin collaboration of "factory + supply chain + market" achieves optimal global resource allocation. The closed-loop data model drives continuous optimization, adapting to dynamic environments such as supply chain fluctuations and shifting market demand. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solution of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0022] Figure 1 This is the data middle platform system architecture diagram based on the digital twin system in the present invention; Figure 2 It is a flow chart of data flow and data processing in the data center of the present invention; Figure 3 This is a flowchart of data processing in the XGBoost-LSTM model in the present invention. DETAILED DESCRIPTION

[0023] See also Figure 1, the present application discloses a data middle platform prediction method based on a digital twin system, wherein the overall architecture of the data middle platform of the digital twin system includes a data acquisition and access layer, a data storage and management layer, a data processing and computing layer, a digital twin modeling layer, and a data service and application layer. Among them, the data acquisition and access layer includes a physical device data part and a transmission module; the data storage and management layer includes a hierarchical storage module and a metadata management module; the data processing and computing layer includes a stream-batch integrated processing engine, an AI model prediction module, and an AI model integration module; the digital twin modeling layer includes 3D model construction technology and physical simulation technology; the data service and application layer includes API service technology, visualization tools, and low-code tools. Among them, the hierarchical storage module includes the original data lake, the time series database, and the data warehouse.

[0024] The data collection and access layer mainly refers to the collection and preprocessing of multi-source heterogeneous data from physical devices, business systems and external environments, and transmission through message queues (Kafka) to ensure that the data enters the middle platform in real time and in full.

[0025] The physical device data mentioned here primarily refers to industrial device data and edge computing nodes. Industrial device data includes raw data generated by sensors, PLCs, CNC machine tools, and other devices. Edge computing nodes are lightweight computing units deployed near the devices. They run filtering and noise reduction algorithms on the devices through AWS IoT Greengrass, only uploading critical data. They also utilize a protocol gateway (Kepware) to convert industrial protocols supported by different devices into the standardized OPC UA protocol. The transmission module mainly refers to the message queue (Kafka) that receives the cleaned data and handles the unified access of heterogeneous device protocols.

[0026] The data storage and management layer is responsible for the hierarchical storage of data, the management and governance of metadata, and ensures that data is accessible, traceable, secure and compliant.

[0027] The function of the hierarchical storage module is to classify and store data in storage systems at different levels according to the data type, access frequency, processing requirements and usage scenarios.

[0028] The raw data lake mainly refers to the Hadoop Distributed File System (Hadoop HDFS), which uses columnar storage (Parquet) to store unprocessed raw data.

[0029] The time series database mentioned here mainly refers to the open source time series database TDengine, which stores frequently accessed real-time data and supports high-speed writing and time window aggregation queries.

[0030] The data warehouse mentioned here primarily refers to Snowflake, which stores cleaned structured data and models it using a star schema. It uses columnar storage and optimized indexes to support complex SQL queries.

[0031] The metadata management module has the function of recording the source, format, lineage and business meaning of the data, ensuring the consistency of data versions at all levels, and ensuring the discoverability, comprehensibility, trustworthiness and reusability of the data.

[0032] The data processing and computing layer is used to perform real-time stream processing, offline batch processing, and AI model calculation on data, converting data into high-value information that can support decision-making and simulation.

[0033] The stream-batch integrated processing engine mainly refers to the same engine supporting the processing of real-time stream data and historical batch data, performing real-time calculation and statistics on dynamic data streams, and issuing timely alarms.

[0034] The AI model prediction module mainly refers to the construction of an XGBoost-LSTM hybrid model, that is, a hybrid model of the XGBoost model and the LSTM fault prediction model, to perform equipment status prediction and equipment fault detection. XGBoost is a decision tree-based machine learning model that excels at processing structured data; LSTM, short for Long Short-Term Memory, is a special recurrent neural network (RNN) that excels at processing raw data and time series data. A hybrid model of the two is constructed, using the XGBoost branch to process structured data and extract static features; and using the LSTM model branch to process raw data, time series data, and extract dynamic features. The outputs of the two are concatenated and input into the fully connected layer of the convolutional neural network for final prediction.

[0035] The AI integration module mainly refers to deploying the trained machine learning model as an API service, predicting the device status in real time through the API, and seamlessly connecting with the data middle platform and digital twin system.

[0036] The digital twin modeling layer mainly refers to the construction of a virtual mapping (digital twin) of a physical entity.

[0037] The 3D model building technology mainly refers to building and rendering the geometric structure and appearance visualization model of the equipment in the Unity 3D platform to support virtual inspection.

[0038] The physical simulation technology mainly refers to importing models into the physical domain digital twin simulation software ANSYS Twin Builder to define the physical rules of objects including physical equipment in the workshop.

[0039] The data service and application layer uses standardized interfaces and visualization tools to provide data and analysis results to ERP and MES systems through APIs to support upper-layer applications.

[0040] The API service technology mentioned above mainly refers to the unified management of API access rights, rate limiting, and monitoring to ensure service stability. Kong is used to limit device status query requests to a maximum of 100 per second.

[0041] The visualization tool mainly refers to a factory monitoring screen developed using Three.js, which intuitively displays data in the form of charts, 3D models, etc.

[0042] The low-code tools mentioned above mainly refer to interactive tools developed using the open source solution AppSmith, which provide a drag-and-drop tool set to quickly build data analysis applications.

[0043] like Figure 1 As shown in the figure, the overall architecture of the digital twin data middle platform has five layers, namely the data collection and access layer, the data storage and management layer, the data processing and computing layer, and the data service and application layer, on this basis, a more detailed internal division is carried out.

[0044] like Figure 1 As shown in the figure, the data acquisition and access layer is responsible for data collection and transmission. The primary source of standardized data is the physical equipment in the workshop, including sensors, PLCs, and CNC machine tool data. Edge computing nodes are deployed near the equipment to perform data cleansing, protocol conversion (to OPC UA), and data compression on the raw data from the physical equipment and structured data from the business system. The cleaned, standardized data is then sent to the message queue (Kafka). The data in the message queue (Kafka) is processed through a protocol gateway (Kepware) to uniformly access heterogeneous device protocols and transmit the data to the data storage and management layer.

[0045] like Figure 1 As shown, the data storage and management layer includes a tiered storage module and a metadata management module. High-frequency sensor data, transaction flows, and other time-series data are stored in the open-source time-series database TDengine, with customizable retention policies. Raw, unprocessed data, such as raw signals, log files, images, and videos, are stored in Hadoop distributed storage, using Parquet columnar storage to optimize analytical performance. Structured data, such as orders and inventory, is stored in the data warehouse and modeled using a star schema to support complex SQL analysis. The metadata management module records the source, format, lineage, and business significance of the data in the tiered storage module, tracing the complete data chain from sensor to API.

[0046] like Figure 1As shown, the data processing and computation layer is responsible for extracting valuable data information. By merging the stream processing engine (Flink) and the batch processing engine (Spark), real-time window aggregation and offline analysis of historical data are simultaneously performed, achieving sub-second response times. The Flink module is responsible for real-time window aggregation and complex event monitoring, such as recording the average temperature of a device every minute and issuing continuous over-temperature alerts. The Spark module is responsible for offline analysis of historical data and generating device health reports from the data. An XGBoost-LSTM hybrid model is constructed. The processed structured data is used to train the XGBoost branch, outputting static features; the processed time series data and the raw data are used to train the LSTM model branch, outputting dynamic features. Finally, the output features of both are concatenated and input into the fully connected layer of the convolutional neural network for final prediction, enabling device status prediction and fault detection. The trained hybrid model is deployed as an API service, connecting the upper and lower layers to implement device status monitoring and fault prediction.

[0047] like Figure 1 As shown in the figure, the digital twin modeling layer is the source of model data. Geometric models of workshop equipment, products, and other components are constructed and rendered using the Unity 3D platform. Once the geometric models are constructed, they are imported into ANSYS Twin Builder to define the physical rules for the virtual workshop, including the physical equipment and products.

[0048] like Figure 1 As shown in Figure 1, the data service and application layers are the "output" of the data middle platform. They expose interfaces for data query and model call through the API gateway, Kong. The Grafana dashboard displays real-time data and analysis results in visual formats such as data charts and 3D models. Low-code tools connect the data middle platform with end users, allowing them to directly call the APIs provided by the data service layer to quickly transform data and business logic into operational applications.

[0049] like Figure 2As shown, physical device data is sent as raw data to edge computing nodes via industrial protocols, and business systems use ETL tools to send structured data to edge computing nodes. Data preprocessing occurs at the edge computing nodes: Kalman filtering for denoising is performed using AWS Greengrass, while the protocol gateway (Kepware) converts the industrial protocol into OPC UA. Data cleansed by the edge computing nodes is transmitted via message queues to the hierarchical storage module. Raw, unprocessed data is stored in the Hadoop distributed system using the Parquet format; time series data is stored in TDengine; and cleaned, structured data is stored in Snowflake and modeled using a star schema. Real-time and historical data requiring analysis and computation are fed into the batch-stream processing engine. The processed data is used to train an XGBoost-LSTM hybrid model. Finally, the data in the hierarchical storage module, the batch-stream processing data, the device status data and fault detection data output by the XGBoost-LSTM hybrid model, and the twin model data are all output to the data center output terminals: the API gateway Kong, visualization tools, and low-code tools, ensuring real-time and transparent data.

[0050] like Figure 3 As shown in the figure, after stream and batch processing, the data is divided into structured data, raw data, and time series data. All of this data is fed into the XGBoost-LSTM hybrid model as a training set to predict device status and equipment failure. The structured data is used to train the XGBoost branch and extract static features of the structured data, including category associations and nonlinear relationships. The raw data and time series data are used to train the LSTM model branch and extract dynamic features of the data, including trends and cycles. These features are concatenated and fed into the fully connected layer of the convolutional neural network to generate the final predictions, including device status and equipment failure. The prediction results are then visualized using visualization tools and low-code tools.

[0051] In an embodiment of the present application, the present invention provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of any of the above methods when executing the program.

[0052] In an embodiment of the present application, the present invention provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of any of the above methods when executed by a processor.

[0053] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.

[0054] Those skilled in the art will readily appreciate other embodiments of the present invention after considering the specification and practicing the invention as disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not invented herein, and the description and examples are to be considered merely as exemplary.

[0055] The above specific implementation methods further illustrate the purpose, technical solutions and beneficial effects of this application in detail. It should be understood that the above are only specific implementation methods of this application and are not intended to limit the scope of protection of this application. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of this application should be included in the scope of protection of this application.

Claims

1. A data center prediction method based on a digital twin system, characterized in that: include: Input the data to be predicted into the trained XGBoost-LSTM hybrid model to obtain static features and dynamic features; Static features and dynamic features are concatenated and input into the fully connected layer of the convolutional neural network to predict the device status and device fault conditions. Among them, the trained XGBoost-LSTM hybrid model includes: The pre-acquired sensor time series data, raw data and structured data are processed to obtain an equipment health report; an XGBoost-LSTM hybrid model is constructed, which includes an LSTM model branch and an XGBoost model branch. The structured data is input into the XGBoost model branch to extract the static features of the structured data; the raw data and time series data are input into the LSTM model branch to extract the dynamic features; the static features and dynamic features are concatenated and used as the input of the fully connected layer in the convolutional neural network, and the data including the equipment status and equipment failure is used as the output of the fully connected layer in the convolutional neural network. The mapping relationship between the static features of the data, the dynamic features of the data, the equipment status and the equipment failure is constructed to obtain a trained convolutional neural network.

2. A data center prediction method based on a digital twin system according to claim 1, It is characterized by: Static features include category associations and nonlinear relationships, and dynamic features include trends and cycles.

3. The data center prediction method based on the digital twin system according to claim 1 is characterized in that: Process the pre-acquired time series data, raw data and structured data to obtain the device health report, including: using edge computing nodes to clean, convert protocols and compress the pre-acquired time series data, raw data and structured data before sending them to the message queue; the data set in the message queue is processed through the protocol gateway to achieve unified access of heterogeneous device protocols, and the data set is transmitted to the data storage and management layer.

4. The data center prediction method based on the digital twin system according to claim 3 is characterized in that: Process the pre-acquired time series data, raw data, and structured data to obtain a device health report, including: The data storage and management layer includes a hierarchical storage module and a metadata management module, which stores time series data in the open source time series database TDengine in the data storage and management layer. Put the original data into Hadoop distributed storage and use Parquet column storage; Put structured data into the data warehouse and model it according to the star schema; The metadata management module records the source, format, lineage and business meaning of the data sets in the hierarchical storage module.

5. The data center prediction method based on the digital twin system according to claim 4 is characterized in that: Process the pre-acquired time series data, raw data, and structured data to obtain a device health report, including: The data processing and computing layer is used to perform real-time calculations on window aggregation, complex event monitoring, offline analysis of data sets, and generation of equipment health reports.

6. The data center prediction method based on the digital twin system according to claim 5 is characterized in that: Processing of pre-time series data, raw data, and structured data, including: Use the Unity 3D platform to build and render geometric models of physical equipment and products in the workshop; Use digital twin simulation software ANSYS Twin Builder to define physical rules including workshop equipment and products.

7. The data center prediction method based on the digital twin system according to claim 1 is characterized in that: The predicted results, including equipment status and equipment failure conditions, are presented in a visual manner including data charts or 3D models.

8. The data center prediction method based on the digital twin system according to claim 1 is characterized in that: Edge computing nodes are used to clean, convert, and compress pre-acquired time series data, raw data, and structured data, and then send them to the message queue, including: Using the edge computing service AWS IoT Greengrass, Kalman filtering is performed on the pre-acquired time series data, raw data, and structured data to denoise them. The protocol gateway is used to uniformly convert the industrial protocol into OPC UA and transmit it to the hierarchical storage module through the message queue.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the method according to any one of claims 1 to 8 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.