Digital twinborn distributed data storage and calculation platform for logistics transfer field
By designing a distributed data storage and real-time computing platform in the logistics transit, combining edge-cloud collaboration and collaborative simulation technology, the problem of single point failure and insufficient real-time computing capabilities of traditional centralized storage systems when processing massive data is solved, and efficient data storage and real-time scheduling are achieved.
Patent Information
- Application Number
- CN202510533966.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-27
AI Technical Summary
The centralized storage system of traditional logistics transit transfers has problems such as single point failure, limited real-time computing power and insufficient collaborative simulation accuracy when processing massive data, which is difficult to meet the real-time access needs of massive heterogeneous data generated by logistics transit transfers equipment and sensors.
A logistics transit digital twin distributed data storage computing platform is designed, using distributed data storage, edge-cloud collaboration, real-time computing and collaborative simulation technology, and the data is balanced distribution and fault-tolerant recovery through consistent hashing algorithms, virtual nodes and Raft protocols, and distributed parallel computing is realized using MapReduce and stream computing frameworks.
It realizes high availability of data and real-time computing capabilities, improves the accuracy of collaborative simulation, meets the real-time data access and scheduling requirements of logistics transit sorting scenarios, and provides a theoretical basis for intelligent scheduling and efficient decision-making.
Smart Images

Figure CN120067219A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of logistics, and particularly to a digital twin distributed data storage and computing platform for a logistics transfer yard. Background Art
[0002] With the rapid development of industrial Internet and Internet of Things, the data sources in traditional logistics transfer yards have become increasingly diverse, including sensor data, video surveillance, manually input data, etc. By using digital twin technology, a real-time digital model of the physical transfer yard can be constructed in a virtual environment, and prediction, monitoring, and scheduling optimization of equipment status, logistics flow, and environmental parameters can be achieved through high-fidelity simulation means. However, traditional centralized storage systems have the following deficiencies in processing massive data: 1. High-concurrency data writing and reading bottlenecks: Centralized systems are prone to single-point failures and are difficult to meet the real-time access requirements of massive heterogeneous data generated by equipment, sensors, etc. in logistics transfer yards.
[0003] 2. Limited real-time computing power: Batch processing systems cannot analyze data immediately and cannot meet real-time warning and scheduling feedback.
[0004] 3. Insufficient collaborative simulation accuracy: There is a lack of efficient coupling between discrete data and simulation models, and synchronous physical and virtual environment feedback cannot be achieved. Summary of the Invention
[0005] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a digital twin distributed data storage and computing platform for a logistics transfer yard, which constructs a scheduling architecture for a logistics sorting digital twin system supported by distributed data storage, edge-cloud collaboration, real-time computing, and collaborative simulation. A data partitioning and real-time status update model is established for the sorting scenario of the logistics transfer yard, and the balanced distribution and fault tolerance recovery of data are realized through the consistent hashing algorithm, virtual nodes, and Raft protocol.
[0006] The purpose of the present invention is achieved through the following technical solutions: A digital twin distributed data storage and computing platform for a logistics transfer yard, comprising: A distributed data storage module, which is used to design a data partitioning strategy to evenly distribute massive heterogeneous data in a distributed environment; achieve high data availability through redundant backup and fault tolerance mechanisms; perform real-time preprocessing and short-term caching on the logistics sorting site at the edge layer side, and then upload it to the cloud asynchronously for long-term persistence by using a cloud distributed file system or object storage; A real-time computing module, which is used to construct a parallel computing architecture to process batch and streaming data; design a global status update model so that each computing node can collaborate on local status and perform global data fusion using a high-performance big data platform to obtain the global status; A co-simulation module is used to build a high-precision digital twin model of the physical transfer yard; it uses continuous-time and discrete-event simulation technologies to achieve multi-node co-simulation, ensuring the consistency between the real scenario and the digital twin model. The cloud generates scheduling decisions and correction amounts based on the global simulation results, and feeds them back to the edge nodes through secure communication. The edge immediately responds to adjust the local data processing and device control parameters, realizing the closed-loop consistency between the real scenario and the digital twin model.
[0007] The beneficial effects of the present invention are as follows: The present invention constructs a scheduling architecture for a logistics sorting digital twin system supported by distributed data storage, edge-cloud collaboration, real-time computing, and co-simulation. It establishes a data partitioning and real-time status update model for the sorting scenario of the logistics transfer yard, and realizes the balanced distribution and fault tolerance recovery of data through the consistent hashing algorithm, virtual nodes, and the Raft protocol. At the same time, real-time data collection, preprocessing, and short-term caching of on-site data are completed at the edge layer, while the cloud is responsible for data aggregation, persistent storage, and global state fusion, and uses the MapReduce and stream computing frameworks to achieve distributed parallel computing. Further combining the digital twin model with the closed-loop feedback control mechanism (using the PID algorithm), it realizes the precise monitoring and dynamic verification of global scheduling decisions and on-site device status, effectively reconciling the balance between data consistency, system throughput, response speed, and scheduling accuracy. Thus, the present invention provides a new theoretical basis and technical path for the intelligent scheduling and efficient decision-making of the logistics sorting system, effectively realizes the dynamic collaborative optimization of key performance indicators, and lays a solid foundation for the subsequent research and application of the intelligent logistics digital twin system. Description of the Drawings
[0008] Figure 1 It is a schematic diagram of the principle of the present invention. Detailed Embodiments
[0009] The technical solutions of the present invention will be further described in detail below with reference to the drawings, but the protection scope of the present invention is not limited to the following.
[0010] As Figure 1 shown, a digital twin distributed data storage and computing platform for a logistics transfer yard includes: 1. A distributed data storage module is used to design a data partitioning strategy to evenly distribute massive heterogeneous data in a distributed environment; achieve high data availability through redundant backup and fault tolerance mechanisms; perform real-time preprocessing and short-term caching on the logistics sorting site at the edge layer side, and then upload it to the cloud through asynchronous replication, and use the cloud distributed file system or object storage to achieve long-term persistence; 1.1 Data Partitioning and Load Balancing: The consistent hashing algorithm is used to evenly distribute the sorting data items of the current logistics transfer yard to each storage node of the digital twin system. Specifically, it includes: assuming that the current digital twin system includes storage nodes, and multiple virtual nodes are configured on each node to improve the balance; defining the hash function , and mapping the sorting data item to the interval ; where represents the set of sorting data items , i = 1, 2,.., I , where I represents the number of sorting data items; assuming that the positions of all virtual nodes are located on the hash ring and the values are between 0 and 1; the virtual node position of the storage node is , where is the number of virtual nodes; Then the main storage node selection rule for the sorting data item is: That is, find all the values of j that satisfy the condition , and take the value of j among them; represents mapping the sorting data item to the interval , and obtaining the hash value; If all do not satisfy , then ; 1.2 Redundant Backup and Fault Tolerance Mechanism: For each data item, outside the main storage node, the system will sequentially select the next storage nodes in the clockwise direction of the main node as the replica storage points. To ensure the consistency of the data status between the main node and the replicas, the system adopts the distributed consensus protocol Raft. Specifically, in the cluster of the current main node and replicas, the Raft protocol will elect a leader, responsible for the scheduling of all write operations; the remaining nodes are follower nodes; when the sorting data write request arrives, the leader first writes the sorting data write request into its own log, and then copies the sorting data write request to other follower nodes; when more than half of the total number of nodes in the cluster confirm receiving the sorting data write request and successfully write it into their own logs, the leader and follower nodes will write and update the node status with the sorting data; Through replicas, when the main node fails, the system can elect a new main node from the remaining replicas, quickly complete the fault recovery, ensure that the data will not be lost, and the business will continue without interruption.
[0011] 1.3 Edge-Cloud Hierarchical Storage Architecture: To meet the dual requirements of low-latency response and big data deep processing in the logistics sorting scenario, this solution adopts an edge-cloud collaborative architecture, including edge-layer storage, cloud storage, data synchronization, and consistency.
[0012] 1.3.1 Edge-layer storage is concentrated on edge servers or intelligent gateways deployed near the logistics sorting site. Its main functions include: Real-time data collection and preprocessing, that is, directly collecting sensor sorting data on-site, including sorting speed, equipment status, video, and RFID information, and using the sliding window algorithm to filter and normalize the original sorting data to obtain: where, is the sliding window size; is the current data processing time, is the offset step within the sliding window, and its value range is , representing the historical time point traced back from the current moment; represents at time the original data value; Short-term caching and local analysis, that is, caching the sorting data within the last few seconds to minutes in local memory or flash memory, and using the local state update model to achieve real-time update: where, represents the state vector of the edge node. The initial value of the vector is usually zero, dynamically reflecting the local evaluation of the edge node in dimensions such as sorting state recognition and operation stability, represents the time step of state update, in seconds, specifically determined by the system sampling frequency or the edge node processing cycle; is the edge-side state mapping function, used to estimate its change rate based on the current state and input data, specifically defined as where, is the weight matrix, is the bias vector, is the ReLU activation function, and the input is the concatenation of . The output result is the estimated value of the state change rate, which is updated to the state vector after multiplying by the time step .
[0013] The training of the weight and bias vectors is as follows: The input feature pair z(t) = , the real state change rate (obtained from historical collected data) ; Definition of loss function: ; Based on the loss function, use Adam to iteratively update the parameters until convergence (i.e., the loss function is less than the set threshold).
[0014] 1.3.2 Cloud storage is responsible for aggregating and persistently storing the data from each edge node, and storing the global state after the global update of the real-time computing module.
[0015] 1.3.3 Data synchronization and consistency. Upload to the cloud through asynchronous replication and message queues to ensure the consistency of the sorting data upload, and use the consistent hashing algorithm and the Raft fault tolerance mechanism mentioned in 1.2 to ensure the synchronization of the data versions between the edge and the cloud.
[0016] 2. Real-time computing module, used to build a parallel computing architecture to process batch and streaming data; design a global state update model so that each computing node can cooperate with local states to use a high-performance big data platform for global data fusion to obtain the global state; The real-time computing module adopts a parallel computing architecture and uses MapReduce and a stream computing framework to process the batch and streaming data uploaded by the edge, which is used to support the distributed computing and global collaborative update of the data: Based on the edge-cloud hierarchical storage architecture, local calculations are first performed dispersedly at each edge node in the Map stage to obtain local states, and then the global state is updated using the local states uploaded by each edge node in the Reduce stage Among them, represents the global state vector, and the initial value of the vector is zero; represents the global input: is the preset weight, represents the original sorting data of the i-th edge node after being filtered and normalized by the sliding window algorithm; is the fusion function, defined as: Among them, is the weight matrix, is the bias vector, is the ReLU activation function. The training of the weight and bias parameters is as follows: Input feature pair = , the target output of the true global state change rate (obtained from historical collected data): Function: Use Adam to iteratively update the parameters based on the loss function until convergence (i.e., the loss function is less than the set threshold).
[0017] 2.2 Global state update: First, the edge nodes update the local state according to the real-time data Then, the cloud fuses the states from each node, corrects the errors, and conducts in-depth data analysis, and finally updates the global digital twin state to achieve seamless collaboration.
[0018] 3. Cooperative simulation module, used to build a high-precision digital twin model of the physical transfer yard; use continuous time and discrete event simulation technologies to achieve multi-node cooperative simulation, ensure the consistency between the real scenario and the digital twin model, the cloud generates scheduling decisions and correction amounts according to the global simulation results, and feeds them back to the edge nodes through secure communication, and the edge immediately responds to adjust the local data processing and device control parameters to achieve the closed-loop consistency between the real scenario and the digital twin model.
[0019] 3.1 Data coordination and synchronization of the high-precision digital twin model of the physical transfer yard: including simulating the dynamic changes of sorting equipment during sorting tasks in continuous time simulation, and using differential equations to describe the equipment state : Where represents the local sorting state information of the sorting equipment. And in discrete event simulation, key events such as scheduling decisions and abnormal events are described, and the state is corrected through event triggering.
[0020] 3.2 Multi-node cooperative simulation: Combine the states of all key scheduling nodes into a global state vector : Wherein is the scheduling and control signal to ensure the coordinated operation of each node.
[0021] 3.3 Cooperative feedback and closed-loop control: To ensure the real-time synchronization of cloud scheduling decisions and on-site control, a closed-loop feedback mechanism is established, including global error calculation and feedback generation and immediate response of edge nodes.
[0022] 3.3.1 Global error calculation and feedback generation: According to the global state updated at each moment , determine the expected state through historical sorting data , and calculate the global error: Based on the error Use the PID algorithm to calculate the feedback correction amount: Among them, is the proportional gain, which directly amplifies the current error; is the integral gain, which is used to eliminate the steady-state error; is the derivative gain, which is used to predict the error change and reduce system oscillation.
[0023] 3.3.2 Edge node instant response: As the part closest to the on-site physical environment, the main task of the edge side is to receive the feedback control information from the cloud and timely adjust the local parameters or device working parameters. Specifically, it includes using the edge node to maintain a persistent connection with the cloud using MQTT to receive feedback messages; and verifying the received messages, using the check bit to verify the message integrity, and confirming the timeliness and version consistency of the messages. After the edge node receives the cloud feedback, update the processing parameters: Among them, is the original sorting data after filtering and normalization by the sliding window algorithm, is the updated parameter. At the same time, the edge side can configure a local buffering mechanism: before receiving new feedback, continue to use the previous parameter, and temporarily store the key data at the same time, and wait for subsequent updates to ensure that the system will not fall into an unstable state due to communication delays.
[0024] The updated parameters are directly applied to the control logic of the edge device. For example, automatically adjust the camera sampling rate, control the running speed of the sorting device, correct the threshold value collected by the sensor, etc.; the edge node also feeds back the updated local state to the cloud for subsequent global fusion and iterative improvement, so as to form a closed-loop control and achieve continuous update of the global and local states of the system and continuous approach to the optimal equilibrium state.
[0025] The overall working principle of the present invention is as follows: It consists of three modules: data storage, real-time computing, and collaborative simulation, forming a tightly linked hierarchical closed-loop structure with each other: 1) The data storage module first collects on-site sorting data at the edge side, performs preprocessing and short-term caching, and uploads the data to the cloud through the consistent hashing and redundant backup mechanisms to achieve efficient partitioning, distributed storage, and version consistency guarantee of multi-source heterogeneous data; 2) The real-time computing module directly calls the structured data and state vectors at the edge and in the cloud, and uses the MapReduce and stream computing frameworks for batch processing and incremental analysis to update the local state at the edge and the global state in the cloud in real time, providing dynamic data support for subsequent simulation and scheduling; 3) The collaborative simulation module runs continuous-time and discrete-event simulation models based on the state information output by the real-time computing module, constructs a digital twin picture of the global logistics sorting scenario, and calculates the global error by comparing with the historical model; 4) Subsequently, a feedback control quantity is generated through the PID algorithm and transmitted to the edge node via the message middleware for real-time adjustment of the device operation parameters and local processing strategies. The updated local state is then fed back to the storage and computing modules to achieve dynamic closed-loop control with edge-cloud collaboration and data-model linkage, and efficient coupling of multiple modules.
[0026] The above is the preferred embodiment of the present invention. It should be understood that the present invention is not limited to the form disclosed herein, and should not be regarded as excluding other embodiments. Instead, it can be used in other combinations, modifications, and environments, and can be changed within the scope of the concept described herein through the above teachings or the technology or knowledge in related fields. Any changes and variations made by those skilled in the art without departing from the spirit and scope of the present invention shall fall within the protection scope of the appended claims of the present invention.
Claims
1. A digital twin distributed data storage and computing platform for logistics transfer sites, characterized by: include: Distributed data storage module, which is used to design data partitioning strategies and add redundant backup and fault-tolerant mechanisms. It performs real-time preprocessing and short-term caching of logistics sorting sites on the edge layer side, and then uploads them to the cloud through asynchronous replication, using the cloud distributed file system or object storage for long-term persistence. The real-time computing module is used to build a parallel computing architecture to process batch and streaming data; the global state update model is designed so that each computing node can coordinate the local state and use the high-performance big data platform to perform global data fusion to obtain the global state; The collaborative simulation module is used to build a high-precision digital twin model of the physical transfer site. It uses continuous time and discrete event simulation technology to achieve multi-node collaborative simulation to ensure the consistency between the real scene and the digital twin model. The cloud generates scheduling decisions and corrections based on the global simulation results, and feeds them back to the edge nodes through secure communication. The edge nodes respond instantly to adjust local data processing and equipment control parameters to achieve closed-loop consistency between the real scene and the digital twin model.
2. According to claim 1, a digital twin distributed data storage and computing platform for logistics transfer sites is characterized by: The distributed data storage module includes a data partition and load balancing submodule, a redundant backup and fault tolerance mechanism submodule, and an edge-cloud layered storage architecture: The data partitioning and load balancing submodule evenly distributes the current logistics transfer site sorting data items to each storage node of the digital twin system through the consistent hashing algorithm: Assume that the current digital twin system includes storage nodes, each node is configured with multiple virtual nodes to improve balance; define the hash function , sorting data items Mapping to interval within; among them, Represents a sorted data item A collection of i =1,2,.., I , where I represents the number of sorted data items; assume that the positions of all virtual nodes are located on the hash ring and the values are between 0 and 1; storage node The virtual node location is ,in is the number of virtual nodes; Then for the sorting data items The primary storage node selection rule is: Find the satisfaction For all the values of j under the condition, take the value of j among them; Indicates that data items will be sorted Mapping to interval , the obtained hash value; If all Not satisfied ,but ; The redundant backup and fault tolerance mechanism submodule, each data item is outside the main storage node, and selects the next one in the clockwise direction of the main node. storage nodes as replica storage points; to ensure that the data status between the primary storage node and the replica storage point is consistent, the distributed consistency protocol Raft is used, which includes: In the cluster of the current primary storage node and replica storage point, the Raft protocol will elect a leader, and the remaining nodes will serve as follower nodes. When a sorting data write request arrives, the leader will first write the sorting data write request to its own log, and then copy the sorting data write request to other follower nodes. When more than half of the nodes in the cluster confirm that they have received the sorting data write request and have successfully written it to their own logs, the leader and follower nodes will write the sorting data and update the node status. pass When the master node fails, a new master storage node is selected from the remaining replica storage points to quickly complete the failure recovery, ensuring that data will not be lost and business continuity is uninterrupted; The edge-cloud layered storage architecture is used to realize edge storage and cloud storage and ensure data synchronization and consistency.
3. According to claim 2, a digital twin distributed data storage and computing platform for logistics transfer sites is characterized by: The edge-cloud layered storage architecture includes an edge layer storage unit, a cloud layer storage unit, and a data synchronization and consistency unit; The edge layer storage unit is concentrated on the edge server or intelligent gateway deployed near the logistics sorting site, and its functions include: Real-time data collection and preprocessing, that is, directly collecting sensor sorting data on site, including sorting speed, equipment status, video, RFID information, and processing the original sorting data The sliding window algorithm is used for filtering and normalization, and the following is obtained: in, is the sliding window size; is the current data processing time, is the number of offset steps in the sliding window, and its value range is , indicating the historical time point looking back to the current moment; Indicates at time The original data value at ; Short-term caching and local analysis, that is, caching the sorting data in the recent period in local memory or flash memory, and using the local state update model to achieve real-time update: in, Represents the state vector of the edge node. The initial value of the vector is zero, which dynamically reflects the local evaluation of the edge node in the dimensions of sorting state recognition and operation stability; is the edge-side state mapping function, which is used to estimate the rate of change based on the current state and input data. It is specifically defined as in, is the weight matrix, is the bias vector, is the ReLU activation function; Indicates the time step of state update in seconds; The cloud storage unit is responsible for aggregating and persistently storing data from each edge node, and storing the global state after the real-time computing module is globally updated; The data synchronization and consistency unit uses a consistent hashing algorithm and a Raft fault-tolerant mechanism to ensure data version synchronization between edge and cloud storage units.
4. According to claim 3, a digital twin distributed data storage and computing platform for logistics transfer sites is characterized by: The real-time computing module adopts a parallel computing architecture, and uses MapReduce and stream computing frameworks to process batch and streaming data uploaded from the edge, to support distributed computing and global collaborative updates of data: Based on the edge-cloud layered storage architecture, each edge node first performs local calculations in a decentralized manner in the Map phase to obtain local states; Then, in the Reduce phase, the local states uploaded by each edge node are used to update the global state, and the updated data is uploaded to the cloud storage unit for storage; in, Represents the global state vector, the initial value of the vector is zero; Represents global input: is the preset weight, Represents the original sorting data of the i-th edge node The result obtained by filtering and normalizing with the sliding window algorithm; is the fusion function, defined as: in, is the weight matrix, is the bias vector, is the ReLU activation function.
5. According to claim 4, a digital twin distributed data storage and computing platform for logistics transfer sites is characterized by: The collaborative simulation module includes: Multiple high-precision physical transfer station digital twin model data collaborative synchronization submodule: including continuous time simulation of the dynamic changes of sorting equipment in sorting tasks, using differential equations to describe the equipment status : in Represent the local sorting status information of the sorting equipment, and describe the key events of scheduling decisions and abnormal events in discrete event simulation, and modify the status through event triggering; Multi-node collaborative simulation submodule combines the states of all key scheduling nodes into a global state vector : in, To schedule and control signals to ensure that all nodes work together; Collaborative feedback and closed-loop control submodule: To ensure real-time synchronization between cloud scheduling decisions and on-site control, a closed-loop feedback mechanism is established, including global error calculation and feedback generation and instant response of edge nodes.
6. A digital twin distributed data storage and computing platform for logistics transfer sites according to claim 5, characterized in that: The collaborative feedback and closed-loop control submodules include: Global error calculation and feedback generation unit: based on the global state updated at each moment , determine the expected state through historical sorting data , and calculate the global error: Based on error The PID algorithm is used to calculate the feedback correction: in, is the proportional gain, which directly amplifies the current error; is the integral gain, used to eliminate steady-state error; is the differential gain, which is used to predict error changes and reduce system oscillation; Edge node instant response unit: As the part closest to the on-site physical environment, the edge side's main task is to receive feedback control information from the cloud and make timely adjustments to local parameters or equipment operating parameters, including: Use the edge node to maintain a persistent connection with the cloud using MQTT to receive feedback messages; verify the received messages, use the check bit to verify the message integrity, and confirm the timeliness and version consistency of the message. After the edge node receives feedback from the cloud, it updates the processing parameters: in, It is the original sorting data after filtering and normalization by the sliding window algorithm. The updated parameters; at the same time, the edge side configures a local buffer mechanism: before receiving new feedback, the previous parameters continue to be used, and key data is temporarily stored and waits for subsequent updates to ensure that the system will not fall into an unstable state due to communication delays; The updated parameters are directly applied to the control logic of the edge device. The edge node also feeds back the updated local status to the cloud for subsequent global fusion and iterative improvement, thus forming a closed-loop control to achieve the system's global and local status constantly updated and approaching the optimal equilibrium state.
Citation Information
Patent Citations
Logistics sorting method based on reinforcement learning and digital twins
CN117114524A
Power distribution network symbiosis simulation system and method based on digital twinning
CN119419954A
Digital twin architecture for satellite-ground collaborative edge network
CN119675752A
Oil and gas production and distribution blockchain systems and methods implementing same
US20250088376A1
Cited By
Distributed asynchronous order processing system and method and storage medium
CN121029770A
Dynamic scene simulation system and method based on digital twinning
CN121093780A
Dynamic scene simulation system and method based on digital twinning
CN121093780B
Real-time monitoring scheduling method and system based on edge cloud collaboration
CN121126449A