Computer-vision-assisted digital twinning auxiliary modeling method for logistics transfer field
By using computer vision-assisted digital twin-assisted modeling method in logistics transit, the traditional monitoring system's shortcomings in logistics transit dynamic monitoring and multi-modal data fusion are solved, and high-precision digital twin modeling and real-time monitoring of logistics transit is realized.
Patent Information
- Application Number
- CN202510559300.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-30
AI Technical Summary
Traditional logistics transit monitoring systems are difficult to fully reflect the operating status, cargo flow and personnel location, the visual data processing capabilities are insufficient, the target detection and tracking are not accurate enough, and the multi-modal data fusion is lacking, resulting in a large deviation in the digital twin model.
Using a digital twin-assisted modeling method based on computer vision assistance, global dynamic monitoring and high-precision digital twin modeling of key areas of logistics transition are achieved through multi-source data acquisition, edge preprocessing and deep learning technologies. Specific steps include noise suppression and quality enhancement of visual data, object detection and feature extraction, image segmentation and scene recognition, non-visual data acquisition and encoding, and multimodal data fusion and digital twin model construction.
It realizes high-precision perception of the status of logistics transit operations, cargo flow and personnel location, improves real-time monitoring capabilities for dynamic changes on the site, reduces deviations from the digital twin model, and provides accurate data support for fault prediction, maintenance warning and intelligent scheduling.
Smart Images

Figure CN120069690A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of logistics, and particularly to a method for digital twin assisted modeling of a logistics transfer yard assisted by computer vision. Background Art
[0002] Currently, in logistics transfer yards (such as sorting lines, warehouse aisles, loading and unloading areas, etc.), due to large logistics volumes and complex operating environments, the on-site dynamic information changes rapidly. Traditional monitoring mostly relies on single-sensor data or low-resolution videos, making it difficult to comprehensively reflect the operating status, goods flow, and personnel positions. Although some technologies have attempted to fuse RFID, environmental sensor, and camera data, there are still the following defects in aspects such as data preprocessing, semantic extraction, and multi-source information fusion: 1. Insufficient visual data processing ability: Existing systems have limited capabilities in video and image noise reduction and image enhancement, and cannot fully utilize the detailed information captured by high-resolution cameras. 2. Inaccurate object detection and tracking: Current object detection and image segmentation algorithms have problems of false detection or missed detection in complex scenarios, making it difficult to accurately obtain goods flow and personnel position information. 3. Lack of multi-modal data fusion: The spatio-temporal dynamic correlation analysis between visual information and RFID and environmental data is not deep enough, resulting in large deviations in the digital twin model and being unable to completely restore the on-site physical structure and dynamic state. Summary of the Invention
[0003] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a digital twin assisted modeling method assisted by computer vision, which realizes global dynamic monitoring and high-precision digital twin modeling of key areas in the logistics transfer yard through multi-source data collection, edge preprocessing, and deep learning techniques.
[0004] The purpose of the present invention is achieved through the following technical solutions: A digital twin assisted modeling method assisted by computer vision includes the following steps: S1. Collect visual data of key areas in the logistics transfer yard, and perform noise suppression and quality enhancement on the collected visual data; S2. Use the trained YOLO network to extract features from the image after noise suppression and quality enhancement to provide visual features for subsequent object detection and multi-modal fusion; S3. Collect and encode non-visual data to obtain non-visual features; S4. On the preprocessed image perform image segmentation, scene recognition, and object tracking using deep learning methods; S5. Perform multi-modal data fusion and digital twin model construction.
[0005] The beneficial effects of the present invention are as follows: The present invention constructs a system for digital twin-assisted modeling for logistics transfer yards. By deploying a variety of non-visual sensors such as high-resolution cameras, RFID, temperature and humidity, and gas, as well as other key monitoring devices in key areas such as sorting lines, warehouse aisles, and loading and unloading areas, real-time synchronous acquisition and preprocessing of on-site video, images, and environmental data are achieved. The system adopts an advanced edge computing platform and uses image enhancement and noise reduction techniques based on generative adversarial networks (GANs) to restore noisy images with high quality, ensuring the input quality of subsequent vision processing modules. On this basis, the pre-trained YOLO network performs object detection and feature extraction on the enhanced images, effectively identifying the boundaries, positions, and category information of each object in the images. At the same time, image segmentation technology is used to achieve pixel-level semantic segmentation to further obtain fine-grained scene information. By vectorizing and deeply fusing the object detection results, image segmentation results, scene recognition information, and non-visual sensor data, the system establishes a comprehensive description vector reflecting the global state of the logistics transfer yard, thus realizing real-time mapping and two-way data synchronization between the on-site physical structure and the digital twin model. This system not only greatly improves the perception ability of on-site operation status, personnel and cargo flow, and environmental changes, but also provides accurate and comprehensive data support for subsequent fault prediction, maintenance warning, and intelligent scheduling, breaking through the deficiencies of traditional periodic maintenance models in real-time monitoring and warning, and providing a new theoretical basis and technical support for the management and maintenance of intelligent logistics equipment. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] Figure 1 The flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0007] The technical solution of the present invention will be further described in detail below with reference to the drawings, but the protection scope of the present invention is not limited to the following.
[0008] As Figure 1 shown, a digital twin-assisted modeling method based on computer vision assistance includes the following steps: S1. Collect visual data of key areas of the logistics transfer yard, and perform noise suppression and quality enhancement on the collected visual data; Visual data collection and preprocessing: Deploy high-resolution cameras in key areas (such as sorting lines, warehouse aisles, loading and unloading areas) of the logistics transfer yard to collect video and image data in real time. Let the original image frame be represented as: where represents the image height, represents the width, is the number of color channels. The edge computing device performs noise suppression and quality enhancement on the image, and the enhanced image is defined as: Among them, represents an image enhancement and noise reduction algorithm based on a generative adversarial network (GAN). The main architecture is as follows: The generator (Generator, G) uses an encoder-decoder (or U-Net structure) to recover a high-quality image from a noisy image; the discriminator (Discriminator, D) is used to distinguish the difference between the generated image and the high-quality real image.
[0009] Image input: The input noisy image and the real high-quality image are respectively denoted as: Generator : The generator takes the noisy image as the input and outputs the enhanced image: Inside the generator, there are generally convolutional layers, downsampling layers (encoder part), corresponding upsampling layers (decoder part), and skip connections to fully utilize local and global features to restore image details.
[0010] Discriminator : The discriminator inputs an image (real image or generator output image) to judge its authenticity, that is, it outputs a probability: Loss function: The global objective of GAN mainly consists of two parts, including adversarial loss and content loss.
[0011] Adversarial loss: It enables the image generated by the generator to "fool" the discriminator. For the generator and the discriminator, there are the following losses respectively: Content loss: To ensure that the generated image is as consistent as possible with the real image in content, an L1 loss (L2 loss can also be used) is introduced: Global generator objective: The objective function of the generator combines adversarial loss and content loss: Among them, is the trade-off coefficient.
[0012] During the actual training process, multiple batches of samples are collected. Each sample contains several sample groups, and each sample group consists of a noisy image and the corresponding real high-quality image; For each batch of samples, a global generator objective is obtained as the loss function, and then the generative adversarial network GAN is optimized by the mini-batch gradient descent method until the generative adversarial network GAN converges and all batches of samples are trained. In the actual processing, the image to be processed is fed into the generator, and the image after noise suppression and quality enhancement is obtained.
[0013] In the embodiment of the present application, the optimization is achieved by Backpropagation + Adam. The training objective of the entire GAN is as follows: ; In actual training, fix G and optimize D (maximize discriminator accuracy), fix D and optimize G (minimize the gap between the generated image and the real image), and alternate until convergence.
[0014] Multiple groups of samples need to be given during training. Each group of samples includes a noisy image and the corresponding real image .
[0015] S2. Use the trained YOLO network to extract features from the image after noise suppression and quality enhancement, and provide high-level semantic information (visual features) for subsequent object detection and multimodal fusion; The step S2 includes: Use the YOLO network to perform object detection and feature extraction on the image processed in step S1, and the output feature map is: Among them, is the number of grids into which the input image is divided; is the number of bounding boxes predicted for each grid; 5 represents the dimension of 5 parameters, and the 5 parameters are: the center coordinates of the bounding box and the width and height as well as the confidence ; is the number of classes; is the network parameter; the information output by the YOLO network also includes whether a target is detected in each bounding box; For the th grid and the th bounding box, the predicted value includes Among them is the class probability vector, satisfying ; is the The confidence, center abscissa, center ordinate, width, and height predicted for each grid and the th bounding box; denotes the c-th element in During the training process of the YOLO network, multiple sets of samples need to be obtained. Each set of samples includes an input image and the expected feature map of the input image. The input image is input into the YOLO network, and the feature map output by the YOLO network is used to calculate the loss function with the expected feature map. Then, based on the loss function, the YOLO network is updated by the method of gradient descent until convergence or all samples are trained, and a trained YOLO network is obtained. Convergence means that the value of the loss function is less than a preset threshold; Among them, the expected feature map also contains the number of grids; the number of bounding boxes predicted for each grid is , and for the th grid and the th bounding box, the parameters included are , respectively representing the confidence, center abscissa, center ordinate, width, height, and class probability vector expected for the th grid and the th bounding box; ; denotes the c-th element in Using the trained YOLO network to process the input image to obtain a visual feature vector; that is, taking the output of the YOLO network as the visual feature vector for subsequent multimodal data fusion: Among them, the loss function of the YOLO network includes three components: coordinate loss, confidence loss, and classification loss; Among them, the coordinate loss is defined as: Among them, denotes the weight of the coordinate loss; if the th grid and the th bounding box detect a target, , otherwise ; At the same time, the width and height are processed by taking the square root: The confidence loss is defined as: Add confidence loss for bounding boxes that do not detect a target: Among them, represents the confidence weight when no target is detected; The classification loss is defined as: Overall loss synthesis: ; S3. Perform non-visual data acquisition and encoding to obtain non-visual features; Perform multi-source sensor data acquisition: Synchronously deploy multiple sensors for data acquisition; Let the time The non-visual data vector is: Among them represents the number of sensors, represents the information collected by the i-th sensor; The sensors include RFID, temperature and humidity, and gas environment sensors; Time synchronization mechanism: Since the visual data has a high frame rate and the non-visual data has a low sampling frequency, the non-visual data is averaged or interpolated within a certain time window to obtain the aligned data: Data normalization and feature encoding: First, perform normalization processing on the data of different sensors, and then map it to a high-dimensional feature space through a fully connected layer encoding to obtain the non-visual feature expression: Among them is the weight matrix, is the bias, is the activation function, represents the encoded feature dimension; Obtain the visual feature and the encoded non-visual feature respectively; Subsequently, fusion will be performed at the feature level to form a unified multi-modal description.
[0016] S4. On the preprocessed image use deep learning methods for image segmentation, scene recognition, and object tracking; S401. Image segmentation and scene recognition: Image segmentation: Use the Mask R-CNN network to perform pixel-level classification on the image and output the segmentation mask: Among them, represents the function of the Mask R-CNN network, Represents the Mask R-CNN network parameters; When training the Mask R-CNN network, a sample set formed by multiple samples needs to be constructed to train the model; each sample contains the image to be segmented and the expected segmentation mask; after training, it is directly used for image segmentation; Scene recognition: Using a scene classification network (such as a CNN network) to map the image to a scene category probability vector : Among them, is the function of the scene classification network, is the parameter of the scene classification network; When training the scene classification network, a sample set formed by multiple samples needs to be constructed to train the model; each sample contains the input image and the expected category probability vector; after training, it is directly used for image mapping; Where is the total number of scene categories, satisfying ; S402. Object tracking: In consecutive frames, calculate the matching cost through the spatial and appearance features of the target in adjacent frames and frame The cost function is defined as Among them, for frame the th appearance feature vector of the detection box This feature vector contains the bounding box parameters of the i th detection box and the feature vector corresponding to the i-th detection box is the Euclidean distance between target positions; is the target appearance feature distance, is the adjustment coefficient; based on the defined cost function, calculate and the matching cost between them, and then construct a cost matrix: Among them, Convert the constructed cost matrix into a bipartite assignment problem using the Hungarian algorithm, and find the optimal one-to-one matching in polynomial time through the following steps: (1) Perform row subtraction on each row: , ensuring that at least one 0 appears in each row; (2) Perform column subtraction for each column: to ensure that at least one 0 appears in each column; (3) Find the minimum number of horizontal / vertical lines to cover all zero elements; if the number of required lines is less than , then let , and execute: for each uncovered element, update the value ; for the elements covered by only one line, keep the value unchanged; for the elements covered by the intersection of two lines, update the value ; repeat until the number of covering lines is equal to .
[0017] (4) Extract one-to-one matches from the covered zero elements, let the corresponding , to minimize . is a binary variable representing the one-to-one matching situation, that is, The algorithm will output a one-to-one pairing scheme, so that each detection box in frame exactly corresponds to a detection box in frame , realizing the target association and continuous tracking of consecutive frames.
[0018] S5. Perform multi-modal data fusion and digital twin model construction.
[0019] (1) Multi-modal feature fusion: Feature-level fusion: Concatenate the visual feature with the non-visual encoded feature at the vector level to form a unified fused feature vector where, " " represents the vector concatenation operation.
[0020] Attention mechanism weighting: Use the attention network to adaptively adjust the contributions of different modal information, by calculating the attention weights: to obtain the final weighted fused feature: where, are the parameters of the attention network.
[0021] Temporal modeling: Since the state of the logistics site has temporal continuity, use the LSTM model to capture the temporal dependence and update the state representation: During the training process of the LSTM model, multiple groups of samples need to be collected for training. Each group of samples is obtained when t takes different values, and the feature of each group of samples is the weighted fusion feature at time t and the logistics site status at time t-1 , and the label is the logistics site status at time t ; After the training is completed, for any time t, only the weighted fusion feature at time t and the logistics site status at time t-1 are input into the LSTM model, and the LSTM model outputs the prediction result of the logistics site status at time t ; (2) Digital twin model construction and dynamic update: Construct the state description vector: Based on object detection, image segmentation, scene recognition, and non-visual data, form a composite description of the logistics transfer field site status: Among them, , is the deep fusion network mapping function, is to vectorize the detection and segmentation results.
[0022] Dynamic state update: Use the exponential smoothing update strategy to achieve smooth transition of state information: Among them, is the smoothing factor. To further improve the prediction accuracy, Kalman filtering is introduced for state correction: Among them, is the state transition matrix, is the control input matrix, is the external control input, is the Kalman gain.
[0023] Based on , , combined with the image segmentation results, scene recognition results in step S402, and the target tracking scheme as an auxiliary, to achieve the construction of a high-precision dynamic digital twin model for the logistics transfer field.
[0024] In the embodiment of the present application, the mapping in the physical-virtual space can consider the coordinate mapping function: To realize the mapping between the on-site physical space and the coordinates of the digital twin model, an affine transformation function is introduced: Among them, is the rotation and scaling matrix, is the translation vector, is the transformation parameter.
[0025] Minimize the mapping error: Use the least squares method to optimize the mapping parameters and construct the objective function: In summary, the present invention performs real-time acquisition and preprocessing of multi-source data: Deploy high-resolution cameras in key areas of the logistics transfer yard and combine with RFID, environmental sensors, etc. to achieve synchronous and real-time acquisition of on-site videos, images and non-visual data, and perform effective noise reduction and image enhancement preprocessing. And perform in-depth semantic extraction of visual data: Use deep learning technology to perform object detection, scene recognition, object tracking and image segmentation on the preprocessed visual data, extract key information such as operation status, goods flow, and personnel location, and ensure high precision and robustness in complex scenarios. Finally, multi-modal data fusion and digital twin modeling: Real-time fuse the results of computer vision processing with other data such as RFID and environmental sensors to build a high-precision digital twin model of the logistics transfer yard, comprehensively restore the on-site physical structure and dynamic operation status, and achieve virtual mapping and real-time feedback.
[0026] The above is the preferred embodiment of the present invention. It should be understood that the present invention is not limited to the form disclosed herein, should not be regarded as excluding other embodiments, but can be used in other combinations, modifications and environments, and can be changed within the scope of the concept described herein through the above teachings or the technology or knowledge in related fields. And the changes and modifications made by those skilled in the art without departing from the spirit and scope of the present invention should all be within the protection scope of the appended claims of the present invention.
Claims
1. A computer vision-assisted digital twin modeling method for logistics transfer sites, characterized in that: The following steps are involved: S1. Collect visual data of key areas in the logistics transfer station, and perform noise suppression and quality enhancement on the collected visual data; S2. Extract features from the noise suppressed and quality enhanced images by training the YOLO network to obtain visual features; S3. collect and encode non-visual data to obtain non-visual features; S4. Image after preprocessing On the other hand, deep learning methods are used for image segmentation, scene recognition and object tracking; S5. Perform multimodal data fusion and build digital twin models.
2. According to claim 1, a computer vision-assisted digital twin-assisted modeling method for logistics transfer sites is characterized by: The key areas include sorting lines, warehouse passages and loading and unloading areas in logistics transfer yards. By deploying high-resolution cameras in key areas, video image data is collected in real time as visual data for the key areas.
3. According to claim 2, a computer vision-assisted digital twin modeling method for logistics transfer sites is characterized by: The step S1 comprises: Assume the original image frame is represented as: in, Indicates the image height, Indicates width, is the number of color channels; the image is subjected to noise suppression and quality enhancement through edge computing devices, and the enhanced image is defined as: in, Represents an image enhancement and denoising algorithm based on generative adversarial networks.
4. According to claim 3, a computer vision-assisted digital twin modeling method for logistics transfer sites is characterized by: The architecture of the image enhancement and denoising algorithm based on the generative adversarial network (GAN) includes: a generator, which is used to restore a high-quality image from a noisy image; a discriminator, which is used to identify the difference between the generator output image and the real image; The model training process of the image enhancement and denoising algorithm based on the generative adversarial network GAN is as follows: A1. Image input: The input noise image and the real image are recorded as: A2. Generator : The generator uses a noisy image As input, output enhanced image: The generator is composed of convolutional layers, downsampling layers, upsampling layers, and skip connections. The convolutional layers and downsampling layers are the encoder part, and the upsampling layers and skip connections are the decoder part. A3. Discriminator : The discriminator inputs a real image or the generator outputs an image to determine its authenticity, that is, the output probability : Represents the discriminator input, which is a real image or the output image of the generator; the discriminator is composed of convolutional layers, pooling layers, and fully connected layers; A4. Loss function: The global objective of GAN consists of two parts, including adversarial loss and content loss, both of which are loss functions designed for a batch of samples; Adversarial loss: There are the following adversarial losses for the generator and discriminator: in, : It is used to optimize the discriminator D, the purpose is to make D recognize the real image, It means to find the discriminant expectation of the real images of the current batch, that is, multiple Corresponding Find the average value; : Used to optimize the generator G, the purpose is to let G "cheat" the discriminator and make the generated image look more realistic. The discriminant expectation of the current batch of real high-quality images is shown as follows: For multiple Corresponding Find the average value; Content loss: To ensure that the generated image is as consistent as possible with the real image in terms of content, L1 loss is introduced. The content loss is recorded as: express The L1 norm of Represents the expected content loss, that is, the expected loss of the current batch of samples Find the average value; Global generator objective: The objective function of the generator combines the adversarial loss and the content loss: in, is the trade-off coefficient; A5. During the actual training process, the training is carried out in the following manner: In the actual training process, multiple batches of samples are collected, each of which contains several sample groups, and each sample group consists of a noise image and a corresponding real high-quality image; For each batch of samples, the global generator objective is obtained as the loss function according to steps A1 to A4, and then the generative adversarial network GAN is optimized by the small batch gradient descent method until the generative adversarial network GAN converges and all batches of samples are trained to obtain a trained generative adversarial network GAN. In the actual processing process, the image to be processed is sent to the generator to obtain an image with noise suppression and quality enhancement.
5. According to claim 1, a computer vision-assisted digital twin-assisted modeling method for logistics transfer sites is characterized by: The step S2 comprises: The YOLO network is used to perform target detection and feature extraction on the image processed in step S1, and the output feature map is: in Represents the processing function of the YOLO network, The number of grids into which the input image is divided; The number of bounding boxes predicted for each grid; 5 represents the dimension of 5 parameters, which are: the center coordinates of the bounding box , Width Height And confidence ; is the number of categories; is the network parameter; the information output by the YOLO network also includes whether the target is detected in each bounding box; For The first bounding boxes, the predicted values include in is a class probability vector satisfying ; For the grid and The confidence, center horizontal coordinate, center vertical coordinate, width and height of the bounding box prediction; express The cth element in ; During the training process of the YOLO network, multiple groups of samples need to be obtained, each group of samples includes an input image and an expected feature map of the input image, the input image is input into the YOLO network, the loss function is calculated by the feature map output by the YOLO network and the expected feature map, and then based on the loss function, the YOLO network is updated by the gradient descent method until convergence or all sample training is completed, and a trained YOLO network is obtained, where convergence means that the loss function value is less than a preset threshold; Among them, the expected feature map also includes, The number of grids; the number of bounding boxes predicted for each grid is , No. The grid A bounding box, including parameters such as , respectively representing the grid and The expected confidence, center horizontal coordinate, center vertical coordinate, width, height and category probability vector of each bounding box; ; express The cth element in ; Use the trained YOLO network to process the input image and obtain the visual feature vector; that is, take the output of the YOLO network as the visual feature vector: Among them, the loss function of the YOLO network includes three components: coordinate loss, confidence loss and classification loss; Among them, the coordinate loss is defined as: in, represents the weight of coordinate loss; if grid and The object is detected in the bounding box. ,otherwise ; At the same time, the width and height are processed by square root: The confidence loss is defined as: Add confidence loss for bounding boxes where no objects are detected: in, Indicates the confidence weight when no target is detected; The classification loss is defined as: Overall loss synthesis: 。 6. According to claim 5, a computer vision-assisted digital twin modeling method for logistics transfer sites is characterized by: The step S3 comprises: Perform multi-source sensor data collection: Synchronously deploy multiple sensors for data collection; set the time The non-visual data vector is: in Indicates the number of sensors, represents the information collected by the i-th sensor; the sensors include RFID, temperature and humidity, and gas environment sensors; Time synchronization mechanism: Since the visual data frame rate is high and the non-visual data sampling frequency is low, the non-visual data is synchronized within a certain time window. Perform averaging or interpolation to obtain aligned data: Data normalization and feature encoding: The data from different sensors are first normalized, and then encoded in a fully connected layer and mapped to a high-dimensional feature space to obtain non-visual feature expressions: in is the weight matrix, is the bias, is the activation function, Represents the feature dimension after encoding; Get visual features respectively and the encoded non-visual features Later, they will be fused at the feature level to form a unified multimodal description.
7. According to claim 1, a computer vision-assisted digital twin-assisted modeling method for logistics transfer sites is characterized by: The step S4 comprises: S401. Image segmentation and scene recognition: Image segmentation: Use the Mask R-CNN network to segment images Perform pixel-level classification and output segmentation mask: in, Represents the function of the Mask R-CNN network, Represents the Mask R-CNN network parameters; Scene recognition: Use scene classification networks to classify images Mapped to scene category probability vector : in, is the function of the scene classification network, Parameters of the scene classification network; in is the total number of scene categories, satisfying ; S402. Object tracking: In consecutive frames, through adjacent frames With frame The spatial and appearance features of the target in the image are used to calculate the matching cost. The cost function is defined as: Among them, for the frame No. The appearance feature vector of the detection box , the feature vector contains i Bounding box parameters for detection boxes The feature vector corresponding to the i-th detection box is the Euclidean distance between target locations; is the target appearance feature distance, is the adjustment coefficient; based on the defined cost function, calculate and The matching cost between them is then used to construct the cost matrix: in, The constructed cost matrix is transformed into a binary assignment problem using the Hungarian algorithm, and the optimal one-to-one matching is obtained in polynomial time through the following steps. The Hungarian algorithm outputs a one-to-one pairing solution as the target tracking solution. Each detection box in the frame corresponds exactly to the A detection box in the image is used to achieve target association and continuous tracking in consecutive frames.
8. The computer vision-assisted digital twin modeling method for logistics transfer sites according to claim 7 is characterized in that: The step S5 comprises: S501. Multimodal feature fusion: Feature-level fusion: visual features With non-visual encoding features Perform vector cascading to form a unified fusion feature vector: in," " indicates a vector concatenation operation; Attention mechanism weighting: Use the attention network to adaptively adjust the contribution of different modal information by calculating the attention weights: Get the final weighted fusion feature: in, is the attention network parameter; Represents element-wise multiplication; Time series modeling: Since the state of the logistics site has time continuity, the LSTM model is used to capture the time series dependency and update the state representation: The logistics site status includes: the real-time location status of the parcel in the logistics transfer yard, the sorting status of the sorting machine, and the location status of the sorting personnel; During the LSTM model training process, multiple groups of samples need to be collected for training. Each group of samples is obtained when t takes different values. The features of each group of samples are the weighted fusion features at time t. and the logistics site status at time t-1 , the label is the logistics site status at time t ; After training is completed, for any time t, it is only necessary to combine the weighted fusion features at time t and the logistics site status at time t-1 Input the LSTM model, and the LSTM model outputs the logistics site status at time t The prediction results; S502. Construction and dynamic update of digital twin model: Constructing a state description vector: Based on target detection, image segmentation, scene recognition, and non-visual data, a composite description of the state of the logistics transfer site is formed: in, , is the deep fusion network mapping function, To vectorize the detection and segmentation results; Dynamic state update: Use exponential smoothing update strategy to achieve smooth transition of state information: in, is the smoothing factor; Introduce Kalman filtering for state correction: in, is the state transfer matrix, is the control input matrix, For external control input, is the Kalman gain; for the state transfer matrix , that is, the real-time location status of the package, the sorting machine sorting status, and the location status change relationship of the sorting personnel from the current moment to the next moment based on the physical model of the sorting system; external control input That is, the input decision information for the logistics transfer station, including the arrangement of parcel sorting slot resources, conveyor belt running speed and sorting personnel scheduling; control input matrix That is, the external control input Under this condition, the operation status of the logistics sorting system is affected; S503.Based on , , combined with the image segmentation results, scene recognition results, and target tracking solutions in step S402 as an aid, a high-precision dynamic digital twin model of the logistics transfer site can be constructed.
Citation Information
Patent Citations
Welding spot detection method based on convolutional neural network
CN113409250A
Vision-based target tracking method and system in intelligent network connection bus scene
CN113724293A
Missile-borne image deblurring method based on generative adversarial network
CN113947589A
Seismic data weak signal enhancement and denoising method based on generative adversarial network
CN118551156A
Parameter sensing and fault processing method of electric energy measuring instrument verification system
CN119577415A
Cited By
Visual interaction system for control and real-time feedback of digital twin cross-platform equipment
CN120277922A