Logistics cargo outer package damage identification method and device based on unmanned trolley
By obtaining cargo images and videos on unmanned trolleys, combined with the neural network of 3D-CNN-LSTM and CNN-Transformer modules, the accuracy of the damage recognition method for the outer packaging of logistics goods is solved, and efficient and accurate identification of cargo damage is achieved.
Patent Information
- Application Number
- CN202510765546.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, the identification accuracy of the outer packaging of logistics goods is low, and it is difficult to fully capture the dynamic change information of the goods in the transportation process, which is greatly affected by light conditions.
Using an unmanned cart-based method, the image of the cargo during loading and unloading and video during transportation is obtained through the image acquisition device, and the neural network architecture of the 3D-CNN-LSTM module, the CNN-Transformer module and the multimodal attention fusion module is established to conduct model training to realize the identification of cargo damage.
It enhances the ability to identify dynamic changes in the damage to the outer packaging of goods, and improves the accuracy of identification, especially in areas that are disturbed by light.
Smart Images

Figure CN120339961A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of logistics management, and particularly relates to a method and device for identifying damage to the outer packaging of logistics goods based on an unmanned vehicle. Background Art
[0002] In the context of the booming development of the modern logistics industry, the efficient and safe transportation of logistics goods is of crucial importance. During the processes of transportation, loading and unloading, etc. of logistics goods, the outer packaging is extremely prone to damage, such as breakage caused by collision, scratches caused by friction, etc.
[0003] Currently, in the technology for identifying damage to the outer packaging of logistics goods, the method of taking a single image for identification is widely used. This method relies on computer vision technology to judge whether there is damage to the outer packaging and the type of damage by performing operations such as feature extraction and classification on a single image.
[0004] However, this existing technology has many defects. A single image can only reflect the state of the outer packaging of the goods at a certain moment and from a certain perspective, and it is difficult to comprehensively capture the dynamic change information of the goods throughout the transportation process. For example, during transportation, due to external forces such as jolts and collisions, the damage may be a gradually developing process, and a single image cannot reflect this dynamic evolution, resulting in misjudgment of the degree of damage. In addition, in the actual logistics environment, the light conditions are variable, and a single image is greatly affected by light, which may lead to poor image quality and further affect the accuracy of damage identification. Therefore, the accuracy of the method for identifying damage to the outer packaging of logistics goods in the existing technology is relatively low. Summary of the Invention
[0005] In view of this, the purpose of the present invention is to provide a method and device for identifying damage to the outer packaging of logistics goods based on an unmanned vehicle, aiming to solve the problem of relatively low accuracy of the method for identifying damage to the outer packaging of logistics goods in the existing technology.
[0006] On the one hand, the present invention proposes a method for identifying damage to the outer packaging of logistics goods based on an unmanned vehicle, and the method includes: When using the unmanned vehicle to transfer goods, obtaining the goods images during the loading and unloading of the goods and the goods video during the transfer process collected by the image acquisition device arranged on the unmanned vehicle; Establishing a neural network architecture composed of a 3D-CNN-LSTM module, a CNN-Transformer module, and a multi-modal attention fusion module, and obtaining a damage identification model obtained by training the model with the neural network architecture; Inputting the goods images and goods video into the damage identification model to obtain the corresponding damage identification information of the goods; Among them, the 3D-CNN-LSTM module is used to extract spatio-temporal features of the cargo video and memorize important long-term damaged features. The CNN-Transformer module is used to capture image features and long-distance dependencies in the cargo image. The multi-modal attention fusion module is used to deeply fuse the spatio-temporal features of the cargo video and the image features of the cargo image.
[0007] Furthermore, in the above-mentioned method for identifying damage to the outer packaging of logistics goods based on an unmanned vehicle, the 3D-CNN-LSTM module includes a 3D convolutional layer, a 3D pooling layer, a Batch Normalization layer, a ReLU activation function, and an LSTM layer connected in sequence; The 3D convolutional layer is used to extract spatio-temporal features of the cargo video, and then the 3D pooling layer is used to downsample the spatio-temporal features to obtain initial features; After the initial features pass through the Batch Normalization layer and the ReLU activation function in sequence, the LSTM layer is used to control the information flow through a gating mechanism to memorize important long-term damaged features; Among them, the calculation formula of the 3D convolutional layer is: ; Among them, is the feature map value of the l th layer at the position ([[]] ), is the convolutional kernel weight, respectively correspond to the position indexes of the convolutional kernel in the three-dimensional space, and are used to perform weighted calculation on the input feature map during the convolution operation, is the bias, , , are the sizes of the convolutional kernel in the width, height, and depth directions respectively, respectively correspond to the position indexes in the three-dimensional space, l represents the number of layers in the network, refers to the th layer feature map at the spatial coordinate position value.
[0008] Furthermore, in the above-mentioned method for identifying damage to the outer packaging of logistics goods based on an unmanned vehicle, the CNN-Transformer module includes a convolutional layer, a pooling layer, a Batch Normalization layer, a ReLU activation function, and a Transformer layer; The convolutional layer is used to extract spatial features of the cargo image, and then the pooling layer is used to downsample the spatial features to obtain initial features; The initial features are flattened after passing through a Batch Normalization layer and a ReLU activation function in sequence and then input into the Transformer layer; The multi-head attention mechanism is used to learn different feature representations, and then a two-layer feed-forward neural network is connected to further transform and fuse the features. Layer normalization layers are added before and after the multi-head attention mechanism and the feed-forward neural network respectively to capture the image features and long-range dependencies in the cargo images.
[0009] Further, in the above method for identifying damaged packages of logistics goods based on an unmanned vehicle, the multi-modal attention fusion module receives the features output by the 3D-CNN-LSTM module and the CNN-Transformer module; The features output by the 3D-CNN-LSTM module and the CNN-Transformer module are respectively subjected to feature alignment and then feature concatenation to obtain concatenated features; Self-attention calculation is respectively performed using the concatenated features, and cross-attention calculation is performed using the features output by the 3D-CNN-LSTM module and the CNN-Transformer module, and then fused cross-attention features are obtained; The fused cross-attention features are subjected to feature transformation through a multi-layer perceptron to achieve deep fusion of the spatio-temporal features of the cargo video and the image features of the cargo images.
[0010] Further, in the above method for identifying damaged packages of logistics goods based on an unmanned vehicle, the calculation formula for performing self-attention calculation using the concatenated features is: ; The calculation formula for performing cross-attention calculation using the features output by the 3D-CNN-LSTM module and the CNN-Transformer module is:
[0011]
[0012] The calculation formula for the fused cross-attention features is: ; Among them, is the concatenated feature, is the feature dimension, is the matrix transpose operation, is the feature after feature alignment of the features output by the CNN-Transformer module, is the feature after feature alignment of the features output by the 3D-CNN-LSTM module, is Guided Feature Attention is Guided Feature Attention and are learnable weight coefficients respectively.
[0013] Furthermore, in the above method for identifying damaged outer packages of logistics goods based on an unmanned vehicle, the step of performing feature transformation on the fused cross-attention features through a multi-layer perceptron includes: ; wherein, is the weight matrix of the first layer of the multi-layer perceptron, is the weight matrix of the second layer of the multi-layer perceptron, is the bias vector of the first layer of the multi-layer perceptron, is the bias vector of the second layer of the multi-layer perceptron, is the fused cross-attention feature.
[0014] Furthermore, in the above method for identifying damaged outer packages of logistics goods based on an unmanned vehicle, the training process of the damage identification model includes: Obtain a training data set composed of historical cargo images, cargo videos, and corresponding damaged identification information, and divide the training data set into a training set, a test set, and a validation set according to a preset ratio; Use the training set to train the model of the neural network architecture, and use the validation set to perform systematic hyperparameter tuning on the trained damage identification model; Use the test set to evaluate the optimized damage identification model, and evaluate the deviation degree between the predicted value and the actual value until the damage identification model tends to be stable.
[0015] Another object of the present invention is to provide a device for identifying damaged outer packages of logistics goods based on an unmanned vehicle, the device includes: An acquisition module, configured to acquire cargo images during loading and unloading and cargo videos during the transfer process of the cargo collected by an image acquisition device arranged on the unmanned vehicle when using the unmanned vehicle for cargo transfer; A building module, configured to build a neural network architecture composed of a 3D-CNN-LSTM module, a CNN-Transformer module, and a multi-modal attention fusion module, and obtain a damage identification model obtained by training the model of the neural network architecture; An identification module, configured to input the cargo images and cargo videos into the damage identification model to obtain corresponding damage identification information of the cargo; Among them, the 3D-CNN-LSTM module is used to extract spatio-temporal features of the cargo video and memorize important long-term damaged features. The CNN-Transformer module is used to capture image features and long-distance dependency relationships in the cargo image. The multi-modal attention fusion module is used to deeply fuse the spatio-temporal features of the cargo video and the image features of the cargo image.
[0016] Another object of the present invention is to provide a readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the above method are implemented.
[0017] Another object of the present invention is to provide an electronic device, including a memory, a processor, and a computer program stored on the memory and running on the processor. When the processor executes the program, the steps of the above method are implemented.
[0018] In the present invention, when using an unmanned vehicle to transfer goods, a cargo image of the goods during loading and unloading and a cargo video of the goods during the transfer process collected by an image acquisition device arranged on the unmanned vehicle are obtained; a neural network architecture composed of a 3D-CNN-LSTM module, a CNN-Transformer module, and a multi-modal attention fusion module is established, and a damaged recognition model obtained by training the model with the neural network architecture is obtained; the cargo image and the cargo video are input into the damaged recognition model to obtain the damaged recognition information of the corresponding cargo; among them, the 3D-CNN-LSTM module is used to extract spatio-temporal features of the cargo video and memorize important long-term damaged features, the CNN-Transformer module is used to capture image features and long-distance dependency relationships in the cargo image, and the multi-modal attention fusion module is used to deeply fuse the spatio-temporal features of the cargo video and the image features of the cargo image. When using the cargo video for damaged recognition, it complements the information with the cargo image, comprehensively captures the dynamic changes of the damage, and enhances the recognition ability of the damaged features in the area affected by light interference. The problem of low recognition accuracy of the existing method for recognizing damaged outer packaging of logistics goods is solved. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 It is a flowchart of the method for recognizing damaged outer packaging of logistics goods based on an unmanned vehicle in the first embodiment of the present invention; Figure 2 It is a structural block diagram of the device for recognizing damaged outer packaging of logistics goods based on an unmanned vehicle in the third embodiment of the present invention.
[0020] The following specific embodiments will further illustrate the present invention in conjunction with the above drawings. SPECIFIC EMBODIMENTS
[0021] To facilitate the understanding of the present invention, the present invention will be described more comprehensively below with reference to the relevant accompanying drawings. Several embodiments of the present invention are shown in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, these embodiments are provided to make the disclosure of the present invention more thorough and comprehensive.
[0022] It should be noted that when an element is referred to as being "fixedly provided on" another element, it can be directly on the other element or there may also be an intermediate element. When an element is considered to be "connected" to another element, it can be directly connected to the other element or there may be an intermediate element at the same time. The terms "vertical", "horizontal", "left", "right" and similar expressions used herein are for illustrative purposes only.
[0023] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs. The terms used herein in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0024] Embodiment 1 Please refer to Figure 1 , which shows a method for identifying damage to the outer packaging of logistics goods based on an unmanned vehicle in the first embodiment of the present invention. The method includes steps S10 to S12.
[0025] Step S10, when using an unmanned vehicle to transfer goods, obtain the goods images during the loading and unloading of the goods collected by the image acquisition device arranged on the unmanned vehicle and the goods video during the transfer process of the goods.
[0026] Among them, the unmanned vehicle is an AGV vehicle, and an image acquisition device is configured on the unmanned vehicle. When using the unmanned vehicle to undertake the task of goods transfer, at this critical link of goods loading and unloading, the image acquisition device will take pictures of the goods to obtain images that can reflect the state of the outer packaging of the goods. These images record the appearance of the goods at the moment of loading and unloading and can be used for subsequent analysis of whether damage to the outer packaging occurs during the loading and unloading process.
[0027] During the transfer process of the goods, the image acquisition device continuously works to collect video information of the goods. This video covers the dynamic change process of the goods during the transfer period, including the state of the goods on the unmanned vehicle, the bumps and collisions that the goods may receive during the driving process of the unmanned vehicle, etc. By obtaining the goods images and goods videos in these two different stages, it can provide a rich data basis for the subsequent identification and analysis of the damage situation of the outer packaging of the goods.
[0028] Step S11, establish a neural network architecture composed of a 3D-CNN-LSTM module, a CNN-Transformer module, and a multi-modal attention fusion module, and obtain a damage recognition model obtained by training the model with the neural network architecture.
[0029] Among them, to achieve accurate recognition of damage to the outer packaging of logistics goods, a neural network architecture is constructed. After the architecture is built, a large amount of labeled data containing goods images and videos can be collected. These data cover normal and damaged samples under different transportation scenarios, goods types, and outer packaging materials. Use these data to train the model of the neural network architecture, continuously adjust the network parameters through the optimizer, minimize the loss function, so that the model can accurately identify the type, degree, and location of damage to the outer packaging of goods, and finally obtain a trained damage recognition model. This model can be used for automatic detection of goods damage in actual logistics scenarios.
[0030] Specifically, the 3D-CNN-LSTM module is responsible for processing the video data of goods during transfer. The 3D-CNN part extracts features through three-dimensional convolution operations, simultaneously capturing the motion information between different frames of the goods and the dynamic changes of damage features in the time and space dimensions. The LSTM part, relying on the gating mechanism, memorizes important damage features over a long time, suppresses irrelevant noise interference, and reflects the evolution of the damage state of the goods in the entire video sequence; More specifically, the 3D-CNN-LSTM module includes a 3D convolutional layer, a 3D pooling layer, a Batch Normalization layer, a ReLU activation function, and an LSTM layer connected in sequence; Use the 3D convolutional layer to extract the spatio-temporal features of the goods video, and then use the 3D pooling layer to downsample the spatio-temporal features to obtain initial features; After the initial features pass through the Batch Normalization layer and the ReLU activation function in sequence, the LSTM layer is used to control the information flow through the gating mechanism to achieve memorizing important damage features over a long time; Among them, the 3D-CNN-LSTM module undertakes the key tasks of deep feature extraction and temporal information processing for goods video data. This module includes a 3D convolutional layer, a 3D pooling layer, a Batch Normalization layer, a ReLU activation function, and an LSTM layer.
[0031] First, the 3D convolutional layer uses 32 3×3×3 convolutional kernels, with a stride of 1 and a padding of 1; in subsequent convolutional layers, the number of convolutional kernels is gradually increased. Specifically, the calculation formula of the 3D convolutional layer is: ; Among them, is thel The feature map value of the layer at position ( ), is the convolutional kernel weight, which respectively correspond to the position indices of the convolutional kernel in the three-dimensional space and are used to perform weighted calculations on the input feature map during the convolution operation, is the bias, , , are respectively the sizes of the convolutional kernel in the width, height, and depth directions, which respectively correspond to the position indices in the three-dimensional space, l represents the number of layers in the network, refers to the value of the feature map of the th layer at the spatial coordinate position The 3D pooling layer performs downsampling operations on the spatio-temporal features output by the 3D convolutional layer. By reducing the size of the feature map, it reduces the amount of data and computational complexity while retaining key spatio-temporal information to obtain initial features. These initial features then sequentially enter the Batch Normalization layer and the ReLU activation function for processing. The Batch Normalization layer normalizes the data, adjusts the data distribution, accelerates the network training process, and improves the stability of the model; the ReLU activation function introduces non-linearity to the data, effectively solving the gradient vanishing problem in the neural network and enhancing the expressive power of the network.
[0032] After the above processing, the data enters the LSTM layer. With its unique gating mechanism (including the input gate, forget gate, and output gate), the LSTM layer precisely controls the flow of information. When processing the temporal data of the goods video, the LSTM layer can selectively retain important features related to the damage of the goods outer packaging over a long time, filter out irrelevant noise information, thereby achieving the memory of important damage features over a long time and providing strong support for accurately identifying the damage state of the goods outer packaging later.
[0033] The CNN-Transformer module mainly targets the goods image. The CNN layer extracts local detailed features of the image, such as textures, edges, etc., through multiple convolutional operations. The Transformer layer uses the multi-head attention mechanism and the feed-forward neural network to model the global semantic information and long-range dependencies of the image, comprehensively capturing the damage features of the goods outer packaging; Exemplarily, the CNN-Transformer module includes a convolutional layer, a pooling layer, a Batch Normalization layer, a ReLU activation function, and a Transformer layer; The convolutional layer is used to extract the spatial features of the cargo image, and then the pooling layer is used to downsample the spatial features to obtain the initial features; The initial features are flattened after passing through the Batch Normalization layer and the ReLU activation function in sequence and then input into the Transformer layer; The multi-head attention mechanism is adopted to learn different feature representations, and then a two-layer feed-forward neural network is connected to further transform and fuse the features. Layer normalization layers are added before and after the multi-head attention mechanism and the feed-forward neural network respectively to capture the image features and long-range dependencies in the cargo image.
[0034] Among them, the CNN-Transformer module is mainly used to deeply mine the key information in the cargo image. First of all, the convolutional layer performs sliding convolutional operations on the cargo image by designing convolutional kernels of different sizes (such as common 3×3 and 5×5 convolutional kernels), and extracts basic spatial features such as texture, edge, and shape in the image, and can keenly capture local features such as fine cracks and scratches on the surface of the cargo outer packaging.
[0035] Subsequently, the pooling layer uses methods such as max pooling or average pooling to downsample the spatial feature map output by the convolutional layer to reduce the data volume and computational complexity, while retaining the main features, and obtains more concise initial features. These initial features are then processed by the Batch Normalization layer and the ReLU activation function in sequence. The Batch Normalization layer normalizes the data, optimizes the data distribution, accelerates the network training process, and improves the model stability; The ReLU activation function introduces non-linearity to the data and enhances the network's ability to express complex features. The processed features are flattened to convert the two-dimensional feature map into a one-dimensional vector as the input of the Transformer layer.
[0036] In the Transformer layer, the multi-head attention mechanism learns multiple representation forms of the cargo image features from different angles through multiple parallel attention heads, and can capture the dependencies between regions far apart in the image. For example, it can identify small damages hidden in large-area stains or associate the damaged features of different parts of the cargo; adding layer normalization layers before and after the multi-head attention mechanism can stabilize the data distribution and ensure the accuracy of the attention calculation. The features processed by the multi-head attention mechanism are then connected to a two-layer feed-forward neural network to further transform and fuse the features, enhance the semantic information of the features, and finally effectively capture the complete image features and long-range dependencies in the cargo image.
[0037] The multi-modal attention fusion module deeply fuses the spatio-temporal features of the cargo video extracted by the 3D-CNN-LSTM module and the image features of the cargo image extracted by the CNN-Transformer module. Through the self-attention and cross-attention mechanisms, it dynamically assigns weights to different modal features and mines the complementary information between them.
[0038] Step S12: Input the cargo image and the cargo video into the damage recognition model to obtain the damage recognition information of the corresponding cargo.
[0039] Among them, in the logistics cargo transfer scenario based on the unmanned vehicle, in the early stage, through the image acquisition device deployed on the unmanned vehicle, the static images of the cargo during the loading and unloading process and the dynamic videos during the transfer process have been obtained. These data contain the key information of the cargo outer packaging. When it is necessary to detect the damage condition of the cargo, the collected cargo image and cargo video are input into the pre-trained damage recognition model together. This model has mastered the internal logic of damage recognition. Therefore, it can accurately identify the damage recognition information of the corresponding cargo, such as the type or location of the damage, etc.
[0040] In summary, the damage recognition method for the outer packaging of logistics cargo based on the unmanned vehicle in the above embodiments of the present invention obtains the cargo image during loading and unloading and the cargo video during the transfer process collected by the image acquisition device deployed on the unmanned vehicle when using the unmanned vehicle for cargo transfer; establishes a neural network architecture composed of a 3D-CNN-LSTM module, a CNN-Transformer module, and a multi-modal attention fusion module, and obtains a damage recognition model obtained by training the model with the neural network architecture; inputs the cargo image and the cargo video into the damage recognition model to obtain the damage recognition information of the corresponding cargo; among them, the 3D-CNN-LSTM module is used to extract spatio-temporal features of the cargo video and memorize long-term important damage features, the CNN-Transformer module is used to capture the image features and long-distance dependencies in the cargo image, and the multi-modal attention fusion module is used to deeply fuse the spatio-temporal features of the cargo video and the image features of the cargo image. When using the cargo video for damage recognition, it complements the information with the cargo image, comprehensively captures the dynamic changes of the damage, and enhances the recognition ability of the damage features in the area affected by light interference. It solves the problem of low recognition accuracy in the existing damage recognition methods for the outer packaging of logistics cargo.
[0041] Embodiment 2 This embodiment also proposes a damage recognition method for the outer packaging of logistics cargo based on the unmanned vehicle. The difference between the damage recognition method for the outer packaging of logistics cargo based on the unmanned vehicle in this embodiment and the damage recognition method for the outer packaging of logistics cargo based on the unmanned vehicle in Embodiment 1 is as follows: The multimodal attention fusion module receives the features output by the 3D-CNN-LSTM module and the CNN-Transformer module; The features output by the 3D-CNN-LSTM module and the CNN-Transformer module are aligned and then concatenated to obtain concatenated features; The concatenated features are used for self-attention calculation, and the features output by the 3D-CNN-LSTM module and the CNN-Transformer module are used for cross-attention calculation, and then the fused cross-attention features are obtained; The fused cross-attention features are transformed through a multi-layer perceptron to achieve a deep fusion of the spatiotemporal features of the cargo video and the image features of the cargo image.
[0042] First, the multimodal attention fusion module is the key hub connecting the 3D-CNN-LSTM module and the CNN-Transformer module, aiming to integrate the complementary information of cargo videos and images to improve the accuracy of damage recognition. The module first receives the spatiotemporal features of cargo videos output by the 3D-CNN-LSTM module and the cargo image features output by the CNN-Transformer module. Since the dimensions, scales, and semantic representations of the features extracted by the two modules may be different, it is necessary to first align the two types of features, unify the feature dimensions and scales through linear transformation, interpolation, or pooling operations, and make them comparable in the same semantic space; The aligned features are spliced to form spliced features that contain spatiotemporal dynamic information and static image details. Then, the modules use the spliced features to perform self-attention calculations, explore the correlation between the various parts of the spliced features, and highlight the key damaged features; At the same time, cross-attention calculation is performed based on the original features output by the 3D-CNN-LSTM module and the CNN-Transformer module. That is, the spatiotemporal characteristics of the video are used to guide the allocation of attention to image features, and the image features are used to enhance the focus on video features, capture the potential connection between video and image, and fuse the results of self-attention and cross-attention to obtain fused cross-attention features.
[0043] Finally, a multi-layer perceptron (MLP) is used to perform nonlinear transformation on the fused cross-attention features, adjust the feature distribution and enhance the feature expression ability, further strengthen the interaction between different modal features, and ultimately achieve a deep fusion of the spatiotemporal features of the cargo video and the cargo image features, outputting a feature vector containing rich multimodal information, providing a comprehensive and high-value feature representation for the subsequent accurate identification of the damaged state of the cargo outer packaging.
[0044] Exemplarily, the calculation formula for self-attention calculation using splicing features is: ; The calculation formula for cross-attention calculation using the features output by the 3D-CNN-LSTM module and the CNN-Transformer module is:
[0045]
[0046] The calculation formula for fusing cross-attention features is: ; Among them, is the splicing feature, is the feature dimension, is the matrix transpose operation, is the feature after feature alignment of the features output by the CNN-Transformer module, is the feature after feature alignment of the features output by the 3D-CNN-LSTM module, is the feature attention guided by, is the feature attention guided by, , are respectively learnable weight coefficients.
[0047] The step of performing feature transformation on the fused cross-attention features through a multi-layer perceptron includes: ; Among them, is the weight matrix of the first layer of the multi-layer perceptron, is the weight matrix of the second layer of the multi-layer perceptron, is the bias vector of the first layer of the multi-layer perceptron, is the bias vector of the second layer of the multi-layer perceptron, is the fused cross-attention feature In addition, in some optional embodiments of the present invention, the training process of the damage recognition model includes: Obtain a training data set composed of historical cargo images, cargo videos and corresponding damage recognition information, and divide the training data set into a training set, a test set and a validation set according to a preset ratio; Use the training set to train the neural network architecture, and use the validation set to perform systematic hyperparameter tuning on the trained damage recognition model; Use the test set to evaluate the optimized damage recognition model, and evaluate the deviation degree between the predicted value and the actual value until the damage recognition model tends to be stable.
[0048] First, collect a large amount of historical cargo images and cargo video data. These data cover the status records of the cargo under different transportation scenarios, different time stages, and different environmental conditions, and correspond them one by one with the corresponding accurate damage recognition information (such as whether there is damage, the type of damage, the location and degree of damage, etc.) to jointly form a training data set.
[0049] In order to scientifically and reasonably train, optimize, and evaluate the model, divide the training data set into three subsets according to a preset ratio (such as the common 7:1:2, that is, 70% training set, 10% test set, 20% validation set): The training set is used to train the neural network architecture composed of 3D-CNN-LSTM module, CNN-Transformer module, and multi-modal attention fusion module. During the training process, the model continuously adjusts the network parameters to learn the mapping relationship between the cargo image and video data and the damage recognition information. The validation set plays an important role in supervision and optimization during the training process. Whenever a training stage (such as an epoch) is completed, use the validation set to test the trained damage recognition model. By observing the performance of the model on the validation set, systematically adjust the hyperparameters (such as learning rate, number of convolutional kernels, number of LSTM hidden units, etc.) to find the parameter combination that makes the model performance optimal and prevent the model from overfitting or underfitting; after the hyperparameter tuning is completed, use the test set to conduct a final evaluation of the optimized damage recognition model. By comparing the predicted damage recognition results (predicted values) of the model with the actual damage recognition information (actual values), calculate the deviation degree between the two (such as evaluation metrics such as accuracy, recall rate, mean square error, etc.), and judge whether the model meets the expected effect according to the evaluation results. Until the performance of the damage recognition model on the test set tends to be stable and can accurately and reliably identify the damage situation of the cargo.
[0050] In summary, in the above embodiments of the present invention, the method for identifying damage to the outer packaging of logistics goods based on an unmanned vehicle obtains the goods images during the loading and unloading of the goods and the goods videos during the transfer process collected by the image acquisition device arranged on the unmanned vehicle when using the unmanned vehicle for goods transfer; establishes a neural network architecture composed of a 3D-CNN-LSTM module, a CNN-Transformer module, and a multi-modal attention fusion module, and obtains a damage identification model obtained by training the model with the neural network architecture; inputs the goods images and goods videos into the damage identification model to obtain the corresponding damage identification information of the goods; among them, the 3D-CNN-LSTM module is used to extract spatio-temporal features of the goods videos and memorize long-term important damage features, the CNN-Transformer module is used to capture the image features and long-distance dependence relationships in the goods images, and the multi-modal attention fusion module is used to deeply fuse the spatio-temporal features of the goods videos and the image features of the goods images. When using the goods videos for damage identification, it complements the information with the goods images, comprehensively captures the dynamic changes of the damage, and enhances the ability to identify damage features in areas disturbed by light. It solves the problem of low identification accuracy in the existing methods for identifying damage to the outer packaging of logistics goods.
[0051] Embodiment III Please refer to Figure 2 , which shows the device for identifying damage to the outer packaging of logistics goods based on an unmanned vehicle proposed in the third embodiment of the present invention. The device includes: An acquisition module 100, configured to obtain the goods images during the loading and unloading of the goods and the goods videos during the transfer process collected by the image acquisition device arranged on the unmanned vehicle when using the unmanned vehicle for goods transfer; A building module 200, configured to establish a neural network architecture composed of a 3D-CNN-LSTM module, a CNN-Transformer module, and a multi-modal attention fusion module, and obtain a damage identification model obtained by training the model with the neural network architecture; An identification module 300, configured to input the goods images and goods videos into the damage identification model to obtain the corresponding damage identification information of the goods; Among them, the 3D-CNN-LSTM module is used to extract spatio-temporal features of the goods videos and memorize long-term important damage features, the CNN-Transformer module is used to capture the image features and long-distance dependence relationships in the goods images, and the multi-modal attention fusion module is used to deeply fuse the spatio-temporal features of the goods videos and the image features of the goods images.
[0052] The functions or operation steps realized when the above modules are executed are substantially the same as those in the above method embodiments, and will not be described in detail here.
[0053] Example 4 On the other hand, the present invention also provides a readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the method described in any one of the above-mentioned Embodiment 1 to Embodiment 2 are implemented.
[0054] Example 5 On the other hand, the present invention also provides an electronic device, which includes a memory, a processor, and a computer program stored on the memory and running on the processor. When the processor executes the program, the steps of the method described in any one of the above-mentioned Embodiment 1 to Embodiment 2 are implemented.
[0055] The technical features of each of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.
[0056] Those skilled in the art can understand that the logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable storage medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable storage medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.
[0057] More specific examples (non-exhaustive list) of computer-readable storage media include the following: an electrical connection portion (electronic device) having one or more wirings, a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable storage medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, then editing, interpreting, or processing it in other suitable ways as necessary, and then storing it in a computer memory.
[0058] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logic functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like.
[0059] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0060] The above-described embodiments merely represent several implementation manners of the present invention, and the description thereof is relatively specific and detailed, but should not be construed as a limitation on the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the appended claims.
Claims
1. A method for identifying damage to the outer packaging of logistics goods based on an unmanned vehicle, characterized in that, The method includes: When using an unmanned vehicle to transfer goods, obtaining the goods images during the loading and unloading of the goods and the goods videos during the transfer process collected by an image acquisition device arranged on the unmanned vehicle; Establishing a neural network architecture composed of a 3D-CNN-LSTM module, a CNN-Transformer module, and a multi-modal attention fusion module, and obtaining a damage recognition model obtained by training the model with the neural network architecture; Inputting the goods images and goods videos into the damage recognition model to obtain the damage recognition information of the corresponding goods; Among them, the 3D-CNN-LSTM module is used to extract spatio-temporal features of the goods videos and memorize long-term important damage features, the CNN-Transformer module is used to capture the image features and long-distance dependence relationships in the goods images, and the multi-modal attention fusion module is used to deeply fuse the spatio-temporal features of the goods videos and the image features of the goods images.
2. The method for identifying damage to the outer packaging of logistics goods based on an unmanned vehicle according to claim 1, characterized in that, The 3D-CNN-LSTM module includes a 3D convolutional layer, a 3D pooling layer, a Batch Normalization layer, a ReLU activation function, and an LSTM layer connected in sequence; Using the 3D convolutional layer to extract the spatio-temporal features of the goods videos, and then using the 3D pooling layer to downsample the spatio-temporal features to obtain initial features; After the initial features pass through the Batch Normalization layer and the ReLU activation function in sequence, the LSTM layer is used to control the information flow through a gating mechanism to realize memorizing long-term important damage features; Among them, the calculation formula of the 3D convolutional layer is: ; Among them, is the feature map value of the l th layer at the position ( ), is the convolutional kernel weight, corresponding to the position index of the convolutional kernel in the three-dimensional space respectively, and is used to perform weighted calculation on the input feature map during the convolution operation, is the bias, , , are the sizes of the convolutional kernel in the width, height, and depth directions respectively, corresponding to the position indices in the three-dimensional space respectively, l represents the number of layers in the network, refers to the value of the feature map of the th layer at the spatial coordinate position.
3. The method for identifying damage of the outer packaging of logistics goods based on an unmanned vehicle according to claim 2, wherein The CNN-Transformer module includes a convolutional layer, a pooling layer, a Batch Normalization layer, a ReLU activation function, and a Transformer layer; Using the convolutional layer to extract the spatial features of the goods images, and then using the pooling layer to downsample the spatial features to obtain initial features; After the initial features pass through the Batch Normalization layer and the ReLU activation function in sequence, they are flattened and input into the Transformer layer; Adopting a multi-head attention mechanism to learn different feature representations, and then connecting a two-layer feed-forward neural network to further transform and fuse the features, and adding layer normalization layers before and after the multi-head attention mechanism and the feed-forward neural network respectively to capture the image features and long-distance dependence relationships in the goods images.
4. The method for identifying damage to the outer packaging of logistics goods based on an unmanned vehicle according to claim 3, wherein The multi-modal attention fusion module receives the features output by the 3D-CNN-LSTM module and the CNN-Transformer module; Performing feature alignment on the features output by the 3D-CNN-LSTM module and the CNN-Transformer module respectively, and then performing feature splicing to obtain spliced features; Performing self-attention calculation on the spliced features respectively, and performing cross-attention calculation on the features output by the 3D-CNN-LSTM module and the CNN-Transformer module, and then obtaining fused cross-attention features; The fused cross-attention features are subjected to feature transformation through a multi-layer perceptron to achieve the deep fusion of the spatio-temporal features of the cargo video and the image features of the cargo image.
5. The method for identifying damage to the outer packaging of logistics goods based on an unmanned vehicle according to claim 4, characterized in that The calculation formula for self-attention calculation using the concatenated features is: ; The calculation formula for cross-attention calculation using the features output by the 3D-CNN-LSTM module and the CNN-Transformer module is: The calculation formula for the fused cross-attention features is: ; Among them, is the splicing feature, is the feature dimension, is the matrix transpose operation, is the feature after feature alignment of the features output by the CNN-Transformer module, is the feature after feature alignment of the features output by the 3D-CNN-LSTM module, is the feature attention guided by is the feature attention guided by and are learnable weight coefficients respectively.
6. The method for identifying damage to the outer packaging of logistics goods based on an unmanned vehicle according to claim 5, characterized in that, The steps of performing feature transformation on the fused cross-attention features through a multi-layer perceptron include: ; Among them, is the weight matrix of the first layer of the multi-layer perceptron, is the weight matrix of the second layer of the multi-layer perceptron, is the bias vector of the first layer of the multi-layer perceptron, is the bias vector of the second layer of the multi-layer perceptron, is the fused cross-attention feature.
7. The method for identifying damage to the outer packaging of logistics goods based on an unmanned vehicle according to claim 1, characterized in that, The training process of the damage recognition model includes: Obtain a training data set composed of historical cargo images, cargo videos, and corresponding damage recognition information, and divide the training data set into a training set, a test set, and a validation set according to a preset ratio; Use the training set to train the neural network architecture, and use the validation set to perform systematic hyperparameter tuning on the trained damage recognition model; Use the test set to evaluate the optimized damage recognition model, and evaluate the deviation degree between the predicted value and the actual value until the damage recognition model tends to be stable.
8. A damaged recognition device for the outer packaging of logistics goods based on an unmanned vehicle, characterized in that, The device includes: An acquisition module, configured to acquire the cargo image during loading and unloading of the cargo and the cargo video during the transfer process of the cargo collected by an image acquisition device arranged on the unmanned vehicle when using the unmanned vehicle for cargo transfer; A building module, configured to build a neural network architecture composed of a 3D-CNN-LSTM module, a CNN-Transformer module, and a multi-modal attention fusion module, and obtain a damage recognition model obtained by training the neural network architecture; An identification module, configured to input the cargo image and the cargo video into the damage recognition model to obtain the corresponding damage recognition information of the cargo; Among them, the 3D-CNN-LSTM module is used to extract spatio-temporal features of the cargo video and memorize long-term important damage features, the CNN-Transformer module is used to capture the image features and long-distance dependencies in the cargo image, and the multi-modal attention fusion module is used to deeply fuse the spatio-temporal features of the cargo video and the image features of the cargo image.
9. A readable storage medium, on which a computer program is stored, characterized in that, When the program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.
10. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and running on the processor. When the processor executes the program, it implements the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Driver behavior identification method based on space-time characteristics
CN110119709A
Container damage identification method, container damage identification device, container damage identification equipment and readable access medium
CN112819793A
Multi-modal lip language recognition method and device based on three-dimensional convolution and visual Transformer, and medium
CN118823881A
Container damage detection method
CN119359657A
Method, device, equipment and system for detecting the operating status of an unmanned vehicle
CN119760575A