Method and system for monitoring stacking safety of goods in goods yard based on computer vision
By applying a computer vision-based cargo stacking safety monitoring method in the cargo yard, semantic features and three-dimensional attitude parameters of the cargo are extracted, and spatial relationships and stress conditions are analyzed, the problem of difficult to monitor the safety and stability of cargo stacking is solved, and high-precision and real-time intelligent monitoring and early warning are achieved.
Patent Information
- Application Number
- CN202510115005.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-24
AI Technical Summary
In the freight yard environment, the safety and stability of cargo stacking are difficult to effectively monitor. Traditional manual inspections are low in efficiency, limited frequency and are susceptible to subjective factors. The existing computer vision-based technology faces technical challenges such as instable lighting, diverse types of goods, complex posture changes and occlusion interference.
The safety monitoring method of cargo stacking in the cargo yard based on computer vision is adopted. By obtaining cargo image data, semantic features are extracted, three-dimensional attitude parameters are calculated, regional weights are adjusted, attitude information of blocked goods is obtained, and the spatial relationship and stress conditions are analyzed between goods are predicted to prevent dangerous situations.
It improves the accuracy, robustness and real-timeness of cargo posture recognition in complex freight yard environments, realizes intelligent monitoring and early warning of cargo stacking, and enhances the safety and stability of cargo stacking.
Smart Images

Figure CN120014555A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of information monitoring, and in particular relates to a computer vision-based method and system for monitoring the safety of cargo stacking in a cargo yard. Background Art
[0002] When cargo is stacked in a freight yard, its posture directly impacts its safety and stability. Traditionally, cargo stacking safety monitoring relies primarily on manual inspections, which suffer from low efficiency, limited frequency, and susceptibility to subjective factors. Computer vision-based cargo posture recognition and monitoring technology offers a new approach to addressing these issues, but practical applications still face numerous technical challenges.
[0003] First, the freight yard environment is complex and ever-changing, with unstable lighting conditions and a wide variety of goods. The appearance, size, material, and packaging of each item vary significantly, making image acquisition and processing difficult. Second, the posture of goods undergoes diverse changes during stacking, including translation, rotation, tilt, and deformation, requiring visual algorithms to accurately capture and quantify these posture changes. Furthermore, goods are often stacked densely, with occlusion and interference between different items, further complicating posture recognition. Furthermore, real-time monitoring of cargo stacking safety requires visual algorithms to quickly process image data and generate judgment results, placing high demands on the algorithm's timeliness. Finally, in practical applications, visual algorithms must be integrated with other sensor data and integrated with freight yard management systems to build a complete cargo stacking safety monitoring system. This poses challenges to the system's robustness, scalability, and ease of use. Summary of the Invention
[0004] To solve the above technical problems, the present invention proposes a computer vision-based cargo stacking safety monitoring method and system in a freight yard, which can improve the accuracy, robustness and real-time performance of cargo posture recognition in complex freight yard environments, and realize intelligent monitoring and early warning of cargo stacking.
[0005] The present invention provides a method for monitoring the safety of cargo stacking in a cargo yard based on computer vision, comprising:
[0006] Obtain cargo image data;
[0007] Acquiring semantic features of the goods based on the goods image data;
[0008] Based on the semantic features, obtaining three-dimensional posture parameters of the goods;
[0009] Based on the three-dimensional posture parameters, the weights of different cargo areas are adjusted to obtain posture information of the obscured cargo;
[0010] Based on the posture information of the obscured goods, the spatial relationship and force conditions between the goods are analyzed to predict dangerous situations.
[0011] Optionally, obtaining semantic features of the goods based on the goods image data includes:
[0012] Preprocessing the cargo image data to obtain high-quality cargo images;
[0013] The semantic features of the goods are obtained according to the high-quality goods image.
[0014] Optionally, preprocessing the cargo image data to obtain high-quality cargo images includes:
[0015] Dynamically adjusting the contrast, brightness, and saturation of the cargo image data using an adaptive algorithm to obtain an adjusted cargo image;
[0016] The adjusted cargo image is evaluated using a convolutional neural network model to determine whether the image quality meets the standards, and the high-quality cargo image is obtained based on the determination result.
[0017] Optionally, obtaining semantic features of the goods based on the high-quality goods image includes:
[0018] Extracting color features, texture features, shape features, and size features of the high-quality cargo image to obtain a multidimensional feature vector;
[0019] A convolutional neural network model is used to perform deep learning and extraction on the multi-dimensional feature vector to obtain the semantic features of the goods.
[0020] Optionally, obtaining three-dimensional posture parameters of the cargo based on the semantic features includes:
[0021] Building a three-dimensional model of the cargo based on the semantic features and pre-defining key points on the cargo, wherein the key points include corner points and edge points;
[0022] Based on predefined key points, detect key points in the cargo image and obtain the pixel coordinates of the key points;
[0023] Convert the pixel coordinates into three-dimensional coordinates, calculate the spatial distance and angular relationship between the key points based on the three-dimensional coordinates, and obtain a spatial relationship matrix between the key points;
[0024] Based on the spatial relationship matrix, the three-dimensional posture parameters of the cargo are estimated.
[0025] Optionally, adjusting the weights of different cargo areas based on the three-dimensional posture parameters to obtain posture information of the obscured cargo includes:
[0026] According to the three-dimensional posture parameters, the attention module adaptively adjusts different acquisition areas in the image to obtain the target cargo area;
[0027] Construct a spatial topological relationship diagram between goods based on the target cargo area;
[0028] A convolutional neural network model is used to extract features from the weighted cargo images to obtain visual feature representations of the cargo.
[0029] The spatial topological relationship graph and visual features are integrated, and the posture information of the occluded goods is inferred through a graph convolutional network.
[0030] Optionally, based on the posture information of the obstructed goods and in combination with the three-dimensional posture parameters of the goods, the spatial relationship and force conditions between the goods are analyzed, and the dangerous situation is predicted including:
[0031] Based on the posture information of the obscured goods, combined with the three-dimensional posture parameters of the goods and the spatial relationship between the goods, a physical simulation engine is used to simulate the force conditions of the goods and evaluate the stability of the goods stacking.
[0032] The present invention also provides a computer vision-based cargo stacking safety monitoring system in a cargo yard, comprising: an image acquisition module, a feature extraction module, a posture estimation module, and a three-dimensional analysis module:
[0033] The image acquisition module is used to acquire images of goods;
[0034] The feature extraction module is used to extract semantic features of the cargo image;
[0035] The posture estimation module is used to estimate the three-dimensional posture of the cargo;
[0036] The three-dimensional analysis module is used to analyze the spatial relationship and stress conditions between goods and predict dangerous situations.
[0037] Compared with the prior art, the present invention has the following advantages and technical effects:
[0038] This invention uses an adaptive image enhancement algorithm to preprocess cargo images to improve image quality. It constructs a multi-level cargo feature extraction model that comprehensively utilizes features such as color, texture, and shape to extract deep semantic features of cargo using a convolutional neural network. It introduces an attention mechanism and contextual information to highlight target cargo, suppress background interference, and infer the posture of obscured cargo. It uses a lightweight posture recognition model to improve model inference speed. It also combines deep information about the cargo yard environment to construct a three-dimensional cargo stacking scene model, analyze the spatial relationships and stress conditions of the cargo, assess stacking stability and safety, and predict possible dangerous situations such as tipping and collapse. This invention improves the accuracy, robustness, and real-time performance of cargo posture recognition in complex cargo yard environments, enabling intelligent monitoring and early warning of cargo stacking. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] The accompanying drawings, which constitute part of the present invention, are provided to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are provided to explain the present invention and do not constitute an undue limitation of the present invention. In the accompanying drawings:
[0040] Figure 1 This is a flow chart of a method for monitoring cargo stacking safety in a freight yard based on computer vision according to an embodiment of the present invention;
[0041] Figure 2 is a data preprocessing flow chart of an embodiment of the present invention;
[0042] Figure 3 This is a flow chart of obtaining semantic features of goods according to an embodiment of the present invention;
[0043] Figure 4 This is a flow chart of obtaining three-dimensional posture parameters of goods according to an embodiment of the present invention;
[0044] Figure 5 This is a flow chart of obtaining posture information of obscured goods according to an embodiment of the present invention;
[0045] Figure 6 is a structural diagram of an attention module according to an embodiment of the present invention;
[0046] Figure 7 This is a structural diagram of a computer vision-based cargo stacking safety monitoring system in a cargo yard according to an embodiment of the present invention. DETAILED DESCRIPTION
[0047] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments of the present invention can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0048] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0049] The following is an explanation of the professional terms involved in the embodiment:
[0050] The new convolutional neural network model is based on the convolutional neural network model, with the improved threshold function as the activation function tReLU. Residual neurons are introduced in the middle layer of the convolutional neural network, and convolution and pooling are alternately connected, and softmax classification is performed throughout the entire structure to generate a new convolutional neural network model RLCNN. The activation function tReLU is shown in formula (1):
[0051]
[0052] The residual neuron is expressed by formula (2):
[0053] F(x)=W2f(W1x+b)+b (2)
[0054] In formula (2), x represents the input of the current layer, F(x) represents the input of the next layer, W1 and W2 represent the weights of the current layer and the next layer respectively, f(.) represents the tReLU activation function, and b represents the bias.
[0055] The new convolutional neural network model is a six-layer convolutional neural network, in which the convolution kernel sizes of the first, second, fifth, and sixth layers are 5×5, 3×3, 3×3, and 3×3, respectively, and the number of convolution kernels is 32, 64, 128, and 256, respectively. Each layer uses a maximum pooling layer with a size of 2×2; the third and fourth layers are residual neuron layers established using residual neurons. The convolution kernel sizes of the residual neurons in the third and fourth layers are 1×1, 3×3, and 1×1, and the corresponding numbers are 16, 16, and 64, respectively; the structural parameters of the six-layer convolutional neural network are shown in Table 1:
[0056] The input of the new convolutional neural network model is the pixel intensity matrix of the grayscale image. The pixel intensity matrix of the grayscale image is obtained by two-dimensional image processing. The specific processing process is: first preprocess the two-dimensional image into a vibration signal dataset, and then convert the data of the vibration signal dataset into the pixel intensity matrix of the grayscale image using formula (3).
[0057]
[0058] Wherein, M represents a vibration data sequence of length M, j=1…M, k=1…M, represents the pixel intensity of the image, and round(·) represents normalizing the pixel value to 0-255.
[0059] A posture estimation algorithm is an algorithm that estimates the posture of an object or a human body through inputs such as sensor data or image data. Posture estimation algorithms are widely used in computer vision, robotics, virtual reality and other fields. Common posture estimation algorithms include: 1. Sensor-based posture estimation algorithm: The posture information of an object or a human body is obtained through sensors such as accelerometers, gyroscopes, and magnetometers, and then posture estimation is performed through algorithms such as filtering and integration. 2. Image-based posture estimation algorithm: The image of an object or a human body is obtained through a camera, and then posture estimation is performed through algorithms such as feature point matching and model fitting. 3. Deep learning-based posture estimation algorithm: The image of an object or a human body is trained through a deep learning model, and then the posture information is predicted through the model. 4. Sensor and image-based fusion posture estimation algorithm: The sensor and image data are fused, and posture estimation is performed through algorithms such as Kalman filtering and extended Kalman filtering. The application of posture estimation algorithms is very extensive. For example, in the field of robotics, posture estimation algorithms can be used for autonomous navigation and operation of robots; in the field of virtual reality, posture estimation algorithms can be used for tracking the user's hand and head posture; in the medical field, posture estimation algorithms can be used to monitor and evaluate the patient's posture, etc.
[0060] The attention module based on the convolutional neural network includes an attention vector generation unit, which is configured to feed the feature vector input by the residual module to a first branch and a second branch; wherein the first branch is configured to perform a deformable convolution operation, a channel attenuation operation, and a global pooling operation on the feature vector in the horizontal direction, and the second branch is configured to perform a deformable convolution operation, a channel attenuation operation, and a global pooling operation on the feature vector in the vertical direction;
[0061] The attention vector generation unit is also configured to concatenate the output of the first branch with the output of the second branch to obtain a concatenated vector, and transform the concatenated vector using a convolution transformation function; feed the transformed concatenated vector to the fully connected layer, and perform convolution operations on the input of the fully connected layer in the horizontal and vertical directions respectively to obtain the attention vector of the input feature vector in the horizontal direction and the attention vector of the input feature vector in the vertical direction.
[0062] The attention module involved in this embodiment uses deformable convolution for horizontal and vertical feature extraction, which facilitates the capture of object location information during subsequent encoding. Specifically, the first branch performs a deformable convolution operation and a channel attenuation operation on the feature vector in the horizontal direction to obtain a feature map in the horizontal direction, and the second branch performs a deformable convolution operation and a channel attenuation operation on the feature vector in the vertical direction to obtain a feature map in the vertical direction, that is, deformable convolution is only performed in one direction on one branch. Compared with the traditional deformable convolution performed in two directions at the same time, this embodiment can enhance the effect of feature extraction and improve the detection and recognition accuracy of the convolutional neural network. In addition, this embodiment uses a dual-branch design to pool the horizontal feature map and the vertical feature map respectively. The attention module can capture long-range dependencies along one spatial direction and retain precise position information along another spatial direction, so that both vertical and horizontal information are retained. Then, after a series of transformations, the attention vector is obtained and multiplied back to the original feature vector as a weight factor. In this way, the spatial attention and channel attention can be integrated, which solves the problem of unified operation of the existing attention mechanism in space and channel, and can improve the accuracy of the convolutional neural network.
[0063] The attention module also includes a weight allocation unit, which is configured to perform weight allocation on the input feature vector based on the attention vector in the horizontal direction and the attention vector in the vertical direction to obtain a weighted feature vector.
[0064] This embodiment proposes a method for monitoring the safety of cargo stacking in a cargo yard based on computer vision. Figure 1 As shown, the specific steps include:
[0065] Obtain cargo image data;
[0066] Based on the cargo image data, obtain the semantic features of the cargo;
[0067] Based on semantic features, obtain the three-dimensional posture parameters of the goods;
[0068] Based on the 3D posture parameters, the weights of different cargo areas are adjusted to obtain the posture information of the obscured cargo.
[0069] Based on the posture information of the obscured goods, the spatial relationship and force conditions between the goods are analyzed to predict dangerous situations.
[0070] Specifically, this embodiment uses an adaptive image enhancement algorithm to preprocess cargo images to improve image quality; constructs a multi-level cargo feature extraction model, comprehensively utilizes features such as color, texture, and shape, and extracts deep semantic features of cargo through a convolutional neural network; introduces an attention mechanism and contextual information to highlight target cargo, suppress background interference, and infer the posture of obscured cargo; adopts a lightweight posture recognition model to improve the model reasoning speed; combines the depth information of the cargo yard environment to construct a three-dimensional scene model of cargo stacking, analyzes the spatial relationship and force conditions of the cargo, evaluates the stacking stability and safety, and predicts possible dangerous situations such as tipping and collapse.
[0071] Furthermore, based on the cargo image data, obtaining the semantic features of the cargo includes:
[0072] Preprocess cargo image data to obtain high-quality cargo images;
[0073] Obtain the semantic features of goods based on high-quality goods images.
[0074] Specifically, cargo images collected in a cargo yard environment are acquired; whether the image quality meets a preset threshold is determined; if not, an adaptive image enhancement algorithm is executed; based on the brightness distribution of the cargo image, the image contrast is dynamically adjusted to improve the image layering; based on the average brightness value of the cargo image, the overall image brightness is adaptively adjusted to make the image brightness moderate; based on the color distribution of the cargo image, the image saturation is dynamically optimized to make the colors more vivid; the enhanced cargo image is evaluated through a convolutional neural network to determine whether the image quality meets the standards; if the image quality meets the standards, the high-quality cargo image is output for subsequent cargo recognition and classification tasks.
[0075] Furthermore, if Figure 2 As shown, preprocessing the cargo image data to obtain high-quality cargo images includes:
[0076] Adopting an adaptive algorithm to dynamically adjust the contrast, brightness and saturation of cargo image data to obtain an adjusted cargo image;
[0077] The convolutional neural network model is used to evaluate the adjusted cargo images to determine whether the image quality meets the standards. Based on the judgment results, high-quality cargo images are obtained.
[0078] Specifically, adaptive image enhancement algorithms adjust parameters based on the image's characteristics. For example, for images with insufficient illumination, a histogram equalization algorithm can be used to redistribute pixel values, thereby improving contrast and brightness. For blurred images, an unsharp masking algorithm can be used to enhance edges and details. Adaptive parameter adjustment can better adapt to image quality issues in different scenarios. For example, histogram equalization can adaptively adjust the mapping of pixel values based on the image's histogram distribution, eliminating the tedious process of manual parameter adjustment and improving the enhancement effect. Dynamically adjusting image contrast based on the brightness distribution of the product image can enhance the image's layering. For example, if a product image has a concentrated brightness distribution, it indicates low contrast and a lack of distinct layering. Contrast can be increased by stretching the image's brightness histogram. Specifically, darker areas in the brightness histogram are darkened, while brighter areas are brightened, thereby expanding the dynamic range of brightness values. This makes dark details clearer and bright details more prominent, thereby enhancing the image's layering. For example, pixel values originally between 50 and 150 are stretched to a range of 0-255, thereby increasing image contrast and enhancing the image's gradation. Based on the average brightness of the product image, the overall image brightness is adaptively adjusted to achieve a moderate brightness. For example, if the average brightness of an image is low, indicating an overall dark image, increasing the overall brightness can improve the visual quality. For example, if the average brightness of an image is 50, the brightness of each pixel can be increased by 30 to bring the average brightness to 80, achieving a moderate brightness. Conversely, if the average brightness is too high, the overall brightness can be reduced. Adaptive brightness adjustment prevents images from being too bright or too dark, ensuring a smooth visual experience. Based on the color distribution of the product image, the image saturation is dynamically optimized to enhance color vividness. For example, if the image has a relatively monotonous color distribution and low saturation, increasing the saturation can make the colors more vivid. For example, for an image of pale fruit, increasing the saturation can make the fruit's color more vivid and closer to its true color. Conversely, if the image's colors are too bright, reducing the saturation can make the colors more muted. Dynamically optimizing saturation can make image colors more natural and realistic. Enhanced cargo images are evaluated using a convolutional neural network to determine whether the image quality meets the quality standards. A convolutional neural network model is trained to score the quality of cargo images. The enhanced images are input into the trained convolutional neural network to obtain an image quality score. For example, images with a score above 9 are considered high-quality, while images with a score below 6 are considered low-quality. This method effectively assesses image quality, and the scoring threshold can be adjusted according to actual needs. If the image quality meets the standards, the high-quality cargo image is output for subsequent cargo identification and classification tasks.For example, after enhancement, cargo images with a quality score of 95 are output. These high-quality images are used to train cargo recognition and classification models or for cargo recognition and classification in real-world scenarios. High-quality images can improve recognition and classification accuracy, thereby enhancing cargo yard management efficiency.
[0079] Furthermore, if Figure 3 As shown in Figure 1, based on high-quality cargo images, the semantic features of the cargo are obtained, including:
[0080] Extract color features, texture features, shape features, and size features of high-quality cargo images to obtain multi-dimensional feature vectors;
[0081] A convolutional neural network model is used to perform deep learning and extraction of multi-dimensional feature vectors to obtain the semantic features of the goods.
[0082] Specifically, the color, texture, shape, and size features of the goods are extracted to obtain a multidimensional feature vector. Based on this multidimensional feature vector, a convolutional neural network model is used to perform deep learning and extraction of the goods' features, resulting in a deep semantic feature representation of the goods.
[0083] These multidimensional features together form the product's feature vector, which is used for subsequent classification and identification. The extracted multidimensional feature vector is then fed into a convolutional neural network (CNN) for deep learning and feature extraction. CNNs can automatically learn both local and global features of an image through multiple layers of convolution and pooling operations. For example, the first convolutional layer of a CNN can learn the product's edge information, the second convolutional layer can learn its texture information, and deeper convolutional layers can learn its shape and structure information. Assume that after CNN extraction, a 1024-dimensional feature vector is obtained. This vector represents the deep semantic features of the product and is more expressive and discriminative than manually designed features. Using the deep semantic features extracted by CNN, a support vector machine (SVM) classification model is constructed.
[0084] Furthermore, if Figure 4 As shown in the figure, based on the semantic features, the three-dimensional posture parameters of the goods are obtained, including:
[0085] Based on the semantic features, a three-dimensional model of the cargo is established, and key points on the cargo are predefined. Key points include corner points and edge points.
[0086] Based on predefined key points, detect key points in the cargo image and obtain the pixel coordinates of the key points;
[0087] Convert pixel coordinates into three-dimensional coordinates, calculate the spatial distance and angle relationship between key points based on the three-dimensional coordinates, and obtain the spatial relationship matrix between key points;
[0088] Based on the spatial relationship matrix, the three-dimensional posture parameters of the cargo are estimated.
[0089] Specifically, a three-dimensional model of the cargo is established based on its deep semantic features, pre-defining key points on the cargo, such as corners and edges. A key point detection algorithm is used to detect key points in the cargo image and obtain their pixel coordinates. Based on the camera calibration parameters, the pixel coordinates of the key points are converted into three-dimensional coordinates in the camera coordinate system. By calculating the spatial distances and angular relationships between the key points, a spatial relationship matrix between the key points is constructed. A pose estimation algorithm, such as the PnP algorithm, is used to estimate the three-dimensional pose parameters of the cargo relative to the camera, including the translation vector and rotation matrix, based on the three-dimensional coordinates of the key points and the spatial relationship matrix. Based on the estimated pose parameters, the cargo is determined to have undergone pose changes such as translation, rotation, tilt, and deformation, and the degree of pose change is quantified. If the pose change exceeds a preset threshold, the cargo pose is considered abnormal, triggering an alarm or taking appropriate action.
[0090] In this step, as an additional embodiment, a PnP algorithm is used to estimate the three-dimensional posture parameters of the cargo relative to the camera, including the translation vector and the rotation matrix, based on the three-dimensional coordinates of the key points and the spatial relationship matrix.
[0091] The camera captures images of the cargo and obtains the 2D coordinates of key points within the cargo image. Based on the pre-calibrated camera intrinsic parameter matrix and distortion coefficients, the 2D coordinates of the key points are corrected to obtain the corrected key point coordinates. Combined with the cargo's 3D model, the 2D coordinates of the key points are converted to 3D coordinates through spatial mapping. A spatial relationship matrix is constructed between the camera and the cargo, including the camera's intrinsic and extrinsic parameter matrices. Using the 3D coordinates of the key points and the spatial relationship matrix as input, the PnP algorithm is used for pose estimation to determine the cargo's rotation matrix and translation vector relative to the camera. Based on the obtained rotation matrix and translation vector, the cargo's 3D pose in the camera coordinate system is calculated to determine its spatial position and orientation relative to the camera. The cargo's 3D pose information is converted to pose parameters in the actual scene to determine its position and orientation in real space, providing a reference for subsequent cargo manipulation.
[0092] Furthermore, if Figure 5 As shown in the figure, based on the 3D posture parameters, the weights of different cargo areas are adjusted to obtain the posture information of the obscured cargo.
[0093] According to the three-dimensional posture parameters, the attention module adaptively adjusts different acquisition areas in the image to obtain the target cargo area, where the attention module is as follows: Figure 6 As shown;
[0094] Construct a spatial topological relationship diagram between goods based on the target cargo area;
[0095] A convolutional neural network model is used to extract features from the weighted cargo images to obtain visual feature representations of the cargo.
[0096] The spatial topological relationship graph and visual features are integrated to infer the posture information of the occluded goods through the graph convolutional network.
[0097] Specifically, the density of the stacked goods determines whether there is significant occlusion and interference. If so, an attention mechanism and contextual information are introduced for processing. An image of the stacked goods scene is captured, and the attention module adaptively adjusts the weights of different goods regions in the image, highlighting the target goods area and suppressing background interference areas. Based on the contextual information of the goods stacking, a spatial topological relationship graph between the goods is constructed to infer the pose of the occluded goods. A convolutional neural network is used to extract features from the weighted goods image to obtain a visual feature representation of the goods. The visual features of the goods are fused with the spatial topological relationships of the context, and the pose of the occluded goods is inferred using a graph convolutional network. Based on the inferred pose of the goods and the visual features, a multi-layer perceptron is used to classify the goods' pose, resulting in a pose recognition result. The pose recognition result is compared with a reference pose of the target goods to calculate the pose recognition accuracy. The model performance is evaluated based on the accuracy, and parameters are tuned as necessary to improve recognition accuracy.
[0098] A convolutional neural network is used to extract features from the weighted cargo images, and the visual feature representations of the cargo obtained include:
[0099] Obtain an image of the goods to be identified and preprocess the image, including size normalization and pixel value normalization. Based on a pre-trained convolutional neural network model, extract features from the pre-processed goods image to obtain a feature vector representing the visual characteristics of the goods. Use a support vector machine classifier to determine the category of the goods based on the extracted feature vector to determine the category to which the goods belong. If the confidence level of the determined goods category is lower than a preset threshold, use a nearest neighbor algorithm based on a pre-established library of goods image features to re-determine the goods category. Based on the result of the goods category determination, obtain the goods attribute information of the corresponding category, including the goods name, specifications, and model. Associate the goods attribute information with the goods image and store it in the goods identification result database. Statistically calculate the accuracy of the recognition results for different goods categories. For goods categories with an accuracy rate lower than the preset threshold, trigger retraining of the convolutional neural network model for the corresponding category.
[0100] More specifically, consider a warehouse environment where the cargo images captured by a camera may include a variety of different types of goods, such as boxed beverages and bagged rice. Image preprocessing is essential. Size normalization can be used to normalize images of varying resolutions to a consistent size, for example, resizing all images to 256x256 pixels for easier processing. Pixel value normalization normalizes the image's pixel values from 0-255 to 0-1. This accelerates model training and improves generalization. Next, a pretrained convolutional neural network (CNN) model is used to extract features from the preprocessed images. For example, a ResNet-50 model is used. This model extracts deep features from the image, generating a set of 1024-dimensional feature vectors. These feature vectors effectively represent the visual characteristics of the goods, such as color, texture, and shape. A support vector machine (SVM) classifier is used to determine the category of the goods based on the extracted feature vectors. Assume there are five categories of goods: beverages, rice, flour, cooking oil, and detergent. The SVM classifier learns the boundaries between different categories through training data. When a new feature vector is input, the SVM can determine the category to which it belongs. For example, a feature vector is classified as a beverage with a confidence level of 85. If the confidence level of the resulting product category falls below a preset threshold, such as 8, further classification is required. In this case, a pre-established product image feature library can be used to perform a further classification using the KNN algorithm. Assuming that the library already contains a large number of labeled product image feature vectors, the KNN algorithm finds the K samples closest to the current feature vector and votes based on the categories of these samples to determine the product category. For example, if four of the five closest samples in the library to the current feature vector belong to the rice category, the product is classified as rice. Based on the product category determination result, the product attribute information for the corresponding category is obtained. Suppose there is a database that stores detailed information on various types of goods. For example, for beverages, attribute information includes brand, volume, and production date. By querying the database, the attribute information of a specific product, such as a 500ml bottle of a certain brand of beverage, can be obtained. Goods attribute information is associated with the goods image and stored in the goods recognition result database. This facilitates subsequent querying and management. For example, in a warehouse management system, detailed information about the corresponding goods can be quickly found through the image. The accuracy of the recognition results is calculated for different goods categories. Assuming that 100 recognitions are performed for each category, the accuracy for beverages is 95%, while the accuracy for rice is only 75%. For goods categories with an accuracy below a preset threshold, such as 80%, retraining of the convolutional neural network model for the corresponding category is triggered. This retraining process can be performed by increasing the number of training samples for that category, adjusting model parameters, and so on.For example, for rice, more rice images with varying lighting and angles can be added, and the ResNet-50 model retrained to improve its accuracy in rice classification. This process not only enables efficient cargo identification but also continuously optimizes the recognition model, improving the accuracy and reliability of the overall system. Size and pixel value normalization ensure the consistency of image input. The feature vectors extracted by the CNN model provide rich visual information. The combined use of the SVM and KNN classifiers improves the accuracy of classification. The association and storage of cargo attribute information facilitates subsequent management. Accuracy statistics and model retraining form a closed-loop optimization mechanism to ensure continuous system improvement. This multi-layered, multi-technique integration not only improves cargo identification accuracy but also enhances the system's robustness and adaptability, enabling it to cope with complex and changing warehouse environments and ensure efficient and accurate cargo management.
[0101] Image data is preprocessed, including image size and pixel normalization. For example, all images are resized to 224x224 pixels, and pixel values are normalized to between 0 and 1 to eliminate size and brightness differences between images and improve model training efficiency. The preprocessed image data is then fed into a pretrained lightweight posture recognition model, such as MobileNetV3. This lightweight model is chosen for its low parameter count and fast computational speed, making it suitable for deployment on resource-constrained edge devices and meeting real-time requirements. The current posture state of the cargo is determined based on the output of the cargo posture recognition model. The model output can be posture angle values, such as Euler angles (yaw, pitch, and roll), or posture categories, such as "upright," "tilted," and "inverted." If the cargo posture deviates significantly from the preset standard posture, for example, if the tilt exceeds 30 degrees, an alert is triggered, generating an abnormal cargo posture message, such as "Cargo A is tilted at 45 degrees." This abnormality information is then sent to the monitoring platform, enabling personnel to take timely action. The model's recognition accuracy in real-world applications is regularly evaluated, for example, by calculating metrics such as precision and recall on a test dataset weekly or monthly. The model structure is continuously optimized and redundant parameters are pruned. For example, model pruning techniques are used to remove neuronal connections that contribute less to recognition accuracy. This reduces model size and computational complexity, improving inference speed while maintaining recognition accuracy. Assuming the initial model's inference speed is 15 frames per second, the pruned model can reach 25 frames per second, meeting real-time requirements. An incremental learning approach is employed to regularly fine-tune and update the model using newly collected cargo pose data. For example, the model is fine-tuned every time 1,000 new cargo pose-annotated images are added to adapt to new cargo pose variations, such as a new cargo stacking method. This incremental learning approach avoids retraining the entire model, saving time and resources. During model inference, the batch size and parallel computing are configured. For example, setting the batch size to 32 fully utilizes the parallel computing capabilities of hardware accelerators such as GPUs, further improving model inference speed and real-time performance. For example, using batch processing and GPU acceleration can increase the model's inference speed from 25 frames per second to 40 frames per second. Multiple specialized gesture recognition sub-models are built for different types of cargo, such as boxed, bagged, and barreled. Because different types of cargo have distinct shapes and gesture characteristics, using specialized models can improve recognition accuracy. Dynamically select the appropriate sub-model for gesture recognition based on the cargo type. For example, identify the cargo type based on its barcode and then select the appropriate sub-model. Deploy distributed edge computing nodes to perform gesture recognition local computations at the cargo monitoring site, such as deploying an edge computing device near each camera. This reduces data transmission latency and improves the system's real-time responsiveness.For example, uploading image data to the cloud for processing might take one second, while processing it on the edge device only takes one second. The central node aggregates and analyzes the recognition results of each edge node, such as counting the number and type of various posture anomalies, to provide decision support for warehouse management.
[0102] Furthermore, based on the posture information of the obscured goods and combined with the three-dimensional posture parameters of the goods, the spatial relationship and force conditions between the goods are analyzed to predict dangerous situations, including:
[0103] Based on the posture information of the obscured goods, combined with the three-dimensional posture parameters of the goods and the spatial relationship between the goods, a physical simulation engine is used to simulate the force conditions of the goods and evaluate the stability of the cargo stacking.
[0104] Specifically, the depth information of the cargo yard environment is obtained, and the three-dimensional spatial data of the cargo yard is collected through devices such as depth cameras or lidar; based on the depth information obtained, three-dimensional reconstruction technologies such as point cloud stitching, surface reconstruction and other algorithms are used to construct a three-dimensional scene model of cargo stacking; in the three-dimensional scene model, the three-dimensional posture information of each cargo is identified, including attributes such as position, direction and size; the spatial relationship between the cargoes in the three-dimensional scene model is analyzed, and parameters such as the distance and contact area between the cargoes are calculated to determine whether the cargo stacking is reasonable; based on the material, weight and other attributes of the cargoes, combined with the spatial relationship between the cargoes, the physical simulation engine is used to simulate the stress conditions of the cargoes and evaluate the stability of the cargo stacking; a safety threshold for the stability of the cargo stacking is set, and when the stability of the cargo stacking is lower than the threshold, it is determined that there are safety hazards such as tipping and collapse; based on the stability analysis results of the cargo stacking, early warnings are issued for cargoes with safety hazards in the three-dimensional scene, and suggestions for optimizing the cargo stacking method are provided to guide staff to restack the cargoes and eliminate safety hazards.
[0105] This embodiment also provides a computer vision-based cargo storage safety monitoring system in a cargo yard. Figure 7 The following modules are included: image acquisition module, feature extraction module, posture estimation module and 3D analysis module:
[0106] Image acquisition module, used to collect cargo images;
[0107] Feature extraction module, used to extract semantic features of cargo images;
[0108] Posture estimation module, used to estimate the three-dimensional posture of the cargo;
[0109] The three-dimensional analysis module is used to analyze the spatial relationship and stress conditions between goods and predict dangerous situations.
[0110] The above are merely preferred embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. A method for monitoring the safety of cargo stacking in a cargo yard based on computer vision, characterized in that: include: Obtain cargo image data; Based on the cargo image data, obtaining semantic features of the cargo; Based on the semantic features, obtaining three-dimensional posture parameters of the goods; Based on the three-dimensional posture parameters, weights of different cargo areas are adjusted to obtain posture information of the obscured cargo; Based on the posture information of the obstructed goods, the spatial relationship and force conditions between the goods are analyzed to predict dangerous situations.
2. The method for monitoring the safety of cargo stacking in a cargo yard based on computer vision according to claim 1 is characterized in that: Based on the cargo image data, obtaining the semantic features of the cargo includes: Preprocessing the cargo image data to obtain high-quality cargo images; The semantic features of the goods are acquired according to the high-quality goods image.
3. The method for monitoring the safety of cargo stacking in a cargo yard based on computer vision according to claim 2 is characterized in that: Preprocessing the cargo image data to obtain high-quality cargo images includes: Adopting an adaptive algorithm to dynamically adjust the contrast, brightness and saturation of the cargo image data to obtain an adjusted cargo image; The adjusted cargo image is evaluated using a convolutional neural network model to determine whether the image quality meets the standard, and the high-quality cargo image is obtained based on the determination result.
4. The method for monitoring the safety of cargo stacking in a cargo yard based on computer vision according to claim 2 is characterized in that: Acquiring semantic features of the goods according to the high-quality goods image includes: Extracting color features, texture features, shape features, and size features of the high-quality cargo image to obtain a multi-dimensional feature vector; A convolutional neural network model is used to perform deep learning and extraction on the multi-dimensional feature vector to obtain the semantic features of the goods.
5. The method for monitoring the safety of cargo stacking in a cargo yard based on computer vision according to claim 1, characterized in that: Based on the semantic features, obtaining the three-dimensional posture parameters of the goods includes: According to the semantic features, a three-dimensional model of the goods is established, and key points on the goods are predefined, wherein the key points include: corner points and edge points; Based on the predefined key points, detect the key points in the cargo image and obtain the pixel coordinates of the key points; Convert the pixel coordinates into three-dimensional coordinates, calculate the spatial distance and angle relationship between key points according to the three-dimensional coordinates, and obtain the spatial relationship matrix between the key points; Based on the spatial relationship matrix, three-dimensional posture parameters of the cargo are estimated.
6. The method for monitoring the safety of cargo stacking in a cargo yard based on computer vision according to claim 1, characterized in that: Based on the three-dimensional posture parameters, the weights of different cargo areas are adjusted to obtain the posture information of the obscured cargo, including: According to the three-dimensional posture parameters, the attention module adaptively adjusts different acquisition areas in the image to acquire the target cargo area; According to the target cargo area, a spatial topological relationship diagram between cargoes is constructed; A convolutional neural network model is used to extract features from the weight-adjusted cargo images to obtain visual feature representations of the cargo; The spatial topological relationship graph and visual features are integrated, and the posture information of the occluded goods is inferred through a graph convolutional network.
7. The method for monitoring the safety of cargo stacking in a cargo yard based on computer vision according to claim 1, characterized in that: Based on the posture information of the obstructed goods and in combination with the three-dimensional posture parameters of the goods, the spatial relationship and force conditions between the goods are analyzed, and the dangerous situation is predicted including: Based on the posture information of the obscured goods, combined with the three-dimensional posture parameters of the goods and the spatial relationship between the goods, a physical simulation engine is used to simulate the force conditions of the goods and evaluate the stability of the goods stacking.
8. Computer vision-based cargo stacking safety monitoring system in cargo yard, characterized by: include: Image acquisition module, feature extraction module, posture estimation module and 3D analysis module: The image acquisition module is used to acquire images of goods; The feature extraction module is used to extract the semantic features of the cargo image; The posture estimation module is used to estimate the three-dimensional posture of the cargo; The three-dimensional analysis module is used to analyze the spatial relationship and stress conditions between goods and predict dangerous situations.
Citation Information
Patent Citations
Three-dimensional attitude acquisition method, model training method and related equipment
CN115471863A
Attitude analysis method and device, electronic equipment and storage medium
CN116959102A
Shielding-oriented human body posture estimation method and device and electronic equipment
CN117671800A
Multidimensional perception identification method and system for goods in carriage
CN118506229A
Goods volume calculation method and system based on panoramic video
CN119068042A
Cited By
Stacked object identification system and method
CN121074869A