Deep learning-based plane parking space state identification method and system

Through deep learning and image processing technology, combined with the improved YOLOv11 object detection and adaptive denoising centroid positioning algorithm, efficient and low-cost library positioning status recognition is achieved, solving the problems of high cost, small range and poor environmental adaptability of library position status recognition in the existing technology, and improving the automation and accuracy of warehouse management.

CN120472205APending Publication Date: 2025-08-12XIAMEN INTRETECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510488400.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The existing warehouse location status recognition technology has problems such as high cost, small monitoring range, poor environmental adaptability and a lot of manual intervention, which affects the efficiency and accuracy of warehousing management.

Method used

The plane library position state recognition method based on deep learning is adopted, combined with the improved YOLOv11 object detection algorithm, adaptive denoising center of mass positioning algorithm and grid library position construction algorithm, and image processing is used to achieve automatic extraction of library position angles and center of mass and library position status recognition.

Benefits of technology

It improves the degree of automation, reduces labor costs and artificial errors, enhances environmental adaptability, expands the monitoring range, reduces hardware investment and power consumption, optimizes the allocation of warehouse location resources, and improves operational efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472205A_ABST
    Figure CN120472205A_ABST
Patent Text Reader

Abstract

The invention provides a plane parking space state recognition method and system based on deep learning. The plane parking space state recognition method comprises the steps that an overall image of a parking space area is collected in real time; an improved YOLOv11 target detection algorithm is adopted to identify the overall image of the storage location area, an image containing storage location angle information is obtained, and the improved YOLOv11 target detection algorithm introduces multi-dimensional dynamic convolution in the feature extraction network; carrying out centroid coordinate calculation on the image containing the storage location angle information through an adaptive denoising centroid positioning algorithm to generate a centroid coordinate point set, and inputting the centroid coordinate point set into a grid type storage location construction algorithm to generate a storage location layout; and according to the storage location layout, obtaining a corresponding storage location area local image in real time, and performing state recognition on the storage location area local image through an image classification algorithm to obtain storage location state information. According to the invention, through integration of deep learning and a target detection algorithm, the storage location angle and the mass center are automatically extracted, the identification accuracy can be improved, and the cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of flat warehouse location status recognition, and specifically to a flat warehouse location status recognition method and system based on deep learning. Background Art

[0002] Warehouse management, as a crucial link in the supply chain, directly impacts logistics efficiency and operating costs. Flat-space management is a warehouse management model characterized by storing goods flat on the ground or in simple stacks, rather than using three-dimensional racks. This model is commonly used in warehousing scenarios involving large equipment, heavy parts, or short-term turnover. Traditional flat-space management still relies heavily on manual intervention. Tasks such as warehousing, shipping, and inventory counting rely on manual operations and management, which not only increases labor intensity but also easily leads to operational errors and inefficient management.

[0003] In recent years, with the rapid development of Industry 4.0 and artificial intelligence (AI) technology, warehouse location management systems are transitioning from traditional manual operations to intelligent and automated systems. Modern intelligent warehousing systems, combining AI with the Internet of Things (IoT) technology, enable automatic identification, precise storage and retrieval, and intelligent scheduling, significantly improving the efficiency and accuracy of warehouse management. Automated guided vehicles (AGVs) and location status recognition are applied to warehouse management. The system identifies the location status and generates location information, which is then transmitted to the AGV. The AGV then moves goods based on this information. Therefore, the software and hardware solutions for location status recognition directly impact key functions such as automatic positioning, real-time data synchronization, and precise inventory control. They also help optimize location resource allocation, reduce human error, and improve overall operational efficiency.

[0004] Currently, solutions for identifying storage location status mainly include: (1) single-point laser radar, which determines the distance to an object by emitting a laser beam and measuring the reflection time, and detects the occupancy status of the storage location; (2) RGB-D cameras, which combine RGB images and depth information to provide both color and distance data. The camera's built-in computing power automatically identifies the presence of goods in the selected area. However, these solutions still have the following shortcomings:

[0005] Manual external parameter calibration: Although these solutions can automatically identify the storage location status, manual calibration of the external parameter identification area is still required, which not only increases labor costs but also may introduce human errors.

[0006] Environmental adaptability issues: Sensor performance may be affected in complex environments, such as those with unstable lighting conditions, high dust concentrations, or the presence of highly reflective objects. RGB-D cameras may experience data deviations in low-light environments or when facing distant objects. LiDAR, which can only emit a single laser beam at a time, may not detect the slightest gaps between cartons or pallets. This can lead to misidentification of storage location status and increase the risk of stacking accidents. In these situations, the sensor's accuracy and reliability may not meet requirements.

[0007] Small monitoring range: The monitoring capability of a single-point lidar is limited to a single storage location, while the effective monitoring distance of an RGB-D camera is limited and the coverage area is small.

[0008] High Cost: When deployed on a large scale, single-point LiDAR systems require a separate sensor for each storage location, and hardware costs rise significantly as the number of locations increases. RGB-D cameras, which need to process high-resolution images and depth data in real time, rely on high-performance hardware, resulting in high device cost and power consumption. Overall, the high hardware requirements and deployment costs of LiDAR and RGB-D cameras make the overall system investment substantial, placing a heavy financial burden on cash-strapped businesses.

[0009] In view of this, the present invention provides an efficient, low-cost, and highly adaptable planar storage location status recognition method and system by integrating deep learning, target detection, and image classification technologies, which can significantly improve the level of intelligent warehouse management. Summary of the Invention

[0010] In order to solve the problems of high cost and small monitoring range of existing storage location status identification, the present invention provides a planar storage location status identification method and system based on deep learning to solve the above technical defects.

[0011] In a first aspect, the present invention proposes a planar storage location status recognition method based on deep learning, the method comprising the following steps:

[0012] S1. Real-time acquisition of an overall image of the storage area, which includes multiple storage areas divided by storage lines;

[0013] S2. Use the improved YOLOv11 target detection algorithm to identify the entire image of the storage area and obtain an image containing storage location angle information. The improved YOLOv11 target detection algorithm introduces multi-dimensional dynamic convolution into the feature extraction network;

[0014] S3, calculating the centroid coordinates of the image containing the storage location angle information using an adaptive denoising centroid positioning algorithm to generate a centroid coordinate point set, and inputting the centroid coordinate point set into a grid-type storage location construction algorithm to generate a storage location layout;

[0015] S4. Acquire the corresponding local image of the storage area in real time according to the storage location layout, and perform status recognition on the local image of the storage area through an image classification algorithm to obtain storage location status information.

[0016] Preferably, in step S2, an improved YOLOv11 target detection algorithm is used to identify the entire image of the storage area to obtain an image containing storage location angle information, which specifically includes the following sub-steps:

[0017] S21, using a multi-dimensional dynamic convolutional feature extraction network to extract features from the entire image of the storage location area, and obtain a first feature map y1, a second feature map y2, and a third feature map y3;

[0018] S22, input the first feature map y1, the second feature map y2 and the third feature map y3 into the feature pyramid network to fuse and exchange information of high-resolution features and low-resolution features, and obtain the enhanced first feature map y′1, the enhanced second feature map y′2 and the enhanced third feature map y′3;

[0019] S23, inputting the enhanced first feature map y′1, the enhanced second feature map y′2, and the enhanced third feature map y′3 into the detection head for predicting whether each feature point contains a storage position angle and corresponding regression parameters;

[0020] S24. Decode the prediction result to correspond to the position of the overall image of the storage area, and perform score sorting and non-maximum suppression screening to obtain the final storage area angle prediction frame, and finally output the image containing the storage area angle information through the prediction frame.

[0021] Preferably, in step S2, the improved YOLOv11 target detection algorithm introduces multi-dimensional dynamic convolution in the feature extraction network, and the expression is:

[0022]

[0023] Where Y (l) Represents the weight of the dynamic convolution kernel; Represents the weight sum of K convolution kernels; * represents the convolution operation; represents the bias sum of K convolution kernels; π k Represents the attention weight value of the kth convolution kernel; s k represents the attention weight value of the kth convolution kernel position space; X (l-1) Represents the output or input feature map of the previous layer; Represents the weight of the k-th convolution kernel; Represents the bias of the k-th convolution kernel.

[0024] In step S21, multi-dimensional dynamic convolution is used to extract features from the overall image of the storage area to obtain a feature map, which specifically includes the following sub-steps:

[0025] S211, performing global average pooling and global maximum pooling operations on the entire image of the storage location area to generate a first eigenvector and a second eigenvector respectively;

[0026] S212, performing full connection layer and activation function processing on the first eigenvector and the second eigenvector to obtain a first branch output and a second branch output respectively;

[0027] S213, use the Sigmoid function to normalize the output of the first branch and generate the position space attention weight s k , use the Softmax function to normalize the output of the second branch and generate the channel space attention weight π k ;

[0028] S214, the position space attention weight s k and channel-space attention weight π k Applied to the corresponding convolution kernels respectively to adjust the weights of the convolution kernels;

[0029] S215. Use the adjusted convolution kernel to perform convolution operation, batch normalization and activation function processing on the input feature map in sequence to generate a feature map.

[0030] Preferably, in step S3, centroid coordinates are calculated for the image containing the storage position angle information by using an adaptive denoising centroid positioning algorithm to generate a centroid coordinate point set, which specifically includes the following sub-steps:

[0031] S311, converting the image containing the storage position angle information into a grayscale image, and performing filtering and denoising processing on the grayscale image;

[0032] S312, performing adaptive threshold binarization processing on the denoised grayscale image to obtain a binary image, and performing morphological processing on the binary image;

[0033] S313 , calculating the image moment of the binary image after morphological processing, and calculating the centroid coordinates of the white area according to the image moment to generate a centroid coordinate point set.

[0034] Preferably, in step S3, the centroid coordinate point set is input into a grid-type storage location construction algorithm to generate a storage location layout, which specifically includes the following sub-steps:

[0035] S321, using a convex hull algorithm to process the centroid coordinate point set to obtain a point sequence constituting an outer boundary, and screening out key points suitable for constructing a quadrilateral from the point sequence;

[0036] S322, calculate the distance between two adjacent key points and obtain the longest side;

[0037] S323, calculating the vertical distances from all points in the centroid coordinate point set to the straight line on which the longest side lies, and dividing the points in the centroid coordinate point set into nearby points and remaining points according to a distance threshold;

[0038] S324. Based on the nearby points and the remaining points, iteratively construct convex quadrilaterals, each convex quadrilateral representing a storage location layout.

[0039] Further preferably, in step S324, convex quadrilaterals are iteratively constructed based on the nearby points and the remaining points, each convex quadrilateral represents a storage location layout, specifically including the following sub-steps:

[0040] a. Determine whether the set of remaining points is empty. If so, return the result directly. If not, determine whether the number of nearby points is greater than 1. If so, execute step b. If not, return the result directly.

[0041] b. Sort the nearby points by x-coordinate, and use the distance function to calculate the Euclidean distance between all nearby points and the remaining points to obtain the distance matrix;

[0042] c. Select the remaining points with the shortest Euclidean distance and not yet matched as matching points in the order of the x-coordinates of the nearby points, and record the indexes of the remaining matched points;

[0043] d. Sequentially combine two adjacent nearby points with their corresponding two nearest matching points to form a convex quadrilateral. The vertex order of the convex quadrilateral is: the first nearby point, the second nearby point, the second nearest matching point, and the first nearest matching point.

[0044] e. Update the set of nearby points and the set of remaining points, remove the remaining matched points, and repeat steps ad until there are no remaining points;

[0045] f. Output a list of convex quadrilaterals, where each convex quadrilateral represents a storage location layout.

[0046] Preferably, in step S4, a local image of the corresponding storage area is acquired in real time according to the storage location layout, and the status of the local image of the storage area is identified by an image classification algorithm to obtain storage location status information, which specifically includes the following sub-steps:

[0047] S41. Acquire a local image of the corresponding storage area in real time according to the storage location layout;

[0048] S42, extracting features from the local image of the storage location area through a convolutional layer to obtain an initial storage location feature map;

[0049] S43, performing a downsampling operation on the initial storage location feature map, and performing feature extraction using the MobileViT block to obtain the final storage location feature map;

[0050] S44. Perform state classification on the final storage location feature map through the convolution layer and the linear layer to obtain the storage location status information.

[0051] In the second aspect, the present invention proposes a planar storage location status recognition system based on deep learning, which includes: an RGB camera, an AI edge computing box and a server;

[0052] The server sends a command to the AI edge computing box to obtain an overall image of the storage area. The AI edge computing box uses multiple connected RGB cameras to collect the overall image of the storage area in real time. The overall image of the storage area includes multiple storage areas divided by storage lines.

[0053] The AI edge computing box uses an improved YOLOv11 target detection algorithm to identify the entire image of the storage area and obtain an image containing storage location angle information. The improved YOLOv11 target detection algorithm introduces multi-dimensional dynamic convolution into the feature extraction network.

[0054] The AI edge computing box calculates the centroid coordinates of images containing storage location angle information using an adaptive denoising centroid positioning algorithm to generate a centroid coordinate point set. This point set is then input into a grid-type storage location construction algorithm to generate a storage location layout.

[0055] The AI edge computing box obtains the corresponding local images of the storage area in real time according to the storage location layout, and uses the image classification algorithm to identify the status of the local images of the storage area to obtain the storage location status information.

[0056] In the third aspect, the present invention proposes a terminal device, comprising a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the computer program, it implements the steps of the planar storage location status identification method based on deep learning as described in any one of the above items.

[0057] In a fourth aspect, the present invention proposes a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps of the planar storage location status recognition method based on deep learning as described in any one of the above items.

[0058] Compared with the prior art, the present invention has the following advantages:

[0059] (1) High degree of automation: Through deep learning and target detection algorithms, the automatic extraction of storage position angle and center of mass is realized, avoiding the tedious manual external parameter calibration and significantly reducing labor costs and human errors.

[0060] (2) Strong environmental adaptability: The RGB camera is combined with an adaptive denoising algorithm to maintain high recognition accuracy even in complex environments with large lighting changes, high dust content, or highly reflective materials.

[0061] (3) Wide monitoring range: By connecting multiple RGB cameras and performing image integration processing by the AI edge computing box, the monitoring range is effectively expanded, overcoming the problem of small monitoring range of single-point lidar and RGB-D camera.

[0062] (4) High cost-effectiveness: Using relatively low-cost RGB cameras to replace expensive lidar and RGB-D cameras greatly reduces hardware investment and power consumption, which is particularly suitable for companies with limited funds.

[0063] (5) Flexibility and accuracy: Through the grid-type storage location construction algorithm and the adaptive denoising centroid positioning algorithm, the system can adapt to different types of storage location layouts, provide sufficient space for manual intervention, and ensure the accuracy and flexibility of the storage location layout.

[0064] (6) Efficient image processing: The improved YOLOv11 target detection algorithm and MobileViT image classification algorithm are used to achieve efficient image processing and storage location status recognition, ensuring the real-time and efficiency of the system.

[0065] (7) Intelligent scheduling and management: By working in collaboration with the server and scheduling system, intelligent inbound and outbound management is achieved, which optimizes the allocation of storage resources, reduces human errors, and improves overall operational efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] Other features, objects and advantages of the present application will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings:

[0067] Figure 1 1 is a flow chart of a planar storage location status recognition method based on deep learning according to the present invention;

[0068] Figure 2 Schematic diagram of the improved YOLOv11 network structure according to the present invention;

[0069] Figure 3 2. It is a schematic diagram of a multi-dimensional dynamic convolution structure according to the present invention;

[0070] Figure 4 2. It is a schematic diagram of the process of the adaptive denoising centroid positioning algorithm according to the present invention;

[0071] Figure 5 This is a schematic diagram of the grid-type storage location construction process according to the present invention;

[0072] Figure 6 is a schematic diagram of the MobileViT network structure according to the present invention;

[0073] Figure 7a The overall image of the storage area collected is shown;

[0074] Figure 7b An image containing storage position angle information is shown;

[0075] Figure 7c A schematic image of the centroid positioning is shown;

[0076] Figure 7d An image of automated storage area generation is shown;

[0077] Figure 7e A schematic diagram of the storage location status results is shown;

[0078] Figure 8 is a structural diagram of a planar storage location status recognition system based on deep learning according to the present invention;

[0079] Figure 9 It is a structural diagram of a computer system suitable for implementing the electronic device of the embodiment of the present application. DETAILED DESCRIPTION

[0080] The present application will be further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are merely for the purpose of explaining the relevant invention and are not intended to limit the invention. It should also be noted that, for ease of description, only portions relevant to the relevant invention are shown in the accompanying drawings.

[0081] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0082] Figure 1 The flow chart of the planar storage location status recognition method based on deep learning of the present invention is shown as follows: Figure 1 As shown, the method includes the following steps:

[0083] S1. Real-time collection of an overall image of the storage area. The overall image of the storage area includes multiple storage areas divided by storage lines.

[0084] S2. Use the improved YOLOv11 target detection algorithm to identify the overall image of the storage area and obtain an image containing storage location angle information. The improved YOLOv11 target detection algorithm introduces multi-dimensional dynamic convolution into the feature extraction network.

[0085] Figure 2The improved YOLOv11 network structure diagram according to the present invention is shown as follows: Figure 2 As shown in the figure, the improved YOLOv11 network structure includes a feature extraction network and a feature pyramid network (FPN), and its processing flow is as follows:

[0086] S21. Utilize a multidimensional dynamic convolutional feature extraction network to extract features from the overall image of the storage area, obtaining a first feature map y1, a second feature map y2, and a third feature map y3. Specifically, the input overall image of the storage area first passes through multiple MD-C3k2 modules, which extract features at different scales layer by layer using multidimensional dynamic convolution operations. Each MD-C3k2 module includes a multidimensional dynamic convolution layer to capture both local and global features in the image.

[0087] S22: Input the first feature map y1, the second feature map y2, and the third feature map y3 into the feature pyramid network to fuse and exchange the information of high-resolution features and low-resolution features, and obtain the enhanced first feature map y′1, the enhanced second feature map y′2, and the enhanced third feature map y′3. That is, the extracted features are fused through the feature pyramid network (FPN). In this process:

[0088] Upsampling: The low-resolution feature map is enlarged to the same size as the high-resolution feature map through an upsampling operation.

[0089] Concatenation: The upsampled low-resolution feature map and the high-resolution feature map are concatenated along the channel dimension to form a fused feature map. This operation integrates feature information at different scales and enhances the expressiveness of features.

[0090] S23: The enhanced first feature map y′1, the enhanced second feature map y′2, and the enhanced third feature map y′3 are input to the detection head to predict whether each feature point contains the storage position angle and the corresponding regression parameters. The fused feature maps are further processed by the C3k2 module and the convolutional (Conv) layer to generate the final three-scale feature maps. These feature maps contain rich multi-scale information for subsequent object detection tasks.

[0091] S24. Decode the prediction result to correspond to the position of the overall image of the storage area, and perform score sorting and non-maximum suppression screening to obtain the final storage area angle prediction frame, and finally output the image containing the storage area angle information through the prediction frame.

[0092] in, Figure 3The figure shows a schematic diagram of the multidimensional dynamic convolution structure according to the present invention. The improved YOLOv11 target detection algorithm introduces multidimensional dynamic convolution into the feature extraction network, so that the feature extraction network can adjust the receptive field of the convolution kernel in a targeted manner according to the characteristics of the input data to better capture detailed information.

[0093] It should be understood that the definition of static convolution is:

[0094] Y (l) =W (l) *X (l-1) +b (l)

[0095] Where Y(l) is the output feature map of the static convolution layer, X(l-1) is the output or input feature map of the previous layer, W(l) is the weight of the static convolution kernel, * indicates the convolution operation, and b(l) is the bias term.

[0096] This application uses multi-dimensional dynamic convolution, and the specific expression is:

[0097]

[0098] Where Y (l) Represents the weight of the dynamic convolution kernel; Represents the weight sum of K convolution kernels; * represents the convolution operation; Represents the bias sum of K convolution kernels; π k Represents the attention weight value of the kth convolution kernel; s k represents the attention weight value of the kth convolution kernel position space; X (l-1) Represents the output or input feature map of the previous layer; Represents the weight of the k-th convolution kernel; Represents the bias of the k-th convolution kernel.

[0099] In step S21, multi-dimensional dynamic convolution is used to extract features from the overall image of the storage area to obtain a feature map, which specifically includes the following sub-steps:

[0100] S211, performing global average pooling and global maximum pooling operations on the entire image of the storage location area to generate a first eigenvector and a second eigenvector respectively;

[0101] S212, performing full connection layer and activation function processing on the first eigenvector and the second eigenvector to obtain a first branch output and a second branch output respectively;

[0102] S213, use the Sigmoid function to normalize the output of the first branch and generate the position space attention weight s k, use the Softmax function to normalize the output of the second branch and generate the channel space attention weight π k ;

[0103] S214, the position space attention weight s k and channel-space attention weight π k Applied to the corresponding convolution kernels respectively to adjust the weights of the convolution kernels;

[0104] S215. Use the adjusted convolution kernel to perform convolution operation, batch normalization and activation function processing on the input feature map in sequence to generate a feature map.

[0105] Figure 4 The schematic diagram of the adaptive denoising centroid positioning algorithm according to the present invention is shown, combined with reference Figure 1 and Figure 4 The planar storage location status recognition method based on deep learning proposed in the present invention further includes the following steps:

[0106] S3. Calculate the centroid coordinates of the image containing the storage location angle information through an adaptive denoising centroid positioning algorithm to generate a centroid coordinate point set, and input the centroid coordinate point set into a grid-type storage location construction algorithm to generate a storage location layout.

[0107] In this embodiment, the centroid coordinates of the image containing the storage position angle information are calculated by an adaptive denoising centroid positioning algorithm to generate a centroid coordinate point set, which specifically includes the following sub-steps:

[0108] S311: Convert the image containing the storage position angle information into a grayscale image, and perform filtering and denoising on the grayscale image. The filtering and denoising process may first use Gaussian blurring to smooth the image and reduce high-frequency noise, and then use median filtering to remove salt and pepper noise and improve edge preservation.

[0109] S312. Perform adaptive threshold binarization on the denoised grayscale image to obtain a binary image. For each pixel (x, y) in the image, calculate the average or weighted average within the local window. Subtract a constant C from this value as the threshold T(x, y) to binarize the pixel value. In contrast, the Otsu method is a global binarization method that determines a unified global threshold by analyzing the grayscale histogram of the entire image. It is suitable for images with uniform illumination, but may not accurately reflect the actual conditions of each area when there are significant local variations.

[0110] Morphological processing is performed on binary images, specifically by first performing an opening operation followed by a closing operation. The opening operation begins with an erosion operation, which helps remove small noise points and isolated white areas in the image; a dilation operation is then performed to restore the basic shape of the target. Next, the closing operation performs a dilation operation to fill small holes in the target area, followed by an erosion operation to restore edge details. This opening-then-closing process effectively improves image connectivity and smoothness, facilitating subsequent target recognition and analysis.

[0111] S313 , calculating the image moment of the binary image after morphological processing, and calculating the centroid coordinates of the white area according to the image moment to generate a centroid coordinate point set.

[0112] The spatial moment of an image is defined as:

[0113]

[0114] Where I(x,y) is the pixel value at pixel (x,y) in the binary image, x and y represent the coordinates of the pixel in the image, and i and j are the corresponding powers that determine the order of the moment.

[0115] Center of mass coordinates (c x ,c y ) is calculated as:

[0116]

[0117] Among them, m 00 is the zero-order moment, representing the sum of white pixels; m 10 represents the x-component of the first-order moment, m 01 Represents the y-component of the first moment.

[0118] Figure 5 The schematic diagram of the grid type storage location construction process according to the present invention is shown, combined with reference Figure 1 and Figure 5 If the storage locations in the storage area are arranged in a grid pattern and the number of locations in each row is the same, the grid-type storage location construction algorithm can be used to automatically construct the storage area based on the input point set. The specific steps include the following:

[0119] S321. Use the convex hull algorithm to process the centroid coordinate point set to obtain a point sequence that constitutes the outer boundary. These points form a convex polygon that represents the outermost contour of the point set.

[0120] And filter out key points suitable for constructing quadrilaterals from the point sequence. The specific screening method is as follows:

[0121] For each point on the convex hull, the algorithm takes the previous point, the current point, and the next point and calculates the angle between them. The calculation method used here is to first calculate the vectors formed by the current point and the previous point, and the current point and the next point, then calculate the angle between these two vectors, and then subtract this angle from 180° to obtain the actual internal angle. The angles are sorted from largest to smallest, and the first four points are selected as the screening results (key points). These four points are used to construct the quadrilateral.

[0122] S322. Calculate the distance between two adjacent key points, obtain the longest side, and use the longest side as the base side. The base side usually reflects the main direction of a row or column in the storage area, providing a reference for subsequent selection of nearby points.

[0123] S323. Calculate the perpendicular distances from all points in the centroid coordinate point set to the line containing the longest side, and divide the points in the centroid coordinate point set into nearby points and remaining points based on a distance threshold. Specifically, a distance threshold (e.g., 20 units) is set, and all points with a distance less than the threshold are classified as "nearby points," while all other points are classified as "remaining points." Nearby points are generally concentrated within the storage location arrangement area, providing a basis for the subsequent construction of the convex quadrilateral.

[0124] S324. Iteratively construct convex quadrilaterals based on nearby points and remaining points. Each convex quadrilateral represents a storage location layout. Specifically, the steps include the following:

[0125] a. Determine whether the set of remaining points is empty. If so, return the result directly. If not, determine whether the number of nearby points is greater than 1. If so, execute step b. If not, return the result directly.

[0126] b. Sort the nearby points by x-coordinate, and use the distance function to calculate the Euclidean distance between all nearby points and the remaining points to obtain the distance matrix;

[0127] c. Select the remaining unmatched points with the closest Euclidean distance as matching points, sorting them by their x-coordinates. Record the indices of the remaining matched points to ensure each remaining point is selected only once. This results in an array of closest points. Similarly, sort the matched closest points by x-coordinate to ensure consistent vertex order when constructing the quadrilateral.

[0128] d. Sequentially combine two adjacent nearby points with their corresponding two nearest matching points to form a convex quadrilateral. The vertex order of the convex quadrilateral is: the first nearby point, the second nearby point, the second nearest matching point, and the first nearest matching point.

[0129] e. Update the set of nearby points and the set of remaining points, remove the remaining matched points (by determining whether each point coincides with a point in the new set of nearby points within the error range), and repeat steps d until there are no remaining points.

[0130] Through this iterative process, new convex quadrilaterals are continuously constructed using the nearest distance matching at each step until all remaining points are processed or no valid matching can be formed, thus completing the construction of the entire point set into multiple convex quadrilaterals.

[0131] f. Output a list of convex quadrilaterals. Each list item represents a convex quadrilateral, and each convex quadrilateral represents a storage location layout. Each convex quadrilateral is represented by four (x, y) coordinate points, which form the location information of the corresponding storage area. These quadrilaterals comprehensively describe the geometric layout of the entire storage area, which is used for subsequent storage location status detection and warehouse management.

[0132] Figure 6 The MobileViT network structure diagram according to the present invention is shown, combined with reference Figure 1 and Figure 6 The planar storage location status recognition method based on deep learning proposed in the present invention further includes the following steps:

[0133] S4: Acquire the corresponding local image of the storage area in real time according to the storage location layout, and perform status recognition on the local image of the storage area through the image classification algorithm to obtain the storage location status information. This specifically includes the following sub-steps:

[0134] S41. Acquire a local image of the corresponding storage area in real time according to the storage location layout;

[0135] S42, extracting features from the local image of the storage location area through a convolutional layer to obtain an initial storage location feature map;

[0136] S43, performing a downsampling operation on the initial storage location feature map, and performing feature extraction using the MobileViT block to obtain the final storage location feature map;

[0137] S44. Perform state classification on the final storage location feature map through the convolution layer and the linear layer to obtain the storage location status information.

[0138] That is, based on the location information of the storage area, the real-time image from the RGB camera is captured. This image can be used to determine the storage location status through traditional digital image processing or image classification models trained through deep learning. However, deep learning models are suitable for complex environments, variable lighting, different angles, and other situations. They have strong generalization capabilities and can adapt to different scenarios. The MobileViT series model is a lightweight visual transformer designed specifically for mobile terminals and edge devices. It combines the local feature extraction capabilities of CNN and the global modeling capabilities of Transformer. Its main features include:

[0139] Fusion of Convolution and Transformer: MobileViT combines the efficiency of convolutional neural networks (CNN) in local feature extraction and the advantages of Transformer in global dependency modeling to achieve efficient and globally perceptual feature expression.

[0140] Lightweight design: The model has low parameter count and computational complexity, making it suitable for mobile devices and edge computing scenarios, ensuring low latency and high energy efficiency.

[0141] High performance: In tasks such as image classification, object detection, and semantic segmentation, the MobileViT model achieves high speed and low memory usage while maintaining high accuracy.

[0142] Modular architecture: Through the local window self-attention mechanism and cross-local region information fusion, the MobileViT model can effectively capture the details and global information in the image and adapt to the requirements of different visual tasks.

[0143] Figure 7a The captured overall image of the storage area is shown, showing the original state of the storage area, which includes multiple storage areas divided by storage lines. Figure 7b An image containing warehouse angle information is shown. The warehouse angles are identified using the improved YOLOv11 object detection algorithm. These warehouse angles are marked with red boxes containing confidence values. Figure 7c The centroid positioning diagram image is shown, which shows the centroid coordinate point set calculated by the adaptive denoising centroid positioning algorithm. These centroid points are marked with blue circles, accurately indicating the location of the storage location angle. Figure 7d The image of automated storage location area generation is shown. Through the grid-type storage location construction algorithm, the storage location layout is automatically generated according to the centroid coordinate point set. The storage location area is presented in the form of a grid of different colors, clearly showing the arrangement and distribution of the storage locations. Figure 7e The diagram shows the result of the storage location status. The image classification algorithm is used to identify the status of the local image of the storage location area. The results are displayed in different colors or logos, which intuitively reflects the occupied or idle status of each storage location. Figures 7a to 7eIt can be seen that the method of the present invention can effectively identify the storage location angle, accurately calculate the center of mass coordinates, automatically generate the storage location layout, and accurately identify the storage location status, thereby realizing the intelligent management of planar storage locations.

[0144] Further references Figure 8 As an implementation of the above method, the present invention, in its second aspect, presents a structural diagram of an embodiment of a deep learning-based planar storage location status recognition system. This system can be specifically applied to various electronic devices. The system comprises an RGB camera, an AI edge computing box, and a server. The RGB camera captures storage location images, the AI edge computing box processes the images to obtain the storage location area and status, and the server issues task instructions, transmitting storage location status information to related systems or devices for subsequent operations or scheduling.

[0145] During the deployment phase, the server sends a command to the AI edge computing box to obtain an overall image of the storage area. The AI edge computing box then uses multiple connected RGB cameras to capture the overall image of the storage area in real time and processes these images to extract storage location information. Each overall image of the storage area contains multiple storage areas divided by storage location lines.

[0146] The AI edge computing box uses an improved YOLOv11 object detection algorithm to identify the entire image of the storage area and obtain an image containing the storage angle information. The improved YOLOv11 object detection algorithm introduces multi-dimensional dynamic convolution into the feature extraction network, adaptively adjusting the convolution kernel weights to enhance adaptability to image lighting changes and improve detection accuracy and stability.

[0147] The AI edge computing box calculates the centroid coordinates of images containing storage location angle information through an adaptive denoising centroid positioning algorithm, generates a centroid coordinate point set, and inputs the centroid coordinate point set into the grid-type storage location construction algorithm to generate a storage location layout. These centroids represent the four corners of the storage location, and then transmits the storage location area image and its corresponding centroid to the server's host computer.

[0148] On the host computer, staff can add, delete, or modify storage locations. If the storage area is arranged in a grid with the same number of locations in each row, the grid location construction algorithm can automatically generate the location layout. If the storage area is randomly dispersed, the center of mass can be manually selected to generate the location layout, ensuring accuracy and flexibility. In this way, the system can maintain the efficiency of automated processing while providing sufficient room for manual intervention to adapt to different storage area layouts.

[0149] After obtaining the location information (location layout) for each storage area, the server transmits this information to the AI edge computing box. The box captures the storage area image from the RGB camera in real time based on the location information of each storage area and inputs the image into the image classification algorithm to detect the storage area status and obtain the storage area status information. The detection results are then reported to the server, which then implements intelligent inbound and outbound management through the scheduling system.

[0150] The solution of the present invention integrates deep learning and target detection algorithms to realize the automatic extraction of storage position angles and center of mass, thereby avoiding tedious manual external parameter calibration, reducing labor costs and human errors; at the same time, the use of RGB cameras with stronger environmental adaptability and lower cost, combined with an adaptive denoising algorithm, can maintain high recognition accuracy even in complex environments with large lighting changes, high dust content or the presence of highly reflective materials; in addition, the system effectively expands the monitoring range by connecting multiple RGB cameras and performing image integration processing by an AI edge computing box, overcoming the problem of the small monitoring range of single-point lidar and RGB-D cameras; and the use of relatively low-cost RGB cameras instead of expensive lidar and RGB-D cameras, coupled with the architecture of collaborative processing of edge computing and servers, can not only ensure efficient image processing and storage position status recognition, but also greatly reduce hardware investment and power consumption.

[0151] In the third aspect, the present invention proposes a terminal device, comprising a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the computer program, it implements the steps of the planar storage location status identification method based on deep learning as described in any one of the above items.

[0152] In a fourth aspect, the present invention proposes a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps of the planar storage location status recognition method based on deep learning as described in any one of the above items.

[0153] Reference below Figure 9 , which shows a structural diagram of a computer system 900 suitable for implementing an electronic device of an embodiment of the present application. Figure 9 The electronic device shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.

[0154] like Figure 9As shown, the computer system 900 includes a central processing unit (CPU) 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage unit 908 into a random access memory (RAM) 903. Various programs and data required for the operation of the computer system 900 are also stored in the RAM 903. The CPU 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0155] The following components are connected to the I / O interface 905: an input section 906 including a keyboard, a mouse, and the like; an output section 907 including a liquid crystal display (LCD) and a speaker; a storage section 908 including a hard disk and the like; and a communication section 909 including a network interface card such as a LAN card or a modem. The communication section 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the I / O interface 905 as needed. A removable medium 911, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 910 as needed, so that a computer program read therefrom can be installed in the storage section 908 as needed.

[0156] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 909, and / or installed from a removable medium 911. When the computer program is executed by the central processing unit (CPU) 901, the above-mentioned functions defined in the method of the present application are performed.

[0157] It should be noted that the computer-readable storage medium described in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. This propagated data signal can take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable storage medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. Program code contained on a computer-readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wireline, optical cable, RF, or any suitable combination thereof.

[0158] Computer program code for performing the operations of the present application can be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as C or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0159] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.

[0160] The above description is merely a preferred embodiment of the present application and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in this application is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also encompasses other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned inventive concept. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this application.

Claims

1. A planar storage location status recognition method based on deep learning, characterized in that: The following steps are involved: S1. Real-time acquisition of an overall image of a storage area, wherein the overall image of the storage area includes a plurality of storage areas divided by storage lines; S2. Using an improved YOLOv11 target detection algorithm to identify the entire image of the storage area to obtain an image containing storage location angle information, wherein the improved YOLOv11 target detection algorithm introduces multidimensional dynamic convolution into the feature extraction network; S3, calculating the centroid coordinates of the image containing the storage location angle information using an adaptive denoising centroid positioning algorithm to generate a centroid coordinate point set, and inputting the centroid coordinate point set into a grid-type storage location construction algorithm to generate a storage location layout; S4. Acquire the corresponding local image of the storage location area in real time according to the storage location layout, and perform status recognition on the local image of the storage location area through an image classification algorithm to obtain storage location status information.

2. The method for plane storage location status recognition based on deep learning according to claim 1 is characterized in that: In step S2, the improved YOLOv11 target detection algorithm is used to identify the entire image of the storage area to obtain an image containing storage location angle information, which specifically includes the following sub-steps: S21, using a multi-dimensional dynamic convolutional feature extraction network to extract features from the entire image of the storage location area to obtain a first feature map y1, a second feature map y2, and a third feature map y3; S22, inputting the first feature map y1, the second feature map y2, and the third feature map y3 into the feature pyramid network to fuse and exchange information of high-resolution features and low-resolution features, thereby obtaining an enhanced first feature map y′1, an enhanced second feature map y′2, and an enhanced third feature map y′3; S23, inputting the enhanced first feature map y′1, the enhanced second feature map y′2 and the enhanced third feature map y′3 into the detection head for predicting whether each feature point contains a storage position angle and corresponding regression parameters; S24. Decode the prediction result to correspond to the position of the overall image of the storage location area, and perform score sorting and non-maximum suppression screening to obtain the final storage location angle prediction frame, and finally output the image containing the storage location angle information through the prediction frame.

3. The planar storage location status recognition method based on deep learning according to claim 2 is characterized in that: In step S2, the improved YOLOv11 target detection algorithm introduces multi-dimensional dynamic convolution into the feature extraction network, which is expressed as: Where Y (l) Represents the weight of the dynamic convolution kernel; Represents the weight sum of K convolution kernels; * represents the convolution operation; Represents the bias sum of K convolution kernels; π k Represents the attention weight value of the kth convolution kernel; s k represents the attention weight value of the kth convolution kernel position space; X (l-1) Represents the output or input feature map of the previous layer; Represents the weight of the k-th convolution kernel; Represents the bias of the kth convolution kernel; In step S21, multi-dimensional dynamic convolution is used to extract features from the overall image of the storage area to obtain a feature map, which specifically includes the following sub-steps: S211, performing global average pooling and global maximum pooling operations on the entire image of the storage location area to generate a first eigenvector and a second eigenvector respectively; S212: Perform full connection layer and activation function processing on the first eigenvector and the second eigenvector to obtain a first branch output and a second branch output respectively; S213, using the Sigmoid function to normalize the output of the first branch to generate the position space attention weight s k , use the Softmax function to normalize the output of the second branch and generate the channel space attention weight π k ; S214, the position space attention weight s k and channel-space attention weight π k Applied to the corresponding convolution kernels respectively to adjust the weights of the convolution kernels; S215. Use the adjusted convolution kernel to perform convolution operation, batch normalization and activation function processing on the input feature map in sequence to generate a feature map.

4. The method for plane storage location status recognition based on deep learning according to claim 1 is characterized in that: In step S3, the centroid coordinates of the image containing the storage position angle information are calculated using an adaptive denoising centroid positioning algorithm to generate a centroid coordinate point set, which specifically includes the following sub-steps: S311, converting the image containing the storage position angle information into a grayscale image, and performing filtering and denoising processing on the grayscale image; S312, performing adaptive threshold binarization processing on the denoised grayscale image to obtain a binary image, and performing morphological processing on the binary image; S313 , calculating the image moment of the binary image after morphological processing, and calculating the centroid coordinates of the white area according to the image moment to generate a centroid coordinate point set.

5. The method for plane storage location status recognition based on deep learning according to claim 1 is characterized in that: In step S3, the centroid coordinate point set is input into a grid-type storage location construction algorithm to generate a storage location layout, which specifically includes the following sub-steps: S321, using a convex hull algorithm to process the centroid coordinate point set to obtain a point sequence constituting an outer boundary, and screening out key points suitable for constructing a quadrilateral from the point sequence; S322: Calculate the distance between two adjacent key points to obtain the longest side; S323, calculating the vertical distances from all points in the centroid coordinate point set to the straight line on which the longest side is located, and dividing the points in the centroid coordinate point set into nearby points and remaining points according to a distance threshold; S324. Iteratively construct convex quadrilaterals based on the nearby points and the remaining points, each of the convex quadrilaterals representing one of the storage location layouts.

6. The method for plane storage location status recognition based on deep learning according to claim 5 is characterized in that: In step S324, convex quadrilaterals are iteratively constructed based on the nearby points and the remaining points, each of which represents one of the storage location layouts. The steps specifically include the following sub-steps: a. Determine whether the set of remaining points is empty. If so, return the result directly. If not, determine whether the number of nearby points is greater than 1. If so, execute step b. If not, return the result directly. b. Sort the nearby points according to their x-coordinates, and calculate the Euclidean distances between all the nearby points and the remaining points using a distance function to obtain a distance matrix; c. Select the remaining points with the shortest Euclidean distance and not yet matched as matching points in the order of the x-coordinates of the nearby points, and record the indexes of the remaining matched points; d. Sequentially combine two adjacent nearby points with their corresponding two nearest matching points to form a convex quadrilateral, where the vertex order of the convex quadrilateral is: the first nearby point, the second nearby point, the second nearest matching point, and the first nearest matching point; e. Update the set of nearby points and the set of remaining points, remove the remaining matched points, and repeat steps d until there are no remaining points; f. Output a list of convex quadrilaterals, where each convex quadrilateral represents one of the storage location layouts.

7. The method for plane storage location status recognition based on deep learning according to claim 1 is characterized in that: In step S4, a local image of the corresponding storage area is acquired in real time according to the storage location layout, and the status of the local image of the storage area is identified by an image classification algorithm to obtain storage location status information. The step S4 specifically includes the following sub-steps: S41, acquiring a local image of the corresponding storage location area in real time according to the storage location layout; S42, extracting features from the local image of the storage location area through a convolutional layer to obtain an initial storage location feature map; S43, performing a downsampling operation on the initial storage location feature map, and performing feature extraction using a MobileViT block to obtain a final storage location feature map; S44. Perform state classification on the final storage location feature map through a convolutional layer and a linear layer to obtain storage location state information.

8. A plane storage location status recognition system based on deep learning, characterized in that: The system includes: an RGB camera, an AI edge computing box, and a server; The server sends an instruction to the AI edge computing box to obtain an overall image of the storage area. The AI edge computing box collects the overall image of the storage area in real time through the multiple connected RGB cameras. The overall image of the storage area includes multiple storage areas divided by storage lines. The AI edge computing box uses an improved YOLOv11 target detection algorithm to identify the entire image of the storage area to obtain an image containing storage location angle information, wherein the improved YOLOv11 target detection algorithm introduces multi-dimensional dynamic convolution into the feature extraction network; The AI edge computing box calculates the centroid coordinates of the image containing the storage location angle information using an adaptive denoising centroid positioning algorithm to generate a centroid coordinate point set, and inputs the centroid coordinate point set into a grid-type storage location construction algorithm to generate a storage location layout; The AI edge computing box obtains the corresponding local image of the storage area in real time according to the storage location layout, and performs status recognition on the local image of the storage area through an image classification algorithm to obtain storage location status information.

9. A terminal device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the planar storage location status recognition method based on deep learning as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the planar storage location status recognition method based on deep learning as described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Storage location state identification method based on target detection and intersection-union ratio

    CN116824264A

  • Contextual visual-based SAR target detection method and apparatus, and storage medium

    US20230184927A1

  • Method for detecting image target in smart home environment

    WO2021244079A1