Storage-based processing method and device, electronic equipment and medium
By rotating the object detection model and feature refining module, the shortcomings of the YOLO model in the location detection are solved, and the accurate detection of obliquely placed or deformed objects is achieved, which improves the accuracy of the location detection and the efficiency of handling equipment scheduling.
Patent Information
- Application Number
- CN202311845122.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-28
- Publication Date
- 2025-07-01
AI Technical Summary
In the prior art, the horizontal detection frame of the YOLO model cannot accurately detect obliquely placed or deformed objects, resulting in inaccurate detection of the warehouse location and affecting the dispatch of the handling equipment.
The rotating object detection model is adopted to detect objects through a rotatable detection box, combining feature refining modules and cross-dimensional interactive attention modules to improve detection accuracy.
It improves the accuracy of warehouse location detection and the efficiency of handling equipment scheduling, can identify non-level placement targets, and enhances production safety monitoring and equipment avoidance capabilities.
Smart Images

Figure CN120236052A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent warehousing, and particularly to a processing method, device, electronic device and medium based on warehousing. Background Art
[0002] In the scenario of intelligent warehousing, handling equipment is used to perform unmanned handling tasks, and the scheduling of the handling equipment depends on relevant information such as storage locations in the warehouse area, such as scheduling the handling equipment according to information such as which storage locations have goods and which storage locations are idle.
[0003] In the prior art, the YOLO (You Only Look Once) model is usually used for storage location detection. However, the detection frames of the YOLO model are set horizontally. For objects placed obliquely or with deformed images, it is impossible to accurately frame them with horizontal detection frames, and there will be more background information other than the object itself, resulting in low accuracy of storage location detection, and thus affecting the scheduling of handling equipment. Summary of the Invention
[0004] In view of the above problems, a processing method, device, electronic device and medium based on warehousing are provided to overcome or at least partially solve the above problems, including:
[0005] A processing method based on warehousing, the method includes:
[0006] Obtain a video stream of a warehousing area collected by a camera;
[0007] Adopt a rotated object detection model to detect the video stream to obtain a detection result for the warehousing area; wherein, the rotated object detection model performs object detection through rotatable detection frames, and the detection result includes storage location related information;
[0008] When a handling task is received, determine a target storage location in the warehousing area according to the storage location related information, so as to schedule the handling equipment to execute the handling task at the target storage location.
[0009] Optionally, the detection result further includes abnormal situation information, and further includes:
[0010] Generate a warning message and give feedback according to the abnormal situation information.
[0011] Optionally, further includes:
[0012] Control the handling equipment to avoid according to the abnormal situation information.
[0013] Optionally, the abnormal situation information includes any one or more of the following:
[0014] Abnormal situation information endangering production safety and abnormal situation information affecting the movement of handling equipment.
[0015] Optionally, it further includes:
[0016] Conduct inventory location quantity statistics based on the inventory location related information to obtain the inventory location quantity statistics result;
[0017] The determining the target inventory location in the storage area according to the inventory location related information includes:
[0018] Determine the target inventory location in the storage area according to the inventory location related information, or according to the inventory location related information and the inventory location quantity statistics result.
[0019] Optionally, the inventory location quantity statistics result includes: the quantity of inventory locations in the stocked state and the quantity of inventory location data in the idle state.
[0020] Optionally, the using the rotation target detection model to detect the video stream to obtain the detection result for the storage area includes:
[0021] Obtain an image frame from the video stream;
[0022] In the rotation target detection model, generate a feature map according to the image frame and reconstruct the feature map to obtain the reconstructed feature map;
[0023] Determine the detection result for the storage area according to the reconstructed feature map.
[0024] Optionally, the video stream includes multiple video streams collected by multiple cameras, and the obtaining an image frame from the video stream includes:
[0025] Obtain local image frames from the multiple video streams;
[0026] Stitch the local image frames in the multiple video streams to obtain a global image frame.
[0027] Optionally, the reconstructing the feature map to obtain the reconstructed feature map includes:
[0028] Obtain a new feature vector of a feature point in the feature map through a rotatable detection frame, and use the new feature vector to replace the previous feature vector of the feature point;
[0029] Traverse the feature points in the feature map to replace the feature vectors to obtain the reconstructed feature map.
[0030] Optionally, before using the new feature vector to replace the previous feature vector of the feature point, it further includes:
[0031] Update the new feature vector through bilinear interpolation.
[0032] Optionally, the rotation target detection model embeds a cross-dimensional interaction attention module, which is used to obtain the features of the fusion of the image spatial dimension and the image channel dimension in the feature map.
[0033] Optionally, the image spatial dimension includes the image width dimension and the image height dimension, and the image channel dimension includes the three primary color channel dimensions of the image and the filter channel dimension.
[0034] Optionally, the rotation target detection model has a feature refinement module, which has a dual-path convolution part and a feature refinement part, and the cross-dimensional interaction attention module is embedded in the dual-path convolution part and the feature refinement part.
[0035] Optionally, the cross-dimensional interaction attention module includes at least three branches, including:
[0036] The first branch is used to obtain the features of the image spatial dimension;
[0037] The second branch is used to obtain the features of the fusion of the image channel dimension and the first image spatial dimension;
[0038] The third branch is used to obtain the features of the fusion of the channel dimension information and the second image spatial dimension;
[0039] Wherein, one of the first image spatial dimension and the second image spatial dimension is the image height dimension, and the other is the image width dimension.
[0040] Optionally, the cross-dimensional interaction attention module is also used to: fuse the features obtained by the at least three branches.
[0041] Optionally, during the training of the rotation target detection model, an inclined intersection over union score loss function is adopted, and the loss value of the inclined intersection over union score loss function is determined according to the intersection over union score loss, the shape loss, and the distance loss.
[0042] Optionally, the intersection over union score loss is determined according to the intersection and union of the ground truth detection box and the predicted detection box;
[0043] The shape loss is determined according to the width and height of the ground truth detection box and the predicted detection box;
[0044] The distance loss is determined based on the width, height, and angle loss of the minimum bounding rectangle of the ground truth detection box and the predicted detection box, and the angle loss is determined based on the height difference between the center points of the ground truth detection box and the predicted detection box and the distance between the center points of the ground truth detection box and the predicted detection box.
[0045] Optionally, the bin-related information includes any one or more of the following:
[0046] Bin location information, bin status information.
[0047] Optionally, the bin status information includes an in-stock status or an idle status.
[0048] A warehousing-based processing device, the device is used for:
[0049] Obtain the video stream of the warehousing area collected by the camera;
[0050] Use a rotated object detection model to detect the video stream to obtain the detection result for the warehousing area; wherein, the rotated object detection model performs object detection through a rotatable detection box, and the detection result includes bin-related information;
[0051] When receiving a handling task, determine the target bin in the warehousing area according to the bin-related information, so as to schedule the handling equipment to execute the handling task at the target bin.
[0052] An electronic device includes a processor, a memory, and a computer program stored on the memory and capable of running on the processor. When the computer program is executed by the processor, the warehousing-based processing method described above is implemented.
[0053] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the warehousing-based processing method described above is implemented.
[0054] The embodiments of the present invention have the following advantages:
[0055] In the embodiments of the present invention, by obtaining the video stream of the warehousing area collected by the camera, using a rotated object detection model to detect the video stream to obtain the detection result for the warehousing area, the rotated object detection model performs object detection through a rotatable detection box, and the detection result includes bin-related information. When receiving a handling task, determine the target bin in the warehousing area according to the bin-related information, so as to schedule the handling equipment to execute the handling task at the target bin, realizing the detection of bins using a rotated object detection model, and the rotated object detection model performs bin detection through a rotatable detection box, improving the accuracy of bin detection, and further improving the accuracy and efficiency of scheduling the handling equipment. Brief Description of the Drawings
[0056] To more clearly illustrate the technical solutions of the present invention, the drawings required for the description of the present invention will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0057] Figure 1 is a flowchart of the steps of a processing method based on warehousing provided by some embodiments of the present invention;
[0058] Figure 2a is a schematic diagram of a system architecture provided by some embodiments of the present invention;
[0059] Figure 2b is a schematic diagram of the detection of a horizontal detection frame provided by some embodiments of the present invention;
[0060] Figure 2c is a schematic diagram of the detection of a rotation detection frame provided by some embodiments of the present invention;
[0061] Figure 2d is a schematic diagram of the IoU score loss provided by some embodiments of the present invention;
[0062] Figure 2e is a schematic diagram of the distance loss provided by some embodiments of the present invention;
[0063] Figure 2f is a schematic diagram of the angle loss provided by some embodiments of the present invention;
[0064] Figure 2g is a schematic diagram of a rotation target detection model provided by some embodiments of the present invention;
[0065] Figure 2h is a schematic diagram of another rotation target detection model provided by some embodiments of the present invention;
[0066] Figure 2i is a schematic diagram of a cross-dimensional interaction attention module provided by some embodiments of the present invention;
[0067] Figure 3 is a schematic diagram of another system architecture provided by some embodiments of the present invention. Detailed Description of the Embodiments
[0068] To make the above objects, features, and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0069] Referring to Figure 1 , a step flowchart of a warehousing-based processing method provided by some embodiments of the present invention is shown. This method can be applied to a server and specifically may include the following steps:
[0070] Step 101, obtain a video stream of the warehousing area collected by a camera.
[0071] In some examples, the camera can be a surveillance camera deployed in the warehousing area. For example, Figure 2a , the video stream collected by the camera can be pulled through the Real Time Streaming Protocol (RSTP) and transmitted to the server through a switch.
[0072] In some examples, multiple cameras are set in the warehousing area. Then, the image data in the video streams collected by the multiple cameras can be globally stitched, and then the globally stitched image data can be detected, or the image data in the video stream collected by each camera can be detected separately.
[0073] Step 102, use a rotated object detection model to detect the video stream to obtain a detection result for the warehousing area; wherein, the rotated object detection model performs object detection through a rotatable detection frame, and the detection result includes information related to storage locations.
[0074] In some examples, the detection frame in the rotated object detection model can be rotated to better fit the actual edges of the object. For example, the rotated object detection model can be the R3Det (Refined single-stage detector with feature refinement for rotating object) model, or other models.
[0075] After obtaining the video stream, the image data in the video stream can be input into a rotated object detection model. The rotated object detection model can perform object detection on the image data through a rotatable detection box. By setting the rotatable detection box, the detection box can be rotated according to the placement position of the object, so that the detection box can better frame the object, and there is less background information in the detection box, resulting in more accurate positioning. For example Figure 2b is a schematic diagram of using a horizontal detection box for object detection. For example Figure 2c is a schematic diagram of using a rotated detection box for object detection. It can be seen that using a rotated detection box has a better effect.
[0076] In the related art, a rotated box angle calculation branch can be added to the head part of the related model, and then rotated object box detection can be realized. However, the rotated box angle calculation branch only roughly calculates the corresponding angle change and still cannot better fit the object edge.
[0077] In the embodiments of the present invention, through a rotated object detection model (such as the R3Det model), the current bounding box information is re-encoded by using a feature refinement module, and feature alignment is achieved by reconstructing the entire feature map, which can better solve the above problems. The rotated detection box can better frame the storage location, and there is less background information in the detection box, resulting in more accurate positioning.
[0078] In some examples, the detection result may include storage location related information. The storage location related information includes any one or more of the following: storage location position information, storage location status information.
[0079] Among them, the storage location position information may be the coordinate information of the storage location, and the storage location status information may be the status of the goods in the storage location, that is, the stocked status or the out-of-stock (idle) status.
[0080] In some embodiments of the present invention, the detection result may further include abnormal situation information. The abnormal situation information may include any one or more of the following:
[0081] Abnormal situation information endangering production safety, abnormal situation information affecting the movement of handling equipment.
[0082] In some embodiments of the present invention, it further includes:
[0083] Generating a warning message and giving feedback according to the abnormal situation information.
[0084] In some embodiments of the present invention, it further includes:
[0085] Controlling the handling equipment to avoid according to the abnormal situation information.
[0086] During the warehousing production process, production safety is essential. However, the relevant storage location detection only focuses on simple storage location detection and lacks safety detection for abnormal situations, such as categories like fire, smoke, and people, which poses certain potential production safety hazards.
[0087] Based on this, in addition to detecting storage locations, the embodiments of the present invention can also detect abnormal situations that endanger production safety, such as fire and smoke, in the warehousing area. Considering that the warehousing area is an unmanned operation mode and the presence of personnel in the warehousing area may affect the movement of handling equipment, the present invention can also detect abnormal situations that affect the movement of handling equipment, such as personnel, in the warehousing area.
[0088] After obtaining the abnormal situation information, a warning message can be generated and transmitted to the operator and the intelligent handling equipment that is performing tasks through the server, enabling the timely discovery of abnormal situations, minimizing safety damage, and ensuring overall life and property safety.
[0089] In some examples, the warning message can be sent to relevant personnel in a timely manner through a program, and the relevant warning message can be output through the audio output interface (playing audio) in the camera, enabling relevant personnel to carry out rescue in a timely manner and ensuring production safety.
[0090] In some examples, when an abnormal situation where there are personnel in the warehousing area is detected during the operation of special vehicles such as handling equipment, a warning message will be sent to the operator of the handling equipment through a program (a warning message is generated when there are personnel on the planned driving path of the handling equipment), for timely warning and suspension of relevant operations until the safety risk is eliminated.
[0091] In some examples, since the abnormal situation information includes abnormal situation information that affects the movement of handling equipment, when the abnormal situation information that affects the movement of handling equipment is detected, the movement of the handling equipment can be controlled to avoid it. For example, when it is detected that there are personnel in the warehousing area, the handling equipment should be controlled to avoid the personnel.
[0092] In some examples, by collecting sample image information and recording the storage location-related information and abnormal situation information of the sample image data, label annotation is performed through annotation software. After annotation, it is sent to a rotated object detection model for a certain degree of training. According to the requirements of the training result accuracy and real-time performance, a suitable rotated object detection model is selected and deployed to the server for real-time rotated object detection.
[0093] During the training of the rotating object detection model, the Skew Intersection OverUnion (SkewIOU) function is adopted. SkewIOU is sensitive to angles, and a slight deviation will cause a rapid decline in SkewIOU. If there is a slight jitter in the predicted detection box during training, it will lead to improper results of SkewIOU, resulting in continuous oscillation of the network and inability to converge.
[0094] Based on this, the vector angle between the ground truth box and the predicted box is introduced, and the relevant loss function is redefined. The loss value of the Skew Intersection OverUnion loss function is determined according to the Intersection over Union cost (IOU cost), Shape cost (Shape cost), and Distance cost (Distance cost), as shown in the following formula:
[0095]
[0096] where L SIoU is the Skew Intersection OverUnion loss, IoU is the Intersection over Union cost, Δ is the Distance cost, and Ω is the Shape cost.
[0097] In the embodiments of the present invention, the vector angle between the ground truth box and the predicted box is further considered through the Skew Intersection OverUnion loss function. The method of considering the matching direction of the Skew Intersection OverUnion loss function greatly helps the training convergence process and effect, because it can quickly move the predicted box to the nearest axis. Adding the angle penalty cost effectively reduces the total degrees of freedom of the loss, and the subsequent method only requires the regression of one coordinate X or Y.
[0098] In some embodiments of the present invention, the Intersection over Union cost is determined according to the intersection and union of the ground truth detection box and the predicted detection box. For example, Figure 2d , the intersection A of the ground truth box and the predicted box and the union B of the ground truth box and the predicted box can be obtained, and the Intersection over Union cost (IOU) can be calculated using the following formula:
[0099]
[0100] In some embodiments of the present invention, the Shape cost is determined according to the width and height of the ground truth detection box and the predicted detection box. The Shape cost Ω can be calculated using the following formula:
[0101]
[0102] where (w, h), (w gt , h gtare the widths and heights of the predicted bounding box and the ground truth bounding box respectively, and θ is used to control the attention to the shape loss. To avoid over-focusing on the shape loss and reducing the movement of the predicted bounding box.
[0103] In some embodiments of the present invention, such as Figure 2e , the distance loss is determined according to the width, height, and angle cost of the minimum enclosing rectangle of the ground truth detection box and the predicted detection box, and can be expressed by the following formula:
[0104]
[0105] where Δ is the distance loss, (c w , c h ) are the width and height of the minimum enclosing rectangle of the ground truth box (B in the figure) and the predicted box (B GT ) respectively, and ρ x can be determined according to the angle loss, and can be expressed by the following formula:
[0106]
[0107] where Λ is the angle loss.
[0108] In some embodiments of the present invention, such as Figure 2f , the angle loss is determined according to the height difference between the centers of the ground truth detection box and the predicted detection box, and the distance between the centers of the ground truth detection box and the predicted detection box, and the angle loss Λ can be determined by the following formula:
[0109]
[0110] where c h is the height difference between the centers of the ground truth box (B in the figure) and the predicted box (B GT ), and σ is the distance between the centers of the ground truth box and the predicted box.
[0111] In some embodiments of the present invention, the rotation object detection model is used to detect the video stream to obtain the detection result for the storage area, including:
[0112] Obtaining an image frame from the video stream; in the rotation object detection model, generating a feature map according to the image frame, and reconstructing the feature map to obtain a reconstructed feature map; determining the detection result for the storage area according to the reconstructed feature map.
[0113] In practical applications, an image frame in a video stream can be input into a rotated object detection model. In the rotated object detection model, a feature map can be generated based on the image frame, and then the feature map can be reconstructed to obtain a reconstructed feature map. After obtaining the reconstructed feature map, object detection is performed based on the reconstructed feature map to determine the detection result for the warehousing area.
[0114] In some examples, some rotated object detectors still use the same feature map to perform multiple classifications and regressions (that is, the final detection (classification + regression) results are output without refining the feature map), without considering the feature misalignment caused by the change in the position of the bounding box (the center point of the horizontal detection box, the rotated detection box after rotation and width / height adjustment is not aligned with the center point of the original horizontal detection box). In the rotated object detection model (such as R3Det) in the embodiments of the present invention, the feature refinement module re-encodes the current bounding box information and realizes feature alignment by reconstructing the entire feature map (which can make the detection box more fitting to the real object in terms of angle, length, width, etc., and align the spatial information and channel information of the image).
[0115] Such as Figure 2g , in the feature refinement (Feature Refinement) module of the rotated object detection model, it may include a two-way convolution part (such as Figure 2g Conv1x1, Conv5x1, and Conv1x5 in
[0116] ), and also includes Refined Bbox (refinement part) and Bbox Filtering (filtering part). Rotatable detection boxes will be used to reconstruct the features output by the two-way convolution part and the features output by Refined Bbox and Bbox Filtering to obtain a reconstructed feature map (RefinedFeature Maps).
[0117] Specifically, new features are obtained by superimposing feature maps (the backbone network processes the input image frame to obtain the feature map) through two-way convolution. In the refinement stage, only the detection box with the highest score for each feature point is retained (during the program detection, false detections are inevitable, and there will be detection boxes for the false detection parts, so they need to be screened according to the highest confidence score criterion) to improve the speed, and at the same time, it is ensured that each feature point only corresponds to one refined detection box.
[0118] In some embodiments of the present invention, the reconstructing the feature map to obtain a reconstructed feature map includes:
[0119] Obtain a new feature vector of a feature point in the feature map through a rotatable detection frame, and use the new feature vector to replace the prior feature vector of the feature point; traverse the feature points in the feature map to replace the feature vectors, and obtain a reconstructed feature map.
[0120] In some embodiments of the present invention, before using the new feature vector to replace the prior feature vector of the feature point, it further includes:
[0121] Update the new feature vector through bilinear interpolation.
[0122] For each feature point of the feature map, according to 5 coordinates in the refined detection frame (one center point and four corner points of the detection frame), obtain the corresponding feature vector (i.e., the new feature vector) on the feature map. In some examples, more accurate feature vectors can be obtained through bilinear interpolation. Then, 5 feature vectors (i.e., the new feature vectors) can be added to replace the previous feature vectors (i.e., the prior feature vectors). After traversing all feature points, the entire feature map will be reconstructed, and then the reconstructed feature map can be added to the original feature map to complete the entire process.
[0123] In some embodiments of the present invention, multiple cameras are deployed in the storage area, and the video stream includes multiple video streams collected by multiple cameras. Obtaining an image frame from the video stream includes:
[0124] Obtain local image frames from the multiple video streams; splice the local image frames in the multiple video streams to obtain a global image frame.
[0125] In practical applications, local image frames can be extracted from each video stream, then the local image frames in the multiple video streams are spliced to obtain a global image frame, and then the global image frame can be input into a rotation target detection model for detection.
[0126] In some embodiments of the present invention, when using a rotation target detection model, the accurate rotation angle of the detection frame becomes very important. Then, a lightweight cross-dimensional interaction attention (TripletAttention, TA) module can be embedded in the rotation target detection model, and the cross-dimensional interaction attention module can be used to obtain the features of the fusion of the image spatial dimension and the image channel dimension in the feature map.
[0127] In some examples, the image spatial dimension may include the image width dimension and the image height dimension, and the image channel dimension may include the three primary color channels dimension of the image and the filter channel dimension.
[0128] For other attention modules, such as SE, CBAM, CA, AFF, etc., they are mainly divided into channel attention modules and spatial attention modules. They all focus on a single direction and do not achieve the interaction and fusion of spatial information and channel information. For example, SE only focuses on channel information, or although CBAM pays attention to both channel information and spatial information, it first focuses on channel information and then on spatial information, there is a sequential problem and there is no interaction operation. In the embodiments of the present invention, the cross-dimensional interaction attention module interacts and fuses the information of channel information and spatial information, enabling the model to better focus on the detailed information of the object to be detected, and can obtain relevant improvements with extremely few parameters, and the cross-dimensional interaction attention module has better metrics.
[0129] During the process of reconstructing the feature map, the cross-dimensional interaction attention module is embedded to obtain a more accurate detection box position and angle. Specifically, the rotation target detection model has a feature refinement module, and the feature refinement module has a dual-path convolution part and a feature refinement part. The cross-dimensional interaction attention module is embedded in the dual-path convolution part and the feature refinement part. As Figure 2h , the cross-dimensional interaction attention module is added to the dual-path convolution part of the feature refinement module, the Refined Bbox (i.e., the feature refinement part), and the Bbox Filtering part. Adding the cross-dimensional interaction attention module to the Refined Bbox can obtain more refined results for the 5 coordinates of the refined bounding box, and the subsequent output added to the original feature map can obtain more refined results. Adding the cross-dimensional interaction attention module after the process of superimposing feature maps in the dual-path convolution can pay more attention to more detailed features such as small targets in the image, and improve the fitting of the subsequent reconstructed feature map added to the original feature map.
[0130] For the cross-dimensional interaction attention module, it can interact and fuse the information of channel information and spatial information, enabling the model to better focus on the detailed information of the object to be detected, and can obtain relevant improvements with extremely few parameters.
[0131] In some embodiments of the present invention, the cross-dimensional interaction attention module may include at least three branches, including: a first branch, a second branch, and a third branch. Among them, two branches (the second branch and the third branch) are respectively used to capture the cross-channel interaction between the channel C dimension (the three primary color channels of the RGB image and subsequent filter channels) and the spatial dimension W (the width dimension of the image) / H (the height dimension of the image), and the remaining one branch (the first score) is the calculation of the spatial attention weight, as Figure 2i , from left to right are the first branch, the second branch, and the third branch.
[0132] The first branch is used to obtain the features of the spatial dimension of the image;
[0133] A second branch for obtaining features of the fused image channel dimension and the first image spatial dimension;
[0134] A third branch for obtaining features of the fused channel dimension information and the second image spatial dimension;
[0135] Wherein, one of the first image spatial dimension and the second image spatial dimension is the image height dimension, and the other is the image width dimension. For example, Figure 2i in, the first image spatial dimension in the second branch (the second branch from left to right) is the image width dimension, and the second image spatial dimension in the third branch (the second branch from left to right) is the image height dimension.
[0136] For the first branch, the channel attention calculation branch, the input feature passes through Z-Pool, then followed by a 7x7 convolution, and finally a Sigmoid activation function to generate the spatial attention weight, as shown in the following formula:
[0137] TA1(x) = x · [δ(BN(Conv 7x7 (Z-Pool(x))))]
[0138] Wherein, x refers to the input image information, δ represents the Sigmoid operation, BN represents the normalization operation, and Conv7 represents the 7x7 convolution operation.
[0139] In some examples, Z-Pool can be used to compress dimension information. The specific operation is to perform MaxPooling (maximum pooling layer) and AvgPooling (average pooling layer) on the input, and output a 2xHxW feature, as shown in the following formula:
[0140] Z-Pool(x) = [MaxPool 0d (x), AvgPool 0d (x)]
[0141] For the second branch, the channel C and spatial dimension W dimension interaction capture branch, the input feature first passes through permute (transpose operation, which is convenient for code calculation and model interaction and fusion, so as to adjust the input format from CxHxW to HxCxW), and becomes a HxCxW dimension feature. Then, Z-Pool is performed on the H dimension, and the subsequent operations are similar. Finally, it needs to be transformed into a CxHxW dimension feature through the permuter operation, which is convenient for subsequent element-wise (adding the relevant feature values at the corresponding positions of the image pixels) addition, as shown in the following formula:
[0142] TA2(x) = Z-Pool(x · [δ(BN(Conv 7x7(Z-Pool(x))))])
[0143] Among them, δ represents performing the Sigmoid operation, BN represents performing the normalization operation, and Conv7 represents performing the 7x7 convolution operation.
[0144] For the third branch, the cross-dimensional interaction attention module for capturing the interaction between the channel C and the spatial H dimensions. The input feature first undergoes a permute to become a feature of the WxHxC dimension, then a Z-Pool is performed in the W dimension, and the subsequent operations are similar. Finally, it needs to undergo a permute to become a feature of the CxHxW dimension for convenient element-wise addition later, as shown in the following formula:
[0145] TA3(x) = Z-Pool(x · [δ(BN(Conv 7x7 (Z-Pool(x))))])
[0146] In some embodiments of the present invention, the cross-dimensional interaction attention module can also be used to: fuse the features obtained from the at least three branches.
[0147] Finally, the output features of the above three branches are added and averaged Avg. After passing through this module, the important channel and spatial information of the image will be relatively more obvious than before processing and will be more easily noticed by the model, thereby improving the model detection accuracy, as shown in the following formula:
[0148]
[0149] Step 103, when receiving a handling task, determine the target storage location in the storage area according to the storage location related information, so as to schedule the handling device to execute the handling task at the target storage location.
[0150] In some examples, the handling device can be an Automated Guided Vehicle (AGV), and the handling device can also be other unmanned operation devices.
[0151] After obtaining the detection result, the detection result can be stored on the server (when there is a new detection result, the new detection result replaces the old detection result). In some examples, different information is stored using different parameters for convenient subsequent calling of the corresponding interfaces and display on the operation interface.
[0152] In practical applications, the detection results stored on the server can represent the real-time information of goods. When receiving a handling task for the goods, such as moving the goods from outside the storage area to an idle storage location, moving the goods in the storage location to outside the storage area, or moving the goods in a certain storage location to another storage location, the optimal algorithm can be used to screen based on factors such as the length of the path and the size of the goods according to the detection results stored on the server. Finally, a suitable target storage location is selected for the scheduling operation of the relevant handling equipment, and then the handling equipment can be scheduled to execute the handling task at the target storage location.
[0153] In some examples, the above determination of the target storage location and the scheduling operation can be performed by a scheduling system.
[0154] In some embodiments of the present invention, it further includes:
[0155] Count the number of storage locations according to the storage location-related information to obtain the storage location quantity statistics result.
[0156] The determining the target storage location in the storage area according to the storage location-related information includes:
[0157] Determine the target storage location in the storage area according to the storage location-related information, or the storage location-related information and the storage location quantity statistics result.
[0158] In some examples, the storage location quantity statistics result may include: the number of storage locations in the stocked state and the number of storage location data in the idle state.
[0159] In practical applications, the storage location-related information in the detection results only contains status information or location information or both, lacking relevant statistics on the quantity information of the storage locations and not being fully utilized.
[0160] Based on this, in the embodiments of the present invention, the number of storage locations can be counted according to the storage location-related information to obtain the storage location quantity statistics result. For example, the storage location quantity statistics result includes the number of stocked storage locations and the number of idle storage locations, and is stored in the server to provide an interface for the subsequent query and use of the scheduling system, that is, the target storage location can be determined by combining the storage location-related information and the storage location quantity statistics result.
[0161] In an embodiment of the present invention, by acquiring a video stream of a warehousing area collected by a camera, and using a rotated object detection model to detect the video stream, a detection result of the warehousing area is obtained. The rotated object detection model performs object detection through a rotatable detection frame, and the detection result includes information related to storage locations. When a handling task is received, the target storage location in the warehousing area is determined according to the information related to the storage location, so as to schedule a handling device to perform a handling task at the target storage location, realizing the detection of storage locations by using a rotated object detection model. The rotated object detection model performs storage location detection through a rotatable detection frame, improving the accuracy of storage location detection, and further improving the accuracy and efficiency of scheduling the handling device.
[0162] Specifically, the beneficial effects of the improved rotated object detection model are as follows:
[0163] 1. Insert a lightweight TA attention module into the feature refinement module to refine the 5 coordinate parameters of the rotated object frame. With the cost of adding a very small number of parameters, the accuracy of the rotated object frame fitting the GT frame is improved.
[0164] 2. Use the SIoU loss function calculation method to replace the SkewIOU loss function, which can effectively solve problems such as slow model convergence caused by angle parameter deviation. The method of considering the matching direction in the SIoU loss function greatly helps the training convergence process and effect.
[0165] Generally speaking, the improved rotated object detection model can effectively identify storage location targets that are not horizontally placed and have a small resolution, with high recognition accuracy and strong robustness.
[0166] The following is an exemplary description of the present invention in conjunction with Figure 3 :
[0167] For example, Figure 3 , the video stream of the warehousing area is collected by multiple cameras, and then the video stream is sent to a server through a switch. The server can detect the video stream through a rotated object detection model to obtain retrieval results, including information related to storage locations and abnormal situation information, and can calculate the statistical result of the number of storage locations, and store this information in the server. When a handling task is received, the scheduling system can generate a scheduling instruction according to the information stored in the server to realize the scheduling of the handling device.
[0168] Among them, the information related to the storage location includes the storage location information (such as Figure 3 the location of storage location No. 1 is (110, 234), the location of storage location No. 2 is (324, 234), and the location of storage location No. 6 is (538, 468)), the storage location status information (such as Figure 3 storage locations No. 1, No. 2, and No. 6 are in the stocked state, and other storage locations are in the idle state), and the abnormal situation information may include personnel information in the warehousing area (such asFigure 3 Among the abnormal situations, there is someone), and the statistical result of the bin quantity can include the bin quantities in different states (such as Figure 3 the in-stock quantity is 3 and the idle quantity is 3).
[0169] It should be noted that for the method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of the present invention are not limited by the described action sequence, because according to the embodiments of the present invention, certain steps can be carried out in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential for the embodiments of the present invention.
[0170] A processing device based on warehousing provided by some embodiments of the present invention, the device is used for:
[0171] Obtain the video stream for the warehousing area collected by the camera;
[0172] Adopt a rotated object detection model to detect the video stream, and obtain the detection result for the warehousing area; wherein, the rotated object detection model performs object detection through a rotatable detection frame, and the detection result includes bin-related information;
[0173] When receiving a handling task, determine the target bin in the warehousing area according to the bin-related information, so as to dispatch the handling device to execute the handling task at the target bin.
[0174] In some embodiments of the present invention, the detection result further includes abnormal situation information, and the device is further used for:
[0175] Generate a warning message and give feedback according to the abnormal situation information.
[0176] In some embodiments of the present invention, the device is further used for:
[0177] Control the handling device to avoid according to the abnormal situation information.
[0178] In some embodiments of the present invention, the abnormal situation information includes any one or more of the following:
[0179] Abnormal situation information endangering production safety, abnormal situation information affecting the movement of the handling device.
[0180] In some embodiments of the present invention, it is further used for:
[0181] Conduct bin quantity statistics according to the bin-related information, and obtain the bin quantity statistical result;
[0182] Determining the target storage location in the storage area according to the storage location related information includes:
[0183] Determining the target storage location in the storage area according to the storage location related information, or according to the storage location related information and the statistical result of the storage location quantity.
[0184] In some embodiments of the present invention, the statistical result of the storage location quantity includes: the quantity of storage locations in the in-stock state and the quantity of data of storage locations in the idle state.
[0185] In some embodiments of the present invention, using the rotation target detection model to detect the video stream to obtain the detection result for the storage area includes:
[0186] Obtaining an image frame from the video stream;
[0187] In the rotation target detection model, generating a feature map according to the image frame and reconstructing the feature map to obtain a reconstructed feature map;
[0188] Determining the detection result for the storage area according to the reconstructed feature map.
[0189] In some embodiments of the present invention, the video stream includes multiple video streams collected by multiple cameras, and obtaining an image frame from the video stream includes:
[0190] Obtaining local image frames from the multiple video streams;
[0191] Stitching the local image frames in the multiple video streams to obtain a global image frame.
[0192] In some embodiments of the present invention, reconstructing the feature map to obtain a reconstructed feature map includes:
[0193] Obtaining a new feature vector of a feature point in the feature map through a rotatable detection frame, and using the new feature vector to replace the prior feature vector of the feature point;
[0194] Traversing the feature points in the feature map to replace the feature vectors to obtain a reconstructed feature map.
[0195] In some embodiments of the present invention, before using the new feature vector to replace the prior feature vector of the feature point, the device is further configured to:
[0196] Updating the new feature vector through bilinear interpolation.
[0197] In some embodiments of the present invention, the rotation target detection model is embedded with a cross-dimensional interaction attention module, and the cross-dimensional interaction attention module is used to obtain the features of the fusion of the image spatial dimension and the image channel dimension in the feature map.
[0198] In some embodiments of the present invention, the image spatial dimension includes the image width dimension and the image height dimension, and the image channel dimension includes the three primary color channel dimensions of the image and the filter channel dimension.
[0199] In some embodiments of the present invention, the rotation target detection model has a feature refinement module, the feature refinement module has a dual-path convolution part and a feature refinement part, and the cross-dimensional interaction attention module is embedded in the dual-path convolution part and the feature refinement part.
[0200] In some embodiments of the present invention, the cross-dimensional interaction attention module includes at least three branches, where:
[0201] The first branch is used to obtain the features of the image spatial dimension;
[0202] The second branch is used to obtain the features of the fusion of the image channel dimension and the first image spatial dimension;
[0203] The third branch is used to obtain the features of the fusion of the channel dimension information and the second image spatial dimension;
[0204] Wherein, one of the first image spatial dimension and the second image spatial dimension is the image height dimension, and the other is the image width dimension.
[0205] In some embodiments of the present invention, the cross-dimensional interaction attention module is further used to: fuse the features obtained by the at least three branches.
[0206] In some embodiments of the present invention, during the training of the rotation target detection model, an inclined intersection over union score loss function is adopted, and the loss value of the inclined intersection over union score loss function is determined according to the intersection over union score loss, the shape loss, and the distance loss.
[0207] In some embodiments of the present invention, the intersection over union score loss is determined according to the intersection and union of the true detection box and the predicted detection box;
[0208] The shape loss is determined according to the width and height of the true detection box and the predicted detection box;
[0209] The distance loss is determined according to the width and height of the minimum bounding rectangle of the true detection box and the predicted detection box, and the angle loss, and the angle loss is determined according to the height difference between the center points of the true detection box and the predicted detection box and the distance between the center points of the true detection box and the predicted detection box.
[0210] In some embodiments of the present invention, the bin location-related information includes any one or more of the following:
[0211] Bin location information, bin status information.
[0212] In some embodiments of the present invention, the bin status information includes an in-stock status or an idle status.
[0213] In an embodiment of the present invention, by acquiring a video stream of a warehousing area collected by a camera and using a rotated object detection model to detect the video stream, a detection result for the warehousing area is obtained. The rotated object detection model performs object detection through a rotatable detection frame. The detection result includes bin location-related information. When a handling task is received, the target bin in the warehousing area is determined according to the bin location-related information, so as to schedule a handling device to execute a handling task at the target bin, realizing the detection of bins by using a rotated object detection model. The rotated object detection model performs bin detection through a rotatable detection frame, improving the accuracy of bin detection, and further improving the accuracy of scheduling the handling device.
[0214] Some embodiments of the present invention further provide an electronic device, which may include a processor, a memory, and a computer program stored on the memory and capable of running on the processor. When the computer program is executed by the processor, the above-mentioned warehousing-based processing method is implemented.
[0215] Some embodiments of the present invention further provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned warehousing-based processing method is implemented.
[0216] For the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the related parts, reference may be made to the partial description of the method embodiments.
[0217] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties. And the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for the user to select to authorize or refuse.
[0218] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments may be referred to each other.
[0219] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, an apparatus, or a computer program product. Therefore, the embodiments of the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0220] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate a device for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0221] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0222] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0223] Although the preferred embodiments of the embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications to these embodiments once they know the basic creative concepts. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.
[0224] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or terminal device. Without more limitations, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or terminal device including the above elements.
[0225] The above provides a detailed introduction to a warehousing-based processing method, device, electronic device and medium. In this article, specific examples are used to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A warehousing-based processing method, characterized in that, The method includes: Obtaining a video stream of a warehousing area collected by a camera; Using a rotated object detection model to detect the video stream to obtain a detection result for the warehousing area; wherein, the rotated object detection model performs object detection through a rotatable detection frame, and the detection result includes information related to storage locations; When a handling task is received, determining a target storage location in the warehousing area according to the information related to the storage location, so as to schedule a handling device to execute the handling task at the target storage location.
2. The method according to claim 1, wherein The detection result further includes abnormal situation information, and further includes: Generating a warning message and giving feedback according to the abnormal situation information.
3. The method according to claim 2, wherein It further includes: Controlling the handling device to avoid according to the abnormal situation information.
4. The method according to claim 2, wherein The abnormal situation information includes any one or more of the following: Abnormal situation information endangering production safety, abnormal situation information affecting the movement of the handling device.
5. The method according to claim 1, wherein It further includes: Performing statistics on the number of storage locations according to the information related to the storage location to obtain a statistical result of the number of storage locations; The determining the target storage location in the warehousing area according to the information related to the storage location includes: Determining the target storage location in the warehousing area according to the information related to the storage location, or the information related to the storage location and the statistical result of the number of storage locations.
6. The method according to claim 5, wherein The statistical result of the number of storage locations includes: the number of storage locations in the stocked state and the number of data of storage locations in the idle state.
7. The method according to any one of claims 1 to 6, characterized in that The using the rotated object detection model to detect the video stream to obtain a detection result for the warehousing area includes: Obtaining an image frame from the video stream; In the rotated object detection model, generating a feature map according to the image frame and reconstructing the feature map to obtain a reconstructed feature map; Determining a detection result for the warehousing area according to the reconstructed feature map.
8. The method according to claim 7, characterized in that, The video stream includes multiple video streams collected by multiple cameras, and the obtaining an image frame from the video stream includes: Obtaining local image frames from the multiple video streams; Stitching the local image frames in the multiple video streams to obtain a global image frame.
9. The method according to claim 7, characterized in that, The reconstructing the feature map to obtain a reconstructed feature map includes: Obtaining a new feature vector of a feature point in the feature map through a rotatable detection frame, and using the new feature vector to replace the prior feature vector of the feature point; Traversing the feature points in the feature map to replace the feature vectors to obtain a reconstructed feature map.
10. The method according to claim 9, wherein Before using the new feature vector to replace the prior feature vector of the feature point, it further includes: Updating the new feature vector through bilinear interpolation.
11. The method according to claim 7, wherein The rotated object detection model is embedded with a cross-dimensional interaction attention module, and the cross-dimensional interaction attention module is used to obtain features of the fusion of the image spatial dimension and the image channel dimension in the feature map.
12. The method according to claim 11, characterized in that, The image spatial dimension includes the image width dimension and the image height dimension, and the image channel dimension includes the three primary color channel dimensions of the image and the filter channel dimension.
13. The method according to claim 11, wherein The rotation target detection model has a feature refinement module, and the feature refinement module has a dual-path convolution part and a feature refinement part. The cross-dimensional interaction attention module is embedded in the dual-path convolution part and the feature refinement part.
14. The method according to claim 11, characterized in that, The cross-dimensional interaction attention module includes at least three branches, where: The first branch is used to obtain the features of the image spatial dimension. The second branch is used to obtain the features of the fusion of the image channel dimension and the first image spatial dimension. The third branch is used to obtain the features of the fusion of the channel dimension information and the second image spatial dimension. Wherein, one of the first image spatial dimension and the second image spatial dimension is the image height dimension, and the other is the image width dimension.
15. The method according to claim 14, wherein The cross-dimensional interaction attention module is also used to: fuse the features obtained by the at least three branches.
16. The method according to claim 1, wherein During the training of the rotation target detection model, an inclined intersection over union (IoU) score loss function is adopted, and the loss value of the inclined IoU score loss function is determined according to the IoU score loss, the shape loss, and the distance loss.
17. The method according to claim 16, characterized in that, The IoU score loss is determined according to the intersection and union of the ground truth detection box and the predicted detection box. The shape loss is determined according to the widths and heights of the ground truth detection box and the predicted detection box. The distance loss is determined according to the widths and heights of the minimum bounding rectangles of the ground truth detection box and the predicted detection box, and the angle loss, and the angle loss is determined according to the height difference between the centers of the ground truth detection box and the predicted detection box and the distance between the centers of the ground truth detection box and the predicted detection box.
18. The method according to claim 1, characterized in that, The bin-related information includes any one or more of the following: Bin location information, bin status information.
19. The method according to claim 18, characterized in that, The bin status information includes the in-stock status or the idle status.
20. A processing device based on a warehouse, characterized in that, The device is used to: Obtain the video stream of the warehousing area collected by the camera. Use the rotation target detection model to detect the video stream to obtain the detection result of the warehousing area; wherein, the rotation target detection model performs target detection through a rotatable detection box, and the detection result includes bin-related information. When receiving a handling task, determine the target bin in the warehousing area according to the bin-related information, so as to schedule the handling equipment to execute the handling task at the target bin.
21. An electronic device, characterized in that, It includes a processor, a memory, and a computer program stored on the memory and capable of running on the processor. When the computer program is executed by the processor, it implements the warehousing-based processing method according to any one of claims 1 to 19.
22. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is executed by the processor, it implements the warehousing-based processing method according to any one of claims 1 to 19.