A method, device, electronic device, and storage medium for detecting the integrity of a shelf

The method enhances shelf completeness detection by integrating region detection, scene classification, and shelf number analysis using supervised and unsupervised learning, addressing the inadequacies of existing methods and improving accuracy and cost-effectiveness.

CN116129271BActive Publication Date: 2025-07-15BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310174523.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-27
Publication Date
2025-07-15
Estimated Expiration
2043-02-27

AI Technical Summary

Technical Problem

The prior art is difficult to effectively detect the integrity of shelf images, which affects the accuracy of item display information.

Method used

By conducting area detection, display scene classification and linear detection of shelf images, combining target area detection information, display scenes and shelf layers, deep learning and semi-supervised learning technology training models are used to generate accurate shelf integrity judgments.

Benefits of technology

It improves the accuracy of shelf integrity detection in complex environments, reduces the cost of manual labeling, and improves the generalization and detection accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116129271B_ABST
    Figure CN116129271B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method, apparatus, electronic device, and storage medium for detecting the integrity of a shelf, relating to the field of artificial intelligence technology, specifically to the fields of deep learning, image processing, and computer vision technology. The method includes: performing region detection on a target shelf image to be detected to obtain target region detection information of the target shelf image; the target region detection information includes target detection frames in the target shelf image and target end types to which the target detection frames belong; performing display scene classification on the target shelf image to obtain a target display scene to which the target shelf image belongs; performing straight line detection on the target shelf image to obtain the number of target shelf layers in the target shelf image; determining whether the shelf in the target shelf image is completely photographed based on at least one of the target region detection information, the target display scene, and the number of target shelf layers. Through the above technical solution, the accuracy of shelf integrity detection can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, specifically to the fields of deep learning, image processing, and computer vision technology. Specifically, it relates to a method, device, electronic device, and storage medium for detecting the integrity of a shelf. Background Art

[0002] With the rapid development of Artificial Intelligence (AI), AI has been widely applied in the fields of computer vision technology, speech recognition technology, natural language processing technology, deep learning, big data processing technology, etc.

[0003] In the fast-moving consumer goods industry, shelf image detection based on computer vision can obtain information on item display in the shelf, such as the occupancy ratio of the shelf surface, the number of empty positions, the out-of-stock rate, the qualification rate, etc. However, whether the shelf image is completely captured determines whether the item display information is true and effective. Therefore, how to detect the integrity of the shelf is very important. Summary of the Invention

[0004] The present disclosure provides a method, device, electronic device, and storage medium for detecting the integrity of a shelf.

[0005] According to one aspect of the present disclosure, a method for detecting the integrity of a shelf is provided, including:

[0006] Performing region detection on a target shelf image to be detected to obtain target region detection information of the target shelf image; the target region detection information includes a target detection frame in the target shelf image and a target end type to which the target detection frame belongs;

[0007] Performing display scene classification on the target shelf image to obtain a target display scene to which the target shelf image belongs;

[0008] Performing straight line detection on the target shelf image to obtain the number of target shelf layers in the target shelf image;

[0009] Determining whether the shelf in the target shelf image is completely captured according to at least one of the target region detection information, the target display scene, and the number of target shelf layers.

[0010] According to another aspect of the present disclosure, a method for training a target region detection model is provided, including:

[0011] Inputting an unlabeled second sample shelf image into a teacher region detection model to obtain pseudo-region detection information of the second sample shelf image; the teacher region detection model is pre-trained using a labeled first sample shelf image;

[0012] It is obtained by fusing the first sample shelf image and the second sample shelf image to obtain a third sample shelf image; the third sample shelf image has an annotated detection frame and an annotated end type.

[0013] Input the third sample shelf image into the student region detection model to obtain the student detection frame and the student end type output by the student region detection model.

[0014] Train the student region detection model according to the annotated detection frame and the annotated end type of the third sample shelf image, as well as the student detection frame and the student end type, and use the training result as the target region detection model.

[0015] According to another aspect of the present disclosure, a method for training a target scene classification model is provided, including:

[0016] Input the unannotated second sample shelf image into the initial scene classification model to obtain the pseudo-display scene of the second sample shelf image; the initial scene classification model is pre-trained using the annotated first sample shelf image.

[0017] It is obtained by fusing the first sample shelf image and the second sample shelf image to obtain a third sample shelf image; the third sample shelf image has an annotated display scene.

[0018] Input the third sample shelf image into the initial scene classification model to obtain the predicted display scene of the third sample shelf image.

[0019] Train the initial scene classification model according to the predicted display scene and the annotated display scene of the third sample shelf image, and use the training result as the target scene classification model.

[0020] According to another aspect of the present disclosure, a detection device for shelf integrity is provided, including:

[0021] A region detection module for performing region detection on the target shelf image to be detected to obtain the target region detection information of the target shelf image; the target region detection information includes the target detection frame in the target shelf image and the target end type to which the target detection frame belongs.

[0022] A scene classification module for classifying the display scene of the target shelf image to obtain the target display scene to which the target shelf image belongs.

[0023] A shelf layer number module for performing straight line detection on the target shelf image to obtain the target shelf layer number in the target shelf image.

[0024] An integrity module for determining whether the shelves in the target shelf image are photographed completely according to at least one of the target area detection information, the target display scenario, and the target number of shelf layers.

[0025] According to another aspect of the present disclosure, there is provided a training device for a target area detection model, including:

[0026] A teacher area detection module for inputting an unlabeled second sample shelf image into a teacher area detection model to obtain pseudo area detection information of the second sample shelf image; the teacher area detection model is pre-trained using a labeled first sample shelf image;

[0027] A sample fusion module for fusing the first sample shelf image and the second sample shelf image to obtain a third sample shelf image; the third sample shelf image has a labeled detection frame and a labeled end type;

[0028] A student area detection module for inputting the third sample shelf image into a student area detection model to obtain a student detection frame and a student end type output by the student area detection model;

[0029] An area detection training module for training the student area detection model according to the labeled detection frame and the labeled end type of the third sample shelf image, and the student detection frame and the student end type, and using the training result as the target area detection model.

[0030] According to another aspect of the present disclosure, there is provided a training device for a target scene classification model, including:

[0031] An initial classification module for inputting an unlabeled second sample shelf image into an initial scene classification model to obtain a pseudo display scenario of the second sample shelf image; the initial scene classification model is pre-trained using a labeled first sample shelf image;

[0032] A sample fusion module for fusing the first sample shelf image and the second sample shelf image to obtain a third sample shelf image; the third sample shelf image has a labeled display scenario;

[0033] A display prediction module for inputting the third sample shelf image into the initial scene classification model to obtain a predicted display scenario of the third sample shelf image;

[0034] A scene classification training module for training the initial scene classification model according to the predicted display scenario and the labeled display scenario of the third sample shelf image, and using the training result as the target scene classification model.

[0035] According to another aspect of the present disclosure, there is provided an electronic device, which includes:

[0036] at least one processor; and

[0037] a memory communicatively connected to the at least one processor; wherein,

[0038] the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method provided in any embodiment of the present disclosure.

[0039] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method provided in any embodiment of the present disclosure.

[0040] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0042] Figure 1a is a flowchart of a method for detecting shelf integrity provided according to an embodiment of the present disclosure;

[0043] Figure 1b is a schematic diagram of the detection principle of shelf integrity provided according to an embodiment of the present disclosure;

[0044] Figure 2a is a flowchart of another method for detecting shelf integrity provided according to an embodiment of the present disclosure;

[0045] Figure 2b is a schematic diagram of the training process of a target area detection model provided according to an embodiment of the present disclosure;

[0046] Figure 3a is a flowchart of yet another method for detecting shelf integrity provided according to an embodiment of the present disclosure;

[0047] Figure 3b is a schematic diagram of the structure of a target area detection model and a target scene classification model provided according to an embodiment of the present disclosure;

[0048] Figure 4 is a flowchart of yet another method for detecting shelf integrity provided according to an embodiment of the present disclosure;

[0049] Figure 5 It is a flowchart of a method for training an object area detection model provided according to an embodiment of the present disclosure;

[0050] Figure 6 It is a flowchart of a method for training an object scene classification model provided according to an embodiment of the present disclosure;

[0051] Figure 7 It is a schematic structural diagram of a detection device for shelf integrity provided according to an embodiment of the present disclosure;

[0052] Figure 8 It is a schematic structural diagram of a training device for an object area detection model provided according to an embodiment of the present disclosure;

[0053] Figure 9 It is a schematic structural diagram of a training device for an object scene classification model provided according to an embodiment of the present disclosure;

[0054] Figure 10 It is a block diagram of an electronic device for implementing the shelf integrity detection method according to an embodiment of the present disclosure. Detailed implementation manners

[0055] Figure 1a It is a flowchart of a method for detecting shelf integrity provided according to an embodiment of the present disclosure. This method is applicable to the situation of detecting whether a shelf image is completely captured. This method can be executed by a detection device for shelf integrity, and this device can be implemented in a software and / or hardware manner and can be integrated into an electronic device. As Figure 1a shown, the method for detecting shelf integrity in this embodiment may include:

[0056] S101, perform area detection on the target shelf image to be detected to obtain target area detection information of the target shelf image; the target area detection information includes a target detection frame in the target shelf image and a target end type to which the target detection frame belongs;

[0057] S102, perform display scene classification on the target shelf image to obtain a target display scene to which the target shelf image belongs;

[0058] S103, perform straight line detection on the target shelf image to obtain the number of target shelf layers in the target shelf image;

[0059] S104, determine whether the shelf in the target shelf image is completely captured according to at least one of the target area detection information, the target display scene, and the number of target shelf layers.

[0060] Figure 1bIt is a schematic diagram of the detection principle of shelf integrity provided according to an embodiment of the present disclosure. Refer to Figure 1b In the embodiment of the present disclosure, region detection and display scenario classification are respectively performed on the target shelf image to obtain the target region detection information of the target shelf image and the target display scenario to which the target shelf image belongs. The target region detection information may include a target detection frame in the target shelf image and the target end type to which the target detection frame belongs, and the target end type is the bottom or the top. The target display scenario may be a shelf, an end rack, a single-door cabinet, a multi-door cabinet, a cold air cabinet, or a floor dragon, etc.

[0061] Refer to Figure 1b In addition, line detection is performed on the target shelf image, and the number of target shelf layers in the target shelf image is obtained according to the line detection result. Strategy processing can be combined with at least one of the target region detection information, the target display scenario, and the number of target shelf layers, and it is determined whether the shelf in the target shelf image is completely photographed according to the processing result. Exemplarily, when it is determined according to the target region detection information that the target shelf image includes a bottom region and a top region, it can be determined that the shelf in the target shelf image is completely photographed; when the target shelf image lacks a bottom region or a top region, the target display scenario and the number of target shelf layers can be introduced as supplementary information to further determine whether the shelf in the target shelf image is completely photographed.

[0062] By using the target region detection information as a reference basis for shelf integrity and introducing the target display scenario and the number of target shelf layers as supplementary information to judge whether the shelf in the target shelf image is completely photographed, the accuracy of the judgment result can be improved in complex situations such as the lack of a top region or a bottom region, that is, the accuracy of shelf integrity can be improved.

[0063] The technical solution provided by the embodiment of the present disclosure, by respectively performing region detection and display scenario classification on the target shelf image to obtain the target region detection information and the target display scenario of the target shelf image; and determining the number of target shelf layers in the target shelf image; combining the target region detection information, the target display scenario, and the number of target shelf layers to judge whether the shelf in the target shelf image is completely photographed, can improve the accuracy of shelf integrity in complex situations such as the lack of a top region or a bottom region.

[0064] Figure 2a It is a flowchart of another method for detecting shelf integrity provided according to an embodiment of the present disclosure. Refer to Figure 2a In this embodiment, the method for detecting shelf integrity may include:

[0065] S201. Detect the region of the target shelf image to be detected through a target region detection model, and obtain the target region detection information of the target shelf image; the target region detection information includes the target detection box in the target shelf image and the target end type to which the target detection box belongs.

[0066] S202. Classify the display scenario of the target shelf image to obtain the target display scenario to which the target shelf image belongs.

[0067] S203. Detect straight lines in the target shelf image to obtain the number of layers of the target shelf in the target shelf image.

[0068] S204. Determine whether the shelf in the target shelf image is completely photographed according to at least one of the target region detection information, the target display scenario, and the number of layers of the target shelf.

[0069] Among them, the target region detection model is trained in the following way: input the unlabeled second sample shelf image into the teacher region detection model to obtain the pseudo-region detection information of the second sample shelf image; the teacher region detection model is pre-trained with the labeled first sample shelf image; fuse the first sample shelf image and the second sample shelf image to obtain a third sample shelf image; the third sample shelf image has a labeled detection box and a labeled end type; input the third sample shelf image into the student region detection model to obtain the student detection box and the student end type output by the student region detection model; train the student region detection model according to the labeled detection box and the labeled end type of the third sample shelf image, and the training result is used as the target region detection model.

[0070] In the embodiments of the present disclosure, the integrity of the target shelf image is judged based on deep learning. Specifically, the target region detection model is used to detect the region of the target shelf image to obtain the target region detection information. The target region detection model can be trained by using target detection technology and semi-supervised learning technology.

[0071] During the training process of the target region detection model, the labeled first sample shelf image can be obtained. The first sample shelf image has a labeled detection box and a labeled end type. The labeled end type is used to represent the end type to which the labeled detection box belongs; the end type is the bottom or the top. Specifically, based on the target detection network, the teacher region detection model can be trained according to the first sample shelf image, the labeled detection box of the first sample shelf image, and the labeled end type. The teacher region detection model can be used to guide the training of the student region detection model.

[0072] Reference Figure 2b, it also obtains unlabeled second-sample shelf images. For example, a large amount of shelf business return data is collected, and pseudo-labels are generated for the second-sample shelf images through a teacher region detection model. The unlabeled second-sample shelf images can be input into the teacher region detection model, and pseudo-detection boxes and pseudo-end types in the second-sample shelf images can be obtained according to the output of the teacher region detection model; the pseudo-end type is also the top or the bottom.

[0073] That is to say, in order to better utilize weakly labeled unsupervised learning, a teacher region detection model can be used to better iterate the student region detection model. The teacher region detection model is a deep network model that has been trained with both labeled and unlabeled samples and has a certain prediction effect. Specifically, weak augmentation can be performed on the unlabeled second-sample shelf images, the prediction results can be obtained by inputting them into the teacher region detection model, and the filtering method in the mean teacher model can be used to obtain pseudo-labels.

[0074] Moreover, for the first-sample shelf images, the labeled detection boxes and labeled end types of the first-sample shelf images; and for the second-sample shelf images, the pseudo-detection boxes and pseudo-end types of the second-sample shelf images are fused to obtain the third-sample shelf images, the labeled detection boxes and labeled end types of the third-sample shelf images. For example, the third-sample shelf images can include the first-sample shelf images and the second-sample shelf images, and the pseudo-detection boxes and pseudo-end types of the second-sample shelf images are used as the corresponding labeled detection boxes and labeled end types.

[0075] Reference Figure 2b , strong augmentation can also be performed on the labeled third-sample shelf images and input into the student region detection model to obtain the student detection boxes and student end types output by the student region detection model; according to the labeled detection boxes and labeled end types of the third-sample shelf images, and the student detection boxes and the student end types of the third-sample shelf images, the student region detection model is trained, and the training result is used as the target region detection model.

[0076] By training the teacher region detection model with the labeled first-sample shelf images, using the teacher region detection model to generate pseudo-region detection information (i.e., pseudo-labels) for the unlabeled second-sample shelf images, and fusing the labeled first-sample shelf images and the second-sample shelf images with pseudo-labels to obtain the third-sample shelf images, the labeled detection boxes and labeled end types of the third-sample shelf images; and by training the student region detection model with the third-sample shelf images and using the training result as the target region detection model, the training samples of the student region detection model can be enriched, thereby improving the accuracy of the target detection information and further improving the accuracy of the shelf integrity detection.

[0077] In an alternative embodiment, training the student region detection model based on the labeled detection box and labeled end type of the third sample shelf image, as well as the student detection box and the student end type, and using the training result as the target region detection model includes: determining the detection box loss of the student region detection model according to the labeled detection box of the third sample shelf image and the student detection box; determining the end type loss of the student region detection model according to the labeled end type of the third sample shelf image and the student end type; training the student region detection model according to the detection box loss and the end type loss of the student region detection model, and using the training result as the target region detection model.

[0078] The loss function of the student region detection model has two parts. One part is the detection box loss between the labeled detection box of the third sample shelf image and the student detection box, and the other part is the end type loss between the labeled end type of the third sample shelf image and the student end type. Training the student region detection model by combining the detection box loss and the end type loss can take into account the position of the target detection box and the accuracy of the target end type.

[0079] In an alternative embodiment, the method further includes: updating the network parameters in the teacher region detection model based on the Exponential Moving Average (EMA) method according to the network parameters in the student region detection model.

[0080] In the embodiments of the present disclosure, the network structures of the student region detection model and the teacher region detection model may be the same. For example, they may both be trained on the ppyoloe-x large model. Refer to Figure 2b , during the training process of the target region detection model, not only the network parameters in the student region detection model are updated, but also the network parameters in the teacher region detection model are updated based on the EMA method using the accelerated average of the gradients during the training process of the student region detection model.

[0081] Exemplarily, in the initial training stage of the student region detection model, the model is unstable, and relatively small weights can be set for the gradients; during the process of the student region detection model tending to be stable, the weights for the gradients can be increased accordingly. By updating the network parameters in the teacher region detection model based on the EMA according to the network parameters in the student region detection model, not only can the stability of model training be improved, but also the performance of the teacher region detection model can be enhanced while continuously optimizing the student region detection model, thereby further promoting the performance improvement of the student region detection model.

[0082] The technical solution provided by the embodiments of the present disclosure can improve the accuracy of target detection information, reduce the manual annotation cost, and enhance the model generalization ability by training a teacher region detection model using object detection and semi-supervised learning techniques, guiding the training of a student region detection model with the teacher region detection model, and synchronously optimizing the teacher region detection model while continuously optimizing the student region detection model, thereby further improving the accuracy of shelf integrity detection.

[0083] Figure 3a It is a flowchart of another method for detecting shelf integrity provided by the embodiments of the present disclosure. This embodiment is an alternative solution based on the above embodiments. Refer to Figure 3a , the method for detecting shelf integrity in this embodiment may include:

[0084] S301, perform region detection on the target shelf image to be detected to obtain target region detection information of the target shelf image; the target region detection information includes a target detection box in the target shelf image and a target end type to which the target detection box belongs;

[0085] S302, classify the display scene of the target shelf image to be detected through a target scene classification model to obtain a target display scene to which the target shelf image belongs;

[0086] S303, perform straight line detection on the target shelf image to obtain the number of target shelf layers in the target shelf image;

[0087] S304, determine whether the shelf in the target shelf image is completely photographed according to at least one of the target region detection information, the target display scene, and the number of target shelf layers.

[0088] In an alternative embodiment, the target scene classification model is trained in the following manner: input an unlabeled second sample shelf image into an initial scene classification model to obtain a pseudo-display scene of the second sample shelf image; the initial scene classification model is pre-trained using a labeled first sample shelf image; fuse the first sample shelf image and the second sample shelf image to obtain a third sample shelf image; the third sample shelf image has a labeled display scene; input the third sample shelf image into the initial scene classification model to obtain a predicted display scene of the third sample shelf image; train the initial scene classification model according to the predicted display scene and the labeled display scene of the third sample shelf image, and use the training result as the target scene classification model.

[0089] In the embodiments of the present disclosure, the target scene classification model can be trained using semi-supervised learning and multi-task learning techniques.

[0090] During the training process of the target scene classification model, labeled first sample shelf images can be obtained. The first sample shelf images have labeled display scenes. And, based on a multi-task network, an initial scene classification model can be obtained according to the first sample shelf images and the labeled display scenes of the first sample shelf images. Second sample shelf images without labels are also obtained. For example, a large amount of shelf business return data is collected, and pseudo-labels are generated for the second sample shelf images through the initial scene classification model. Specifically, the unlabeled second sample shelf images are input into the initial scene classification model, and pseudo-display scenes in the second sample shelf images are obtained according to the output of the initial scene classification model.

[0091] Moreover, the labeled display scenes of the first sample shelf images and the pseudo-display scenes of the second sample shelf images are fused to obtain the third sample shelf images and the labeled display scenes of the third sample shelf images. For example, the third sample shelf images can include the first sample shelf images and the second sample shelf images, and the pseudo-display scenes of the second sample shelf images are used as the corresponding labeled display scenes.

[0092] Also, the third sample shelf images are input into the initial scene classification model to obtain the predicted display scenes output by the initial scene classification model. According to the labeled display scenes and the predicted display scenes of the third sample shelf images, the initial scene classification model is trained, and the training result is used as the target scene classification model. The target scene classification model can be trained on the feature classification model EfficientNetB1. Specifically, the scene classification loss can be determined according to the labeled display scenes and the predicted display scenes of the third sample shelf images, and the initial scene classification model is trained according to the scene classification loss.

[0093] By using the labeled first sample shelf images to train the initial scene classification model, using the initial scene classification model to generate pseudo-display scenes (i.e., pseudo-labels) for the unlabeled second sample shelf images, and fusing the labeled first sample shelf images and the second sample shelf images with pseudo-labels to obtain the third sample shelf images and the labeled display scenes of the third sample shelf images; and using the third sample shelf images to train the initial scene classification model, and using the training result as the target scene classification model, it is possible to enrich the training samples of the target scene classification model, thereby improving the accuracy of scene classification and further improving the accuracy of shelf integrity detection.

[0094] In addition, it should be noted that with reference to Figure 3b, the training tasks of the target area detection model 31 and the target scene classification model 32 can also be integrated into one model for learning. The feature extraction layers of the two are reused to share parameters, reducing the training time consumption and keeping the time consumption of the sample shelf image within 50 ms.

[0095] The technical solution provided by the embodiments of the present disclosure trains an initial scene classification model by adopting semi-supervised learning and multi-task learning techniques, and enriches the training samples of the target scene classification model by using the initial scene classification model, thereby improving the accuracy of scene classification and further improving the accuracy of shelf integrity detection.

[0096] It should also be noted that the embodiments of the present disclosure do not specifically limit the direct detection algorithm. Exemplarily, the Canny (edge detection) operator can be used to extract the line information in the shelf image to obtain a binary image; input it into the Hough transform function. Considering the length of the shelf, minLineLength (the minimum length of the line) can be set to 120 pixels; after obtaining the lines, the messy lines are filtered to obtain coherent shelf lines, and the number of shelf lines is obtained as the number of target shelf layers.

[0097] Figure 4 is a flowchart of another method for detecting shelf integrity provided by the embodiments of the present disclosure. Refer to Figure 4 , the method for detecting shelf integrity in this embodiment may include:

[0098] S401, perform area detection on the target shelf image to be detected to obtain the target area detection information of the target shelf image; the target area detection information includes the target detection frame in the target shelf image and the target end type to which the target detection frame belongs;

[0099] S402, perform display scene classification on the target shelf image to obtain the target display scene to which the target shelf image belongs;

[0100] S403, perform straight line detection on the target shelf image to obtain the number of target shelf layers in the target shelf image;

[0101] S404, determine whether the target shelf image has a bottom area and a top; or, whether the target shelf image has a top area and the target display scene is the earthworm scene; if so, continue to execute S405; otherwise, jump to execute S406;

[0102] S405, determine that the shelf in the target shelf image is photographed completely;

[0103] S406, determine whether the shelf in the target shelf image is photographed completely according to the target display scene and the number of target shelf layers.

[0104] When it is determined according to the target area detection information that the target shelf image includes the bottom area and the top area, it is determined that the shelf in the target shelf image is completely photographed; when it is determined according to the target area detection information that the target shelf image includes the top area but does not include the bottom area, and the target display scenario is the earthworm scenario, it is stated that the target shelf image includes the bottom boundary, and it is also determined that the shelf in the target shelf image is completely photographed. It is also possible to determine whether the abscissa of the center point of the bottom area and the top area is less than the threshold, so as to further improve the accuracy of the shelf integrity detection.

[0105] In other cases, the target display scenario and the target shelf layer number can be combined to determine whether the shelf in the target shelf image is completely photographed. That is to say, for the complex environment where the shelf top cannot be obtained, the target display scenario and the target shelf layer number are used for functional supplementation, so as to improve the accuracy of the shelf integrity detection in the complex environment.

[0106] In an alternative embodiment, determining whether the shelf in the target shelf image is completely photographed according to the target display scenario and the target shelf layer number includes: determining the target shelf layer number interval associated with the target display scenario based on the association relationship between the preset candidate display scenarios and the candidate shelf layer number intervals; the target shelf layer number interval includes a first shelf layer number threshold and a second shelf layer number threshold; when the target shelf layer number is greater than the first shelf layer number threshold, the shelf in the target shelf image is completely photographed; when the target shelf layer number is less than the second shelf layer number threshold, the shelf in the target shelf image is not completely photographed.

[0107] In the embodiments of the present disclosure, when the top area and the bottom area are missing in the target shelf image, the target display scenario and the target shelf layer number are used for functional complementation. Among them, the association relationship between the candidate display scenarios and the candidate shelf layer number intervals can be constructed according to the prior knowledge of each candidate display scenario; the candidate shelf layer number interval of each candidate display scenario includes a first shelf layer number threshold and a second shelf layer number threshold, and the first shelf layer number threshold is greater than the second shelf layer number threshold. When the target shelf layer number is greater than the first shelf layer number threshold, it can be determined that the shelf in the target shelf image is completely photographed; when the target shelf layer number is less than the second shelf layer number threshold, it can be determined that the shelf in the target shelf image is not completely photographed.

[0108] Exemplarily, in the case where the candidate display scenario is a single - door standing cabinet, if the number of target shelf layers exceeds 5, the shelf is photographed completely; if the number of target shelf layers is less than 3, the shelf is photographed incompletely; in the case where the candidate display scenario is a double - door standing cabinet, if the number of target shelf layers exceeds 6, the shelf is photographed completely; if the number of target shelf layers is less than 4, the shelf is photographed incompletely; in the case where the candidate display scenario is a shelf, if the number of target shelf layers exceeds 10, the shelf is photographed completely; if the number of target shelf layers is less than 4, the shelf is photographed incompletely; in the case where the candidate display scenario is an end - cap, if the number of target shelf layers exceeds 8, the shelf is photographed completely; if the number of target shelf layers is less than 3, the shelf is photographed incompletely.

[0109] When the target shelf image has a top boundary and a bottom boundary, it can be directly determined that the shelf is photographed completely; when the target shelf image lacks the top boundary and the bottom boundary, based on the prior knowledge of each candidate display scenario, combined with the target display scenario and the number of target shelf layers for comprehensive judgment, the recall rate can be improved to a certain extent, and the accuracy rate can reach 95.7% in complex business sets.

[0110] The technical solution provided by the embodiments of the present disclosure directly determines that the shelf is photographed completely when the target shelf image has a top boundary and a bottom boundary, and the confidence level is relatively high; and when the top area and the bottom area are not detected, the target display scenario and the number of target shelf layers are used to assist in the judgment, which can improve the accuracy of shelf integrity detection in complex environments.

[0111] Figure 5 It is a flowchart of a training method for a target area detection model provided by an embodiment of the present disclosure. The target area detection model can be used for area detection of shelf images, and the area detection result can be used for integrity detection of shelf images. This method can be executed by a training device of the target area detection model, and the device can be implemented in software and / or hardware and can be integrated into an electronic device. As Figure 5 shown, the training method of the target area detection model in this embodiment may include:

[0112] S501, input an unlabeled second - sample shelf image into a teacher area detection model to obtain pseudo - area detection information of the second - sample shelf image; the teacher area detection model is pre - trained with a labeled first - sample shelf image;

[0113] S502, fuse the first - sample shelf image and the second - sample shelf image to obtain a third - sample shelf image; the third - sample shelf image has a labeled detection frame and a labeled end type;

[0114] S503. Input the third sample shelf image into the student area detection model to obtain the student detection box and the student end type output by the student area detection model.

[0115] S504. Train the student area detection model according to the labeled detection box and the labeled end type of the third sample shelf image, as well as the student detection box and the student end type, and use the training result as the target area detection model.

[0116] During the training process of the target area detection model, obtain the labeled first sample shelf image, which has a labeled detection box and a labeled end type. Based on the target detection network, train the teacher area detection model according to the first sample shelf image and the labeled detection box and the labeled end type of the first sample shelf image. The teacher area detection model can be used to guide the training of the student area detection model. Also obtain the unlabeled second sample shelf image, and generate pseudo-labels for the second sample shelf image through the teacher area detection model. The unlabeled second sample shelf image can be input into the teacher area detection model, and the pseudo-detection box and the pseudo-end type in the second sample shelf image can be obtained according to the output of the teacher area detection model.

[0117] Moreover, fuse the first sample shelf image, the labeled detection box and the labeled end type of the first sample shelf image; and the second sample shelf image, the pseudo-detection box and the pseudo-end type of the second sample shelf image to obtain the third sample shelf image, the labeled detection box and the labeled end type of the third sample shelf image. Train the student area detection model according to the labeled detection box and the labeled end type of the third sample shelf image, as well as the student detection box and the student end type of the third sample shelf image, and use the training result as the target area detection model.

[0118] In an alternative embodiment, the step of training the student area detection model according to the labeled detection box and the labeled end type of the third sample shelf image, as well as the student detection box and the student end type, and using the training result as the target area detection model includes: determining the detection box loss of the student area detection model according to the labeled detection box and the student detection box of the third sample shelf image; determining the end type loss of the student area detection model according to the labeled end type and the student end type of the third sample shelf image; training the student area detection model according to the detection box loss and the end type loss of the student area detection model, and using the training result as the target area detection model.

[0119] During the training process of the target region detection model, the loss function of the student region detection model has two parts. One part is the detection box loss between the labeled detection box of the third sample shelf image and the student detection box, and the other part is the end type loss between the labeled end type of the third sample shelf image and the student end type. By combining the detection box loss and the end type loss to train the student region detection model, the position of the target detection box and the accuracy of the target end type can be taken into account.

[0120] In an alternative embodiment, the method further includes: based on the exponential moving average (EMA) method, updating the network parameters in the teacher region detection model according to the network parameters in the student region detection model.

[0121] During the training process of the target region detection model, not only are the network parameters in the student region detection model updated, but also the network parameters in the teacher region detection model are updated based on the EMA method using the accelerated average of the gradients during the training process of the student region detection model.

[0122] The technical solution of the embodiments of the present disclosure can improve the accuracy of target detection information, reduce the manual annotation cost, and enhance the model generalization ability by training the teacher region detection model using object detection and semi-supervised learning techniques, guiding the training of the student region detection model using the teacher region detection model, and synchronously optimizing the teacher region detection model while continuously optimizing the student region detection model.

[0123] Figure 6 is a flowchart of a method for training a target scene classification model provided by an embodiment of the present disclosure. The target scene classification model can be used to perform scene classification on a shelf image to obtain a display scene, and the display scene can be used to perform integrity detection on the shelf image. This method can be executed by a training device of the target scene classification model, and the device can be implemented in a software and / or hardware manner and integrated into an electronic device. As Figure 6 shown, the method for training the target scene classification model in this embodiment may include:

[0124] S601, inputting the unlabeled second sample shelf image into the initial scene classification model to obtain the pseudo-display scene of the second sample shelf image; the initial scene classification model is pre-trained using the labeled first sample shelf image;

[0125] S602, fusing the first sample shelf image and the second sample shelf image to obtain a third sample shelf image; the third sample shelf image has a labeled display scene;

[0126] S603. Input the third sample shelf image into the initial scene classification model to obtain the predicted display scene of the third sample shelf image.

[0127] S604. Train the initial scene classification model according to the predicted display scene of the third sample shelf image and the labeled display scene, and use the training result as the target scene classification model.

[0128] During the training process of the target scene classification model, labeled first sample shelf images with labeled display scenes can be obtained. And, based on a multi-task network, an initial scene classification model can be obtained according to the first sample shelf images and their labeled display scenes. Unlabeled second sample shelf images are also obtained, and pseudo-labels are generated for the second sample shelf images through the initial scene classification model.

[0129] Moreover, for the first sample shelf images, their labeled display scenes; and, for the second sample shelf images, their pseudo-display scenes are fused to obtain the third sample shelf images and their labeled display scenes. Also, the third sample shelf images are input into the initial scene classification model to obtain the predicted display scenes output by the initial scene classification model. According to the labeled display scenes and the predicted display scenes of the third sample shelf images, the initial scene classification model is trained, and the training result is used as the target scene classification model.

[0130] The technical solution of the embodiments of the present disclosure improves the accuracy of scene classification by training the initial scene classification model using semi-supervised learning and multi-task learning techniques and enriching the training samples of the target scene classification model with the initial scene classification model.

[0131] Figure 7 It is a schematic structural diagram of a shelf integrity detection device provided according to an embodiment of the present disclosure. This embodiment is applicable to the situation of detecting whether a shelf image is completely captured. The device can be implemented in software and / or hardware. As Figure 7 shown, the shelf integrity detection device 700 of this embodiment may include:

[0132] An area detection module 710, configured to perform area detection on the target shelf image to be detected to obtain the target area detection information of the target shelf image; the target area detection information includes the target detection box in the target shelf image and the target end type to which the target detection box belongs.

[0133] A scene classification module 720, configured to perform display scene classification on the target shelf image to obtain the target display scene to which the target shelf image belongs.

[0134] The shelf layer number module 730 is used to perform line detection on the target shelf image to obtain the number of target shelf layers in the target shelf image;

[0135] The integrity module 740 is used to determine whether the shelf in the target shelf image is completely photographed according to at least one of the target area detection information, the target display scenario, and the number of target shelf layers.

[0136] In an alternative embodiment, the area detection module 710 is specifically configured to: perform area detection on the target shelf image to be detected through a target area detection model;

[0137] The above-mentioned shelf integrity detection device 700 further includes an area detection training module, and the area detection training module includes:

[0138] The teacher area detection unit is used to input the unlabeled second sample shelf image into the teacher area detection model to obtain the pseudo area detection information of the second sample shelf image; the teacher area detection model is pre-trained with the labeled first sample shelf image;

[0139] The sample fusion unit is used to fuse the first sample shelf image and the second sample shelf image to obtain a third sample shelf image; the third sample shelf image has a labeled detection frame and a labeled end type;

[0140] The student area detection unit is used to input the third sample shelf image into the student area detection model to obtain the student detection frame and the student end type output by the student area detection model;

[0141] The area detection training unit is used to train the student area detection model according to the labeled detection frame and the labeled end type of the third sample shelf image, and the student detection frame and the student end type, and use the training result as the target area detection model.

[0142] In an alternative embodiment, the area detection training unit includes:

[0143] The detection frame loss sub-unit is used to determine the detection frame loss of the student area detection model according to the labeled detection frame of the third sample shelf image and the student detection frame;

[0144] The end loss sub-unit is used to determine the end type loss of the student area detection model according to the labeled end type of the third sample shelf image and the student end type;

[0145] The region detection training subunit is used to train the student region detection model according to the detection box loss and the end type loss of the student region detection model, and use the training result as the target region detection model.

[0146] In an alternative embodiment, the region detection training module 710 further includes:

[0147] The teacher update subunit is used to update the network parameters in the teacher region detection model based on the exponential moving average (EMA) method according to the network parameters in the student region detection model.

[0148] In an alternative embodiment, the scene classification module 720 is specifically configured to: classify the display scene of the target shelf image to be detected through the target scene classification model;

[0149] The above-mentioned shelf integrity detection device 700 further includes a scene classification training module, and the scene classification training module includes:

[0150] The initial classification unit is used to input the unlabeled second sample shelf image into the initial scene classification model to obtain the pseudo-display scene of the second sample shelf image; the initial scene classification model is pre-trained using the labeled first sample shelf image;

[0151] The sample fusion unit is used to fuse the first sample shelf image and the second sample shelf image to obtain a third sample shelf image; the third sample shelf image has a labeled display scene;

[0152] The display prediction unit is used to input the third sample shelf image into the initial scene classification model to obtain the predicted display scene of the third sample shelf image;

[0153] The scene classification training unit is used to train the initial scene classification model according to the predicted display scene and the labeled display scene of the third sample shelf image, and use the training result as the target scene classification model.

[0154] In an alternative embodiment, the integrity module 740 is specifically configured to:

[0155] When it is determined that the target shelf image has a bottom region and a top region according to the target region detection information; or when it is determined that the target shelf image has a top region according to the target region detection information and the target display scene is the earthworm scene, it is determined that the shelf in the target shelf image is photographed completely;

[0156] Otherwise, determine whether the shelf in the target shelf image is completely captured according to the target display scenario and the target shelf layer number.

[0157] In an alternative embodiment, the integrity module 740 is specifically configured to:

[0158] Based on the association relationship between the preset candidate display scenarios and the candidate shelf layer number intervals, determine the target shelf layer number interval associated with the target display scenario; the target shelf layer number interval includes a first shelf layer threshold and a second shelf layer threshold;

[0159] When the target shelf layer number is greater than the first shelf layer threshold, the shelf in the target shelf image is completely captured;

[0160] When the target shelf layer number is less than the second shelf layer threshold, the shelf in the target shelf image is not completely captured.

[0161] The technical solution of the embodiments of the present disclosure can determine whether the shelf in the target shelf image is completely captured by combining the target area detection information, the target display scenario, and the target shelf layer number, which can improve the accuracy of shelf integrity in complex situations such as the lack of the top area or the bottom area. By using the semi-supervised end-to-end object detection algorithm (soft-teacher mode) for weak annotation, the manual annotation cost is reduced, and at the same time, the generalization of the model is improved, enabling it to be used in more complex scenarios.

[0162] Figure 8 It is a schematic structural diagram of a training device for a target area detection model provided by an embodiment of the present disclosure. This embodiment is applicable to the situation of performing area detection on shelf images. The device can be implemented in software and / or hardware. As Figure 8 shown, the training device 800 for the target area detection model of this embodiment may include:

[0163] A teacher area detection module, configured to input an unlabeled second sample shelf image into a teacher area detection model to obtain pseudo area detection information of the second sample shelf image; the teacher area detection model is pre-trained using a labeled first sample shelf image;

[0164] A sample fusion module, configured to fuse the first sample shelf image and the second sample shelf image to obtain a third sample shelf image; the third sample shelf image has a labeled detection box and a labeled end type;

[0165] A student area detection module, configured to input the third sample shelf image into a student area detection model to obtain a student detection box and a student end type output by the student area detection model;

[0166] An area detection training module, configured to train the student area detection model according to the labeled detection box and the labeled end type of the third sample shelf image, as well as the student detection box and the student end type, and use the training result as the target area detection model.

[0167] In an alternative embodiment, the area detection training module includes:

[0168] A detection box loss unit, configured to determine the detection box loss of the student area detection model according to the labeled detection box of the third sample shelf image and the student detection box;

[0169] An end loss unit, configured to determine the end type loss of the student area detection model according to the labeled end type of the third sample shelf image and the student end type;

[0170] An area detection training unit, configured to train the student area detection model according to the detection box loss of the student area detection model and the end type loss of the student area detection model, and use the training result as the target area detection model.

[0171] In an alternative embodiment, the above training device 800 of the target area detection model further includes:

[0172] A teacher update unit, configured to update the network parameters in the teacher area detection model based on the exponential moving average (EMA) method according to the network parameters in the student area detection model.

[0173] The technical solution of the embodiments of the present disclosure can improve the accuracy of target detection information, reduce the manual annotation cost, and enhance the model generalization ability by training the teacher area detection model using object detection and semi-supervised learning techniques, guiding the training of the student area detection model using the teacher area detection model, and synchronously optimizing the teacher area detection model while continuously optimizing the student area detection model.

[0174] Figure 9 It is a schematic structural diagram of a training device of a target scene classification model provided by an embodiment of the present disclosure. This embodiment is applicable to the situation of scene classification of shelf images. The device can be implemented in software and / or hardware. As Figure 9 shown, the training device 900 of the target scene classification model of this embodiment may include:

[0175] An initial classification module 910 is configured to input an unlabeled second sample shelf image into an initial scene classification model to obtain a pseudo-display scene of the second sample shelf image; the initial scene classification model is pre-trained using labeled first sample shelf images;

[0176] A sample fusion module 920 is configured to fuse the first sample shelf image and the second sample shelf image to obtain a third sample shelf image; the third sample shelf image has a labeled display scene;

[0177] A display prediction module 930 is configured to input the third sample shelf image into the initial scene classification model to obtain a predicted display scene of the third sample shelf image;

[0178] A scene classification training module 940 is configured to train the initial scene classification model according to the predicted display scene and the labeled display scene of the third sample shelf image, and use the training result as a target scene classification model.

[0179] The technical solution of the embodiment of the present disclosure trains an initial scene classification model by adopting semi-supervised learning and multi-task learning techniques, and enriches the training samples of the target scene classification model by using the initial scene classification model, thereby improving the accuracy of scene classification.

[0180] In the technical solution of the present disclosure, the acquisition, storage, and application of user personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0181] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0182] Figure 10 It is a block diagram of an electronic device for implementing the shelf integrity detection method of the embodiment of the present disclosure.

[0183] Figure 10 FIG. shows a schematic block diagram of an exemplary electronic device 1000 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0184] As Figure 10As shown, the electronic device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1002 or a computer program loaded from a storage unit 1008 into a random access memory (RAM) 1003. In the RAM 1003, various programs and data required for the operation of the electronic device 1000 can also be stored. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.

[0185] Multiple components in the electronic device 1000 are connected to the I / O interface 1005, including: an input unit 1006, such as a keyboard, a mouse, etc.; an output unit 1007, such as various types of displays, speakers, etc.; a storage unit 1008, such as a magnetic disk, an optical disc, etc.; and a communication unit 1009, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1009 allows the electronic device 1000 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0186] The computing unit 1001 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 1001 executes the various methods and processes described above, such as the method for detecting shelf integrity. For example, in some embodiments, the method for detecting shelf integrity can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into the RAM 1003 and executed by the computing unit 1001, one or more steps of the method for detecting shelf integrity described above can be executed. Alternatively, in other embodiments, the computing unit 1001 can be configured to execute the method for detecting shelf integrity by any other appropriate means (e.g., by means of firmware).

[0187] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems-on-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0188] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0189] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0190] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball), by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and the input received from the user can be in any form (including acoustic input, speech input, or tactile input).

[0191] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0192] A computer system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, or a server of a distributed system, or a server combined with a blockchain.

[0193] Artificial intelligence is a discipline that studies to make a computer simulate certain thinking processes and intelligent behaviors of humans (such as learning, reasoning, thinking, planning, etc.), and there are both hardware-level technologies and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, machine learning / deep learning technology, big data processing technology, and knowledge graph technology.

[0194] Cloud computing refers to a technical system that accesses an elastic and scalable shared physical or virtual resource pool through a network. The resources can include servers, operating systems, networks, software, applications, storage devices, etc., and the resources can be deployed and managed in a demand-based and self-service manner. Through cloud computing technology, it is possible to provide efficient and powerful data processing capabilities for the application and model training of technologies such as artificial intelligence and blockchain.

[0195] It should be understood that various forms of processes shown above can be used, and steps can be reordered, added, or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. There is no limitation here.

[0196] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A method for detecting the integrity of a shelf, comprising: Performing region detection on a target shelf image to be detected to obtain target region detection information of the target shelf image; The target region detection information includes a target detection frame in the target shelf image and a target end type to which the target detection frame belongs; Performing display scene classification on the target shelf image to obtain a target display scene to which the target shelf image belongs; Performing straight line detection on the target shelf image to obtain the number of target shelf layers in the target shelf image; Determining whether the shelf in the target shelf image is completely photographed according to at least one of the target region detection information, the target display scene, and the number of target shelf layers, including: When it is determined according to the target region detection information that the target shelf image includes a bottom region and a top region, determining that the shelf in the target shelf image is completely photographed; When the target shelf image lacks a bottom region or a top region, determining whether the shelf in the target shelf image is completely photographed according to the target display scene and the number of target shelf layers.

2. The method according to claim 1, wherein The performing region detection on the target shelf image to be detected includes: Performing region detection on the target shelf image to be detected through a target region detection model; The target region detection model is trained in the following manner: Inputting an unlabeled second sample shelf image into a teacher region detection model to obtain pseudo-region detection information of the second sample shelf image; the teacher region detection model is pre-trained with a labeled first sample shelf image; Fusing the first sample shelf image and the second sample shelf image to obtain a third sample shelf image; the third sample shelf image has a labeled detection frame and a labeled end type; Inputting the third sample shelf image into a student region detection model to obtain a student detection frame and a student end type output by the student region detection model; Training the student region detection model according to the labeled detection frame and the labeled end type of the third sample shelf image, and the student detection frame and the student end type, and using the training result as the target region detection model.

3. The method according to claim 2, wherein, The training the student region detection model according to the labeled detection frame and the labeled end type of the third sample shelf image, and the student detection frame and the student end type, and using the training result as the target region detection model includes: Determining a detection frame loss of the student region detection model according to the labeled detection frame and the student detection frame of the third sample shelf image; Determining an end type loss of the student region detection model according to the labeled end type and the student end type of the third sample shelf image; Training the student region detection model according to the detection frame loss and the end type loss of the student region detection model, and using the training result as the target region detection model.

4. The method according to claim 2, further comprising: Based on the exponential moving average (EMA) method, update the network parameters in the teacher region detection model according to the network parameters in the student region detection model.

5. The method according to claim 1, wherein The classification of the display scenario for the target shelf image includes: Classify the display scenario of the target shelf image to be detected through the target scenario classification model; The target scenario classification model is trained in the following way: Input the unlabeled second sample shelf image into the initial scenario classification model to obtain the pseudo-display scenario of the second sample shelf image; the initial scenario classification model is pre-trained with the labeled first sample shelf images; Fuse the first sample shelf image and the second sample shelf image to obtain a third sample shelf image; the third sample shelf image has a labeled display scenario; Input the third sample shelf image into the initial scenario classification model to obtain the predicted display scenario of the third sample shelf image; Train the initial scenario classification model according to the predicted display scenario and the labeled display scenario of the third sample shelf image, and use the training result as the target scenario classification model.

6. The method according to any one of claims 1-5, wherein, When the target shelf image lacks the bottom region or the top region, determine whether the shelf in the target shelf image is completely captured according to the target display scenario and the number of target shelf layers, including: When it is determined according to the target region detection information that the target shelf image has a top region and the target display scenario is the earthworm scenario, determine that the shelf in the target shelf image is completely captured; Otherwise, determine whether the shelf in the target shelf image is completely captured according to the target display scenario and the number of target shelf layers.

7. The method according to claim 6, wherein, The determination of whether the shelf in the target shelf image is completely captured according to the target display scenario and the number of target shelf layers includes: Based on the association relationship between the preset candidate display scenarios and the candidate shelf layer number intervals, determine the target shelf layer number interval associated with the target display scenario; the target shelf layer number interval includes a first shelf layer threshold and a second shelf layer threshold; When the number of target shelf layers is greater than the first shelf layer threshold, the shelf in the target shelf image is completely captured; When the number of target shelf layers is less than the second shelf layer threshold, the shelf in the target shelf image is not completely captured.

8. A training method for a target region detection model, including: Input the unlabeled second sample shelf image into the teacher region detection model to obtain the pseudo-region detection information of the second sample shelf image; The teacher region detection model is pre-trained with the labeled first sample shelf images; Fuse the first sample shelf image and the second sample shelf image to obtain a third sample shelf image; the third sample shelf image has a labeled detection frame and a labeled end type; Input the third sample shelf image into the student region detection model to obtain the student detection frame and the student end type output by the student region detection model; Train the student region detection model based on the labeled detection box and labeled end type of the third sample shelf image, as well as the student detection box and the student end type, and use the training result as the target region detection model.

9. The method according to claim 8, wherein, The training of the student region detection model based on the labeled detection box and labeled end type of the third sample shelf image, as well as the student detection box and the student end type, and using the training result as the target region detection model includes: Determine the detection box loss of the student region detection model according to the labeled detection box of the third sample shelf image and the student detection box; Determine the end type loss of the student region detection model according to the labeled end type of the third sample shelf image and the student end type; Train the student region detection model according to the detection box loss and the end type loss of the student region detection model, and use the training result as the target region detection model.

10. The method according to claim 8, further comprising: Updating the network parameters in the teacher region detection model based on the exponential moving average (EMA) method according to the network parameters in the student region detection model.

11. A method for training a target scene classification model, comprising: Input an unlabeled second sample shelf image into an initial scene classification model to obtain a pseudo-display scene of the second sample shelf image; The initial scene classification model is pre-trained using labeled first sample shelf images; Fuse the first sample shelf image and the second sample shelf image to obtain a third sample shelf image; the third sample shelf image has a labeled display scene; Input the third sample shelf image into the initial scene classification model to obtain a predicted display scene of the third sample shelf image; Train the initial scene classification model according to the predicted display scene and the labeled display scene of the third sample shelf image, and use the training result as the target scene classification model.

12. A detection device for shelf integrity, comprising: A region detection module for performing region detection on a target shelf image to be detected to obtain target region detection information of the target shelf image; The target region detection information includes a target detection box in the target shelf image and a target end type to which the target detection box belongs; A scene classification module for classifying the display scene of the target shelf image to obtain a target display scene to which the target shelf image belongs; A shelf layer module for performing line detection on the target shelf image to obtain the number of target shelf layers in the target shelf image; An integrity module for determining whether the shelf in the target shelf image is completely photographed according to at least one of the target region detection information, the target display scene, and the number of target shelf layers; The integrity module includes: A first integrity unit for determining that the shelf in the target shelf image is completely photographed when it is determined according to the target region detection information that the target shelf image includes a bottom region and a top region; A second integrity unit, configured to determine whether the shelf in the target shelf image is completely captured according to the target display scenario and the number of target shelf layers when the bottom area or the top area of the target shelf image is missing.

13. The device according to claim 12, wherein, The area detection module is specifically configured to: perform area detection on the target shelf image to be detected through a target area detection model; The device further includes an area detection training module, and the area detection training module includes: A teacher area detection unit, configured to input an unlabeled second sample shelf image into a teacher area detection model to obtain pseudo area detection information of the second sample shelf image; the teacher area detection model is pre-trained by using the labeled first sample shelf image; A sample fusion unit, configured to fuse the first sample shelf image and the second sample shelf image to obtain a third sample shelf image; the third sample shelf image has a labeled detection frame and a labeled end type; A student area detection unit, configured to input the third sample shelf image into a student area detection model to obtain a student detection frame and a student end type output by the student area detection model; An area detection training unit, configured to train the student area detection model according to the labeled detection frame and the labeled end type of the third sample shelf image, and the student detection frame and the student end type, and use the training result as the target area detection model.

14. The device according to claim 13, wherein, The area detection training unit includes: A detection frame loss sub-unit, configured to determine a detection frame loss of the student area detection model according to the labeled detection frame of the third sample shelf image and the student detection frame; An end loss sub-unit, configured to determine an end type loss of the student area detection model according to the labeled end type of the third sample shelf image and the student end type; An area detection training sub-unit, configured to train the student area detection model according to the detection frame loss of the student area detection model and the end type loss of the student area detection model, and use the training result as the target area detection model.

15. The device according to claim 13, wherein the area detection training module further includes: A teacher update sub-unit, configured to update the network parameters in the teacher area detection model according to the network parameters in the student area detection model based on the exponential moving average (EMA) method.

16. The device according to claim 12, wherein The scene classification module is specifically configured to: perform display scene classification on the target shelf image to be detected through a target scene classification model; The device further includes a scene classification training module, and the scene classification training module includes: An initial classification unit, configured to input an unlabeled second sample shelf image into an initial scene classification model to obtain a pseudo display scene of the second sample shelf image; the initial scene classification model is pre-trained by using the labeled first sample shelf image; A sample fusion unit, configured to fuse the first sample shelf image and the second sample shelf image to obtain a third sample shelf image; the third sample shelf image has a labeled display scene; An arrangement prediction unit, configured to input the third sample shelf image into an initial scene classification model to obtain a predicted arrangement scene of the third sample shelf image; A scene classification training unit, configured to train the initial scene classification model according to the predicted arrangement scene of the third sample shelf image and the labeled arrangement scene, and use the training result as the target scene classification model.

17. The device according to any one of claims 12 - 16, wherein, The second integrity unit is specifically configured to: When it is determined according to the target area detection information that the target shelf image has a top area and the target arrangement scene is an earthworm scene, determine that the shelf in the target shelf image is completely photographed; Otherwise, determine whether the shelf in the target shelf image is completely photographed according to the target arrangement scene and the number of target shelf layers.

18. The device according to claim 17, wherein, The integrity module is specifically configured to: Based on the association relationship between the preset candidate arrangement scenes and the candidate shelf layer number intervals, determine the target shelf layer number interval associated with the target arrangement scene; the target shelf layer number interval includes a first shelf layer number threshold and a second shelf layer number threshold; When the number of target shelf layers is greater than the first shelf layer number threshold, the shelf in the target shelf image is completely photographed; When the number of target shelf layers is less than the second shelf layer number threshold, the shelf in the target shelf image is not completely photographed.

19. A training device for a target area detection model, comprising: A teacher area detection module, configured to input an unlabeled second sample shelf image into a teacher area detection model to obtain pseudo area detection information of the second sample shelf image; The teacher area detection model is pre-trained using the labeled first sample shelf image; A sample fusion module, configured to fuse the first sample shelf image and the second sample shelf image to obtain a third sample shelf image; the third sample shelf image has a labeled detection frame and a labeled end type; A student area detection module, configured to input the third sample shelf image into a student area detection model to obtain a student detection frame and a student end type output by the student area detection model; An area detection training module, configured to train the student area detection model according to the labeled detection frame and the labeled end type of the third sample shelf image, and the student detection frame and the student end type, and use the training result as the target area detection model.

20. The apparatus according to claim 19, wherein, The area detection training module includes: A detection frame loss unit, configured to determine a detection frame loss of the student area detection model according to the labeled detection frame of the third sample shelf image and the student detection frame; An end loss unit, configured to determine an end type loss of the student area detection model according to the labeled end type of the third sample shelf image and the student end type; An area detection training unit, configured to train the student area detection model according to the detection frame loss of the student area detection model and the end type loss of the student area detection model, and use the training result as the target area detection model.

21. The device according to claim 19, further comprising: A teacher update unit for updating network parameters in the teacher region detection model based on the Exponential Moving Average (EMA) method according to the network parameters in the student region detection model.

22. A training device for a target scene classification model, comprising: An initial classification module for inputting an unlabeled second sample shelf image into an initial scene classification model to obtain a pseudo-display scene of the second sample shelf image; The initial scene classification model is pre-trained using labeled first sample shelf images; A sample fusion module for fusing the first sample shelf image and the second sample shelf image to obtain a third sample shelf image; the third sample shelf image has a labeled display scene; A display prediction module for inputting the third sample shelf image into the initial scene classification model to obtain a predicted display scene of the third sample shelf image; A scene classification training module for training the initial scene classification model according to the predicted display scene and the labeled display scene of the third sample shelf image, and using the training result as the target scene classification model.

23. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-11.

24. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause a computer to execute the method according to any one of claims 1-11.

25. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-11.

Citation Information

Patent Citations

  • Goods shelf obstacle recognition method, device and apparatus and readable storage medium

    CN110472486A

  • Commodity display position positioning method and device and storage medium

    CN110705424A