Shelf vacancy recognition method, machine learning model training method, and related devices

By using a pre-trained machine learning model and instance segmentation technology, the sub-areas of the shelf space below the empty space are identified and marked, which solves the problem of inaccurate empty space boundaries in the existing technology and improves the recognition accuracy.

CN116259011BActive Publication Date: 2026-01-02LENZTECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310274841.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-20
Publication Date
2026-01-02
Estimated Expiration
2043-03-20

AI Technical Summary

Technical Problem

In existing automatic shelf space identification technologies, the boundary markings of empty space areas are inaccurate, resulting in low identification accuracy.

Method used

By using a pre-trained machine learning model, the system identifies adjacent shelf sub-regions below empty shelf spaces, generates rectangular boundaries using instance segmentation techniques, and combines feature pyramid networks, region generation networks, and region of interest alignment methods to filter and label shelf sub-regions.

Benefits of technology

It improves the accuracy of empty shelf space identification and effectively avoids the problem of inaccurate identification results caused by the difficulty in determining the boundary line.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116259011B_ABST
    Figure CN116259011B_ABST
Patent Text Reader

Abstract

The application discloses a shelf vacancy recognition method, a machine learning model training method and related equipment. The shelf vacancy recognition method comprises the following steps: acquiring a first image of a shelf, the first image comprising at least one shelf board region; identifying a shelf board sub-region adjacent to the shelf vacancy in the first image; marking the boundary of the shelf board sub-region adjacent to the shelf vacancy; the shelf board sub-region is all or part of a shelf board region, and there is no commodity in the shelf region adjacent to the top of the shelf board sub-region. The application realizes the labeling of the shelf vacancy region by marking the shelf board sub-region adjacent to the bottom of the shelf vacancy in the first image. The image information of the shelf board is single and the shape of the region is fixed, which facilitates the determination of the region boundary line and the division of the region, effectively avoiding the situation that the accuracy of the shelf vacancy recognition result is low due to the difficulty in determining the region boundary line of the shelf vacancy region. Therefore, the application can effectively improve the recognition accuracy of the shelf vacancy.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of image processing, more particularly, to a shelf vacancy recognition method, a machine learning model training method and related equipment. BACKGROUND

[0002] With the development of technology, many intelligent technologies have been applied in the field of supermarket sales. One of them is automatic recognition of shelf vacancies through shelf images.

[0003] The existing technology for automatically recognizing shelf vacancies is to train a machine learning model using shelf images labeled with the boundaries of vacancy regions, and to use the trained machine learning model to recognize and mark the boundaries of vacancy regions in the shelf image. As shown in Figure 1 A and B are shelves, the black thick line indicates a shelf vacancy, and the shaded object is a commodity around the shelf vacancy.

[0004] The boundaries of the vacancy regions marked by the existing technology for automatically recognizing shelf vacancies are often inaccurate, resulting in low accuracy of recognizing shelf vacancy regions. SUMMARY

[0005] Therefore, the present application provides a shelf vacancy recognition method, a machine learning model training method and related equipment to solve the problem of low accuracy of recognizing shelf vacancies.

[0006] To achieve the above purpose, the present scheme is as follows:

[0007] A shelf vacancy recognition method applied to a pre-trained machine learning model, the method comprising:

[0008] Obtaining a first image of a shelf, the first image comprising at least one shelf board region;

[0009] Recognizing a shelf board sub-region adjacent to the shelf vacancy in the first image, and marking the boundaries of the shelf board sub-region adjacent to the shelf vacancy, wherein the shape of the boundaries is rectangular, the shelf board sub-region is all or part of the shelf board region, and there is no commodity in the shelf region adjacent to the top of the shelf board sub-region, and the training data of the machine learning model comprises a plurality of shelf images, the shelf images comprising at least one shelf board region and being labeled with the boundaries of the shelf board sub-region adjacent to the shelf vacancy.

[0010] Optionally, the recognition of the shelf board sub-region adjacent to the shelf vacancy in the first image comprises:

[0011] Dividing the first image to obtain at least one initial region;

[0012] identify image features in the initial region, the image features including: a shelf space above a shelf panel in the initial region and an image feature of the shelf panel in the initial region; and

[0013] perform a classification prediction on objects in the proposal region to obtain a classification result, the classification result being a class of each object in the proposal region;

[0014] determine at least one region in which an object of a shelf panel class in the proposal region is located as the panel sub-region.

[0015] Optionally, the identifying image features in the initial region and screening at least one proposal region from the initial region based on the identification result comprises:

[0016] classifying the initial region according to image features in the initial region to obtain a panel sub-region adjacent below a shelf space;

[0017] determining a classification confidence of the classified initial region, the classification confidence being a probability of the initial region having the panel sub-region;

[0018] arranging the initial regions in descending order of the classification confidence;

[0019] selecting a first preset number of initial regions from the initial regions in descending order of the arrangement as the proposal regions.

[0020] Optionally, the method further comprises:

[0021] calculating an overlap rate of a proposal region with the highest classification confidence and other proposal regions;

[0022] eliminating the proposal region with an overlap rate greater than an overlap rate threshold.

[0023] Optionally, the determining at least one region in which an object of a shelf panel class in the proposal region is located as the panel sub-region comprises:

[0024] determining a region in which an object of a shelf panel class in the proposal region is located as a candidate region;

[0025] determining a confidence of each candidate region, and screening a panel sub-region adjacent below a shelf space from the candidate regions based on the confidence.

[0026] Optionally, the dividing at least one initial region in the first image comprises:

[0027] perform feature extraction on the first image to obtain at least one image feature layer;

[0028] fuse the at least one image feature layer to obtain at least one fused feature layer;

[0029] generate at least one initial region of different sizes with each pixel point of the fused feature layer as a center point, and the initial region has a rectangular shape.

[0030] A machine learning model training method, the method comprising:

[0031] obtaining training data, the training data comprising: at least one shelf image, the shelf image comprising at least one shelf board region and the shelf image being labeled with a boundary of a shelf board sub-region adjacent to below a shelf vacancy, the shelf board sub-region being all or part of the shelf board region and a shelf region adjacent to above the shelf board sub-region being free of goods;

[0032] training a machine learning model based on the training data.

[0033] Optionally, the training of the machine learning model based on the training data comprises:

[0034] obtaining the shelf image from the training data and dividing the shelf image to obtain at least one initial region;

[0035] training the machine learning model based on a real region to filter a proposal region from the initial region, the real region being a region surrounded by the labeled boundary of the shelf board sub-region, and the proposal region being a region of the initial region containing the shelf board sub-region;

[0036] training the machine learning model based on the proposal region to classify specific objects and filter a candidate region;

[0037] training the machine learning model based on the candidate region to obtain a shelf board sub-region adjacent to below a shelf vacancy in the shelf image.

[0038] A shelf vacancy recognition device, the device comprising:

[0039] an obtaining unit configured to obtain a first image of a shelf, the first image comprising at least one shelf board region;

[0040] The segmentation unit is configured to identify a sub-rack region adjacent to the shelf vacancy in the first image, and mark a boundary of the sub-rack region adjacent to the shelf vacancy, wherein the boundary is in a rectangular shape, the sub-rack region is a whole or partial region of the rack region, and there is no commodity in the sub-rack region adjacent to the shelf region above the sub-rack region. The training data of the machine learning model comprises a plurality of shelf images, the shelf images comprising at least one rack region, and the shelf images being marked with the boundary of the sub-rack region adjacent to the shelf vacancy.

[0041] An electronic device includes a memory and a processor;

[0042] The memory is configured to store a program.

[0043] The processor is configured to execute the program to implement each step of any of the shelf vacancy identification methods and / or implement each step of any of the machine learning model training methods.

[0044] The present application can mark a sub-rack region below a shelf vacancy in an image through a shelf vacancy identification method and a machine learning model training method, and the marking of the sub-rack region represents the marking of the shelf vacancy region. The image information of the shelf rack is single and the shape of the region is fixed, which facilitates the determination of the region boundary line and the division of the region, and effectively avoids the situation that the accuracy of the shelf vacancy identification result is low due to the difficulty in determining the region boundary line of the shelf vacancy region. Therefore, the present application can effectively improve the identification accuracy of the shelf vacancy. BRIEF DESCRIPTION OF DRAWINGS

[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only embodiments of the present application, and those skilled in the art can obtain other drawings according to the provided drawings without creative labor.

[0046] Figure 1 A schematic diagram of shelf vacancy marking;

[0047] Figure 2 A flowchart of a shelf vacancy identification method provided by an embodiment of the present application;

[0048] Figure 3 A schematic diagram of a current shelf vacancy marking defect provided by an embodiment of the present application;

[0049] Figure 4 A schematic diagram of semantic segmentation and instance segmentation provided by an embodiment of the present application;

[0050] Figure 5 A schematic diagram of the mask principle provided for the embodiments of the present application is shown in FIG. 1.

[0051] Figure 6 A schematic diagram of the mask effect provided for the embodiments of the present application is shown in FIG. 2.

[0052] Figure 7 A flowchart of another shelf vacancy recognition method provided for the embodiments of the present application is shown in FIG. 3.

[0053] Figure 8 A flowchart of a machine learning model training method provided for the embodiments of the present application is shown in FIG. 4.

[0054] Figure 9 A structural schematic diagram of a shelf vacancy recognition device provided for the embodiments of the present application is shown in FIG. 5.

[0055] Figure 10 A hardware structural block diagram of an electronic device provided for the embodiments of the present application is shown in FIG. 6. DETAILED DESCRIPTION

[0056] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.

[0057] As shown in FIG. 1, the embodiments of the present application provide a shelf vacancy recognition method applied to a pre-trained machine learning model, which can include: Figure 2

[0058] S10, acquiring a first image of a shelf, the first image including at least one shelf board region;

[0059] S11, recognizing a shelf board sub-region adjacent to the below of a shelf vacancy in the first image, and marking a boundary of the shelf board sub-region adjacent to the below of the shelf vacancy, wherein the shape of the boundary is a rectangle, the shelf board sub-region is all or part of a shelf board region and there is no commodity in the shelf region adjacent to the above of the shelf board sub-region, and the training data of the machine learning model includes: a plurality of shelf images, the shelf images including at least one shelf board region and the shelf images being marked with the boundary of the shelf board sub-region adjacent to the below of the shelf vacancy.

[0060] ​The first image of the shelf can include goods, shelf vacancies and shelf shelves. The first image can be obtained by camera shooting or obtained from the monitoring file of the supermarket. The reason why the existing automatic identification of shelf vacancies has low accuracy in identifying vacancy regions is that the upper boundary of part of the vacancy region is not easy to determine, resulting in inaccurate boundary of the marked vacancy region. As shown in Figure 3 The upper part of the shelf B is empty, which makes it difficult to determine the upper boundary of the vacancy region (i.e. the black thick line region). Since the shape of the shelf shelf is fixed and easy to mark compared to the uncertainty of the shelf vacancy, the shelf vacancy can be marked by marking the shelf sub-region adjacent to the lower part of the shelf vacancy. The present embodiment can obtain the shelf sub-region adjacent to the lower part of the shelf vacancy by instance segmentation of the first image using a pre-trained machine learning model. Instance segmentation is an image segmentation technique based on deep learning, which essentially combines object detection and semantic segmentation to achieve the effect of instance segmentation. Semantic segmentation is also an image segmentation technique. The difference between semantic segmentation and instance segmentation is that semantic segmentation segments objects according to categories, using the same label for the same category, while instance segmentation uses different labels for different individuals of the same category. As shown in Figure 4 There are two clouds in an image, semantic segmentation will label the clouds according to categories, both clouds use the same label, which is white, while instance segmentation uses different labels for each cloud based on categories, one uses white label and the other uses black label. The present embodiment can generate a mask of the first image after instance segmentation of the first image of the shelf. The mask can also be called a mask, and the mask can have multiple layers, each layer representing a category. As shown in Figure 5 The mask is a matrix, and the computer recognizes the image as a matrix. The image is operated by placing a mask on the image, and the image matrix and the matrix represented by the mask are multiplied to obtain the desired result. As shown in Figure 6 There are two flowers in the image, and we only want one flower. We can separate one flower from the image by using a mask. Since the mask can have multiple layers when generated, only the layer of the mask with the category of the shelf shelf is needed in the present embodiment to separate the shelf sub-region adjacent to the lower part of the shelf vacancy from the first image, thereby marking the shelf shelf adjacent to the lower part of the shelf vacancy.

[0061] The embodiment can mark the sub-rail region below the shelf vacancy in the image, and the marking of the sub-rail region represents the marking of the shelf vacancy region. The image information of the shelf rail is single and the region shape is fixed, which facilitates the determination of the region boundary line and the division of the region, and effectively avoids the situation that the shelf vacancy recognition result is low in accuracy due to the difficulty in determining the region boundary line of the shelf vacancy region. Therefore, the method can effectively improve the recognition accuracy of the shelf vacancy.

[0062] As shown in Figure 7 Another shelf vacancy recognition method provided by the embodiment of the application can include the following steps. Figure 2 The process of identifying the adjacent sub-rail region below the shelf vacancy in the first image in step 11 can include:

[0063] S12, dividing the first image to obtain at least one initial region;

[0064] S13, identifying the image features in the initial region and selecting at least one suggestion region from the initial region based on the identification result, the image features including the shelf vacancy feature above the shelf rail in the initial region and the image feature of the shelf rail in the initial region;

[0065] S14, performing classification prediction on the objects in the suggestion region to obtain a classification result, the classification result being the category of each object in the suggestion region;

[0066] S15, determining at least one region in which the object with the category of the shelf rail in the suggestion region as the sub-rail region.

[0067] In the embodiment, the instance segmentation method used in the embodiment can mainly include three parts: a feature pyramid network (FPN), a region proposal network (RPN), and a region of interest align (ROIAlign). When generating the initial region, the first image needs to be first converted into multiple feature image layers by the FPN, and then the initial region is generated on each feature image layer by the RPN. The center point of each initial region can be a pixel point in the first image. The RPN can also classify the initial region according to the image features in the initial region to facilitate subsequent screening to obtain a proposal region. The proposal region obtained by screening is then sent to the ROIAlign part for normalization of the region size. Because the region size required for subsequent classification prediction and mask generation is fixed, the ROIAlign needs to normalize the size of the proposal region by bilinear interpolation. The instance segmentation will finally be divided into two branches, one branch is used for classification prediction, and the other branch is used for generating a mask for segmentation. In the embodiment, the proposal region can be normalized to two region sizes, one is 7x7, and the other is 14x14 (not that the proposal region is divided into two parts, one part is adjusted to 7x7, and the other part is adjusted to 14x14; but the size of the entire proposal region is adjusted to 7x7 and the size of the entire proposal region is adjusted to 14x14). The embodiment can perform classification prediction on the 7x7 proposal region and generate a mask for the 14x14 proposal region. The 14x14 proposal region can generate a 28x28x80 mask (28x28 is the size of the generated mask, and 80 represents that the mask has 80 categories, that is, there are 80 layers, and each layer corresponds to a category). In the embodiment, the two branches of the instance segmentation can not be performed simultaneously. In the embodiment, the classification result of the specific object in the proposal region can be obtained first, and then the corresponding mask is obtained based on the classification result. For example, the range of a proposal region is large, and it includes not only the shelf panel but also the goods C and D on the panel. For each good, classification prediction is performed, and the final classification prediction result is the shelf panel, the goods C, and the goods D. The mask generated according to the proposal region has three layers, one layer represents the mask of the shelf panel, one layer represents the mask of the goods C, and one layer represents the mask of the goods D. In this way, the mask representing the shelf panel can be extracted from the proposal region according to the category.

[0068] According to another shelf vacancy identification method provided in the embodiment of the application, Figure 7 The step S13 can include steps one to four:

[0069] Step one: classify the initial region according to the image features in the initial region, and obtain the adjacent shelf plate sub-region under the shelf vacancy;

[0070] Step two: determine the classification confidence of the classified initial region, and the classification confidence is the probability of the initial region having the shelf plate sub-region;

[0071] Step three: arrange the initial regions from high to low according to the classification confidence;

[0072] Step four: select the first preset number of initial regions as the suggestion region from the initial region in descending order.

[0073] In the embodiment, the initial region is classified according to the image features in the region identified in the initial region. The classification is not specifically implemented to the specific category of the object contained in the initial region, but is to determine whether the initial region contains the shelf plate sub-region according to whether the initial region has the image features of the shelf vacancy and the image features of the adjacent shelf plate under the shelf vacancy. If it contains, it is classified into one category, and if it does not contain, it is classified into another category. In the embodiment, the initial region is classified by RPN. RPN classification is a binary classification, that is, the initial region is classified into foreground and background. The foreground indicates that the initial region may have a shelf plate sub-region, and the other is background. After classifying the initial region in the embodiment, the classification result of each initial region is scored by a softmax function. The score is a probability, that is, the classification confidence in the embodiment. The classification confidence represents the probability of the initial region having the shelf plate sub-region, and also represents the degree of trust in the RPN binary classification. Since the number of initial regions is too large and a large part of the initial regions is not needed in the subsequent processing process, the classification confidence is used for the first initial region screening. In the embodiment, the initial regions are sorted by the classification confidence, and the first preset number of initial regions are selected from high to low. The first preset number in the embodiment is 2000.

[0074] According to another shelf vacancy identification method provided in the embodiment of the application, the method can further include steps five to six:

[0075] Step five: calculate the overlap rate of the suggestion region with the highest classification confidence and other suggestion regions;

[0076] Step six: remove the suggestion region with an overlap rate greater than the overlap rate threshold.

[0077] Wherein, since the number of initial regions is still too large after the first screening of the initial regions by the classification confidence, the first preset number of initial regions can be screened for the second time by the non-maximum suppression (NMS), and the second screening is mainly aimed at the case that multiple initial regions may contain the adjacent shelf board sub-regions under one shelf space. In the second screening, a preset overlap rate threshold is used to remove the multiple redundant initial regions of the adjacent shelf board sub-regions under one shelf space. In the second screening, the initial region with the highest classification confidence in the first preset number is selected first, the overlap rate between the initial region and all other initial regions is calculated, and if the overlap rate is greater than the preset overlap rate threshold, the initial region is removed. In addition to the initial region with the highest classification confidence, the initial region with the second highest classification confidence is selected from the remaining initial regions, the overlap rate between the initial region and other initial regions is calculated again, and if the overlap rate is greater than the preset overlap rate threshold, the initial region is removed. In this way, the second screening of the first preset number of initial regions is completed. After the screening by the NMS, the initial regions are sorted again by the classification confidence, and a part of the initial regions are obtained according to the order. In the embodiment, the number of the final obtained initial regions is 300.

[0078] According to another shelf space identification method provided by the embodiment of the application, Figure 7 The step S15 can include steps seven to eight:

[0079] Step seven: determining the region where the object with the shelf board category in the proposed region as the candidate region;

[0080] Step eight: determining the confidence of each candidate region, and screening the adjacent shelf board sub-regions under the shelf space from the candidate regions based on the confidence.

[0081] Wherein, after obtaining the category of each object in the proposed region, the corresponding mask of the shelf board category is obtained (since the mask generated by the proposed region in the embodiment can be multi-layer, one category is one layer), and the mask region is the candidate region in the embodiment. Since there can be a case that multiple masks correspond to one adjacent shelf board sub-region under the shelf space when the mask is obtained, the NMS is needed for the final screening. When the screening is performed, a preset overlap rate threshold is needed, which is different from the overlap rate threshold preset in the second screening. The specific screening operation is also to select the proposed region with the highest confidence in the current proposed region, calculate the overlap rate between the proposed region and other proposed regions, and if the overlap rate is greater than the overlap rate threshold, the proposed region is removed.

[0082] According to another shelf space identification method provided by the embodiment of the application, Figure 7The step S12 shown can include steps nine to eleven:

[0083] Step nine: feature extraction is performed on the first image to obtain at least one image feature layer;

[0084] Step ten: at least one image feature layer is fused to obtain at least one fused feature layer;

[0085] Step eleven: at least one initial region of different sizes is generated with each pixel point of the fused feature layer as a center point, and the shape of the initial region is a rectangle.

[0086] In this embodiment, the FPN uses the feature extraction network ResNeXt101 to extract the image feature layer. Different levels of features of the image feature layer have different expression abilities. The shallow features mainly reflect details such as brightness and edges, and the deep features reflect more rich overall structures. The fusion between the image feature layers in this embodiment is to add each corresponding element in the upper image feature layer after upsampling and 1x1 convolution to the original lower image feature layer. In order to prevent the problem of insufficient fusion, this embodiment will use a 3x3 convolution to smooth the fused feature layer after the image feature layer fusion, thereby obtaining a fused feature layer with more sufficient fusion. The multiple fused feature layers in this embodiment include both shallow features and deep features, and have rich expression abilities. Different levels of features of the feature map have different expression abilities. The shallow features mainly reflect details such as brightness and edges, and the deep features reflect more rich overall structures. When generating the initial region on each fused feature layer by the RPN, this embodiment predefines three areas of initial regions, and each predefined area corresponds to three fixed length-width ratios. Therefore, nine different initial regions can be generated with a pixel point as the center. The generation method of the initial region can be to generate each pixel as the center, or to set a threshold and generate the initial region with every threshold pixel point as the center. After obtaining multiple proposal regions by the RPN, the proposal regions need to be sent to the ROIAlign part for region size normalization. Since the input requirement of the ROIAlign part is the image feature layer, before sending the proposal region to the ROIAlign part, the image feature layer to which the proposal region belongs needs to be determined. In this embodiment, the image feature layer to which the proposal region belongs is determined by the formula:

[0087]

[0088] The image feature layer to which each region belongs is calculated. k represents the number of image feature layers to which the proposed region belongs, w represents the width of the proposed region, h represents the height of the proposed region, k0 represents the number of image feature layers mapped when the width of the proposed region is 224 and the height of the proposed region is 224, and in this embodiment, the value of k0 is 4. For example, if there is a proposed region of 112*112, the value of k obtained by substituting the formula is 3, that is, the proposed region belongs to the third image feature layer, and the third image feature layer is taken as the input of the ROIAlign part.

[0089] As shown in Figure 8 The embodiment of the present application also provides a training method of a machine learning model, which can include:

[0090] S16, obtaining training data, the training data including: at least one shelf image, the shelf image including at least one baffle region and the shelf image being labeled with the boundary of a baffle sub-region adjacent to the shelf space below, the baffle sub-region being all or part of a baffle region and there being no goods in the shelf region adjacent to the baffle sub-region above;

[0091] S17, training the machine learning model based on the training data.

[0092] Wherein, the classification, screening and mask generation of the machine learning model are trained through the data of the labeled baffle sub-region, that is, the instance segmentation of the machine learning model is trained according to the data.

[0093] According to another training method provided in the embodiment of the present application, Figure 8 As shown in step S17, the method can include steps twelve to fifteen:

[0094] Step twelve: obtaining a shelf image from the training data and dividing to obtain at least one initial region;

[0095] Step thirteen: screening a proposed region from the initial region based on a real region to train the machine learning model, the real region being a region surrounded by the labeled baffle sub-region boundary, and the proposed region being a region containing the baffle sub-region in the initial region;

[0096] Step fourteen: classifying specific objects and screening candidate regions based on the proposed region to train the machine learning model;

[0097] Step fifteen: obtaining the baffle sub-region adjacent to the shelf space below in the shelf image based on the candidate region to train the machine learning model.

[0098] Wherein, in the embodiment, when labeling the shelf board sub-region in the shelf image, the shelf image needs to be marked by a labeling tool labelme, and the final marking result is a multi-layer mask, which is used as a training sample of the model prediction mask. The number of layers of the mask is determined according to the objects contained in the shelf image, one object represents one category, and one category represents one layer of the mask. In the embodiment, when generating the real region, the labeling tool can calculate the minimum bounding rectangle of the mask while generating the mask, and the rectangle is used as the real region in the shelf image. The shape of the real region is not limited, and a rectangle is used in the embodiment. In the embodiment, when generating the initial region for training the shelf image, a plurality of initial regions are generated in the shelf image with each pixel point on each layer of the fusion feature layer as the center. In the embodiment, when obtaining the training classification sample, the overlap rate of each initial region and the real region can be calculated first, and then the initial region is divided into positive samples and negative samples according to the overlap rate. If the overlap rate is greater than 0.7, the initial region is assigned a positive sample label; if the overlap rate is less than 0.3, the initial region is assigned a negative sample label, and the remaining initial regions are not considered. Because the number of samples is too large, 128 positive samples and 128 negative samples are randomly selected from the positive samples and the negative samples for training of the model during the training process of the model. After obtaining the positive samples and the negative samples, the real region is used as a supervision label to continuously train the RPN two-classification of the model. In the embodiment, the process of training the classification is to let the model learn the image features in the initial region that may have a shelf board sub-region, so as to perform two-classification on the initial region based on the image features in the subsequent test link. In addition to training the two-classification ability of the model, the regression of the model also needs to be trained. The initial region regression is to find a relationship, so that the position of the initial region is mapped to a region that is closer to the position of the real region through the relationship. In the embodiment, the model classification is trained based on the suggestion region obtained through the last screening, and the classification is specific to the category of each object in the suggestion region. The model mask prediction is trained through the comparison between the prediction mask in the candidate region and the real mask.

[0099] The training method of the machine learning model provided in the embodiment can effectively avoid the situation that the accuracy of the shelf space recognition result is low due to the difficulty in determining the region boundary line of the shelf space region, and effectively improve the recognition accuracy of the trained machine learning model in recognizing the shelf space. Corresponding to the shelf space recognition method provided in the embodiment, the application also provides a shelf space recognition device.

[0100] AsFigure 9 As shown, the shelf vacancy recognition device provided by the embodiments of the present application can include:

[0101] The acquisition unit 100 is configured to acquire a first image of a shelf, and the first image includes at least one shelf board region.

[0102] The segmentation unit 110 is configured to identify a shelf board sub-region adjacent to a shelf vacancy in the first image, and mark a boundary of the shelf board sub-region adjacent to the shelf vacancy, wherein the boundary is in a rectangular shape, the shelf board sub-region is all or part of a shelf board region, and there is no commodity in a shelf region adjacent to the shelf board sub-region above the shelf board sub-region, and the training data of the machine learning model includes a plurality of shelf images, the shelf images include at least one shelf board region, and the shelf images are marked with the boundary of the shelf board sub-region adjacent to the shelf vacancy.

[0103] In another shelf vacancy recognition device provided by the embodiments of the present application, the segmentation unit 110 can include an identification sub-unit and a marking sub-unit,

[0104] The marking sub-unit is configured to mark the boundary of the shelf board sub-region adjacent to the shelf vacancy.

[0105] The identification sub-unit can include:

[0106] The division sub-unit is configured to divide the first image to obtain at least one initial region.

[0107] The region screening sub-unit is configured to identify an image feature in the initial region and screen at least one proposal region from the initial region based on an identification result, and the image feature includes a shelf vacancy feature above a shelf board of the initial region and an image feature of the shelf board of the initial region.

[0108] The classification prediction sub-unit is configured to perform classification prediction on an object in the proposal region to obtain a classification result, and the classification result is a category of each object in the proposal region.

[0109] The region obtaining sub-unit is configured to determine at least one region in which an object of a category of a shelf board is located in the proposal region as the shelf board sub-region.

[0110] In another shelf vacancy recognition device provided by the embodiments of the present application, the region screening sub-unit can include:

[0111] The region classification sub-unit is configured to classify the initial region according to the image feature in the initial region to obtain the shelf board sub-region adjacent to the shelf vacancy.

[0112] a classification confidence subunit configured to determine a classification confidence of the initial region obtained by the classification, the classification confidence being a probability that the initial region has the sub-region of the shelf panel;

[0113] a sorting subunit configured to arrange the initial regions in descending order of the classification confidence;

[0114] a region selection subunit configured to select a first preset number of initial regions from the initial regions as the suggested regions in descending order of the arrangement.

[0115] According to another shelf vacancy recognition device provided by the embodiments of the present application, the device can further include:

[0116] a calculation unit configured to calculate an overlap rate of a suggested region with the highest classification confidence and other suggested regions;

[0117] a rejection unit configured to reject the suggested region with an overlap rate greater than an overlap rate threshold.

[0118] According to another shelf vacancy recognition device provided by the embodiments of the present application, the region obtaining subunit can include:

[0119] a candidate region subunit configured to determine a region in which an object with a category of a shelf panel is located in the suggested region as a candidate region;

[0120] a confidence subunit configured to determine a confidence of each candidate region and filter a sub-region of a shelf panel adjacent to a shelf vacancy from the candidate regions based on the confidence.

[0121] According to another shelf vacancy recognition device provided by the embodiments of the present application, the division subunit can include:

[0122] a feature extraction subunit configured to perform feature extraction on the first image to obtain at least one image feature layer;

[0123] a fusion subunit configured to fuse the at least one image feature layer to obtain at least one fused feature layer;

[0124] a region generation subunit configured to generate at least one initial region of different sizes with each pixel point of the fused feature layer as a center point, the initial region being in a rectangular shape.

[0125] The shelf vacancy recognition device provided by the application can mark the sub-rail region below the shelf vacancy in the image, and the marking of the sub-rail region reflects the marking of the shelf vacancy region. The image information of the shelf rail is single and the region shape is fixed, so that the region boundary line can be determined and the region can be divided, and the situation that the shelf vacancy recognition result is low in accuracy due to the difficulty in determining the region boundary line of the shelf vacancy region can be effectively avoided. Therefore, the application can effectively improve the recognition accuracy of the shelf vacancy.

[0126] Corresponding to the machine learning model training method provided by the embodiment of the application, the application further provides a machine learning model training device.

[0127] The machine learning model training device can include:

[0128] The training data obtaining unit is configured to obtain training data, and the training data includes at least one shelf image, the shelf image includes at least one rail region, and the shelf image is marked with the boundary of the adjacent sub-rail region below the shelf vacancy. The sub-rail region is all or part of a rail region, and there is no commodity in the adjacent shelf region above the sub-rail region.

[0129] The training unit is configured to train the machine learning model based on the training data.

[0130] Optionally, the training unit can be specifically configured to:

[0131] Obtain at least one initial region by dividing the shelf image obtained from the training data; filter the suggestion region from the initial region based on the real region and the machine learning model, the real region is the region surrounded by the marked sub-rail region boundary, and the suggestion region is the region containing the sub-rail region in the initial region; train the machine learning model based on the suggestion region to classify specific objects and filter the candidate region; and train the machine learning model based on the candidate region to obtain the adjacent sub-rail region below the shelf vacancy in the shelf image.

[0132] The embodiment of the application further provides an electronic device, Figure 10 The hardware structure block diagram of the electronic device is shown, and the hardware structure of the electronic device can include a memory 1 and a processor 2 Figure 10

[0133] The memory 1 is configured to store a program.

[0134] The processor 2 is configured to execute the program, implement each step of any of the above shelf vacancy recognition methods, and / or implement each step of any of the above machine learning model training methods.

[0135] ​The computer program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other processing devices to cause a series of operational steps to be performed on the computer, other programmable apparatus or other processing devices to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions specified in the flowchart block or blocks. Figure 1 The flowchart and / or block diagram in the Figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various embodiments. In this regard, each block in the flowchart and / or block diagrams can represent a module, segment, or portion of code, which comprises one or more executable Figure 1 The flowchart and / or block diagram in the Figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various embodiments. In this regard, each block in the flowchart and / or block diagrams can represent a module, segment, or portion of code, which comprises one or more executable

[0136] In one typical configuration, the device includes one or more processors (CPUs), memory, and a bus. The device can also include an input / output interface, a network interface, and the like.

[0137] The memory can include non-persistent memory and / or volatile memory, e.g., random access memory (RAM) comprising a number of memory locations that can be read and / or written on the fly. The memory can also include non-volatile memory, e.g., read-only memory (ROM), electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), flash memory (flash RAM), and the like. The memory includes at least one memory chip. The memory is an example of computer-readable media.

[0138] Computer-readable media includes permanent and non-permanent, removable and non-removable media implemented in any method or technology for storage of information such as computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible to computing devices. According to the definition herein, computer-readable media does not include transitory media such as modulated data signals and carriers.

[0139] Those skilled in the art will appreciate that embodiments of the present application can be devised for a variety of other systems which are currently developed or later developed. Practitioners skilled in the art will recognize the equivalents of the various features from the preceding description and drawings. There is no intent, therefore, to limit the scope of the present application to these embodiments. The intent is to cover all alternatives, modifications and equivalents.

[0140] It is to be understood that the terms "including", "comprising", "consisting" and "consisting essentially of" to be used in the specification in their broadest sense of meaning that they allow for instances where undertaking an adding, an inclusion, a combination or a collection of any two or more of the instances of functional operations of an entity or operational steps of a process to instantiate an embodiment of the process, an item, a system, or an apparatus. As used herein, "and / or" means and.

[0141] Each of the embodiments described in the specification illustrates by way of example a method, system, or computer program product that can be implemented as a process, method of processing, accomplished on or using one or more computer systems. Each of the disclosed embodiments can be implemented as one or more computer programs or program components running in or on one or more systems, which embody the functionality for the application. The embodiments disclosed herein can each also be implemented using the components, processes, and data stores described herein and / or through the use of devices which are remotely located from each other and from the respective embodiments, wherein some or all data may

[0142] The foregoing description of various embodiments of the application has been presented for the purposes of illustration and description. It is not intended to be exhaustive or to limit the application to the precise form disclosed. Many modifications and variations are possible in light of the above teaching. It is intended that the scope of the application be limited not with this detailed description, but rather by the claims appended hereto.

Claims

1. A method of shelf space identification, characterized in that, The method is applied to a pre-trained machine learning model, and the method comprises: obtaining a first image of a shelf, the first image comprising at least one baffle region; identifying a baffle sub-region adjacent to a shelf vacancy in the first image, and marking a boundary of the baffle sub-region adjacent to the shelf vacancy, wherein the boundary is in the shape of a rectangle, the baffle sub-region is all or part of the baffle region, and there is no product in a shelf region adjacent to the baffle sub-region, and training data of the machine learning model comprises: a plurality of shelf images, the shelf images comprising at least one baffle region and the shelf images being marked with the boundary of the baffle sub-region adjacent to the shelf vacancy; the identification of the baffle sub-region adjacent to the shelf vacancy in the first image comprises: dividing the first image to obtain at least one initial region; identifying image features in the initial region and screening at least one proposal region from the initial region based on the identification result, wherein the image features comprise: a shelf vacancy feature above a shelf baffle of the initial region and an image feature of the shelf baffle of the initial region; performing classification prediction on objects in the proposal region to obtain a classification result, wherein the classification result is a category of each object in the proposal region; determining at least one region in which an object of the category of a shelf baffle is located in the proposal region as the baffle sub-region.

2. The method of claim 1, wherein, the identification of the image features in the initial region and the screening of at least one proposal region from the initial region based on the identification result comprises: classifying the initial region according to the image features in the initial region to obtain a baffle sub-region adjacent to a shelf vacancy; determining a classification confidence of the classified initial region, wherein the classification confidence is a probability that the initial region has the baffle sub-region; arranging the initial regions from high to low according to the classification confidence; selecting a first preset number of initial regions from the initial regions in descending order as the proposal regions.

3. The method of claim 2, wherein, further comprising: calculating an overlap rate of the proposal region with the highest classification confidence and other proposal regions; eliminating the proposal regions with an overlap rate greater than an overlap rate threshold.

4. The method of claim 1, wherein, the determination of at least one region in which an object of the category of a shelf baffle is located in the proposal region as the baffle sub-region comprises: determining a candidate region in which an object of the category of a shelf baffle is located in the proposal region; determining a confidence of each candidate region, and screening a baffle sub-region adjacent to a shelf vacancy from the candidate regions based on the confidence.

5. The method of claim 1, wherein, the division of the first image to obtain at least one initial region comprises: performing feature extraction on the first image to obtain at least one image feature layer; fusing the at least one image feature layer to obtain at least one fused feature layer; generating at least one initial region of different sizes with each pixel point of the fused feature layer as a center point, wherein the initial region is in the shape of a rectangle. 6.A machine learning model training method, characterized in that, the method comprises: Obtaining training data, the training data comprising: at least one shelf image, the shelf image comprising at least one baffle region and the shelf image being labeled with a boundary of a baffle sub-region adjacent to below a shelf vacancy, the baffle sub-region being all or part of one of the baffle regions and there being no goods in a shelf region adjacent to above the baffle sub-region; Obtaining the shelf image from the training data and dividing to obtain at least one initial region; Training the machine learning model based on a true region to filter a proposal region from the initial region, the true region being a region enclosed by the labeled boundary of the baffle sub-region, and the proposal region being a region of the initial region containing the baffle sub-region; Training the machine learning model based on the proposal region to perform specific object classification and filter a candidate region; Training the machine learning model based on the candidate region to obtain a baffle sub-region adjacent to below a shelf vacancy in the shelf image.

7. A shelf space identification apparatus, characterized by, The apparatus comprises: An obtaining unit configured to obtain a first image of a shelf, the first image comprising at least one baffle region; A dividing unit configured to identify a baffle sub-region adjacent to below a shelf vacancy in the first image and mark a boundary of the baffle sub-region adjacent to below the shelf vacancy, wherein the boundary is in the shape of a rectangle, the baffle sub-region is all or part of one of the baffle regions and there is no goods in a shelf region adjacent to above the baffle sub-region, and training data of a machine learning model comprises: a plurality of shelf images, the shelf images comprising at least one baffle region and the shelf images being labeled with a boundary of a baffle sub-region adjacent to below a shelf vacancy; The dividing unit comprises an identifying sub-unit and a marking sub-unit; The marking sub-unit is configured to mark the boundary of the baffle sub-region adjacent to below the shelf vacancy; The identifying sub-unit comprises a dividing sub-unit, a region filtering sub-unit, a classification prediction sub-unit, and a region obtaining sub-unit: The dividing sub-unit is configured to divide the first image to obtain at least one initial region; The region filtering sub-unit is configured to identify image features in the initial region and filter at least one proposal region from the initial region based on the identification result, the image features comprising: a shelf vacancy feature above a shelf baffle of the initial region and an image feature of the shelf baffle of the initial region; The classification prediction sub-unit is configured to perform classification prediction on objects in the proposal region to obtain classification results, the classification results being categories of each object in the proposal region; The region obtaining sub-unit is configured to determine at least one region in which an object of a category of a shelf baffle in the proposal region as the baffle sub-region.

8. An electronic device, comprising: Comprise a memory and a processor; The memory is configured to store a program; The processor is configured to execute the program to implement each step of the shelf vacancy identification method according to any one of claims 1-5, and / or implement each step of the machine learning model training method according to claim 6.

Citation Information

Patent Citations

  • Goods shelf obstacle recognition method, device and apparatus and readable storage medium

    CN110472486A

  • Commodity display analysis method, device and equipment and storage medium

    CN112990095A