An article detection and identification method and device, an electronic device, and a storage medium

CN117935128BActive Publication Date: 2026-09-29GUANGZHOU HUINA DIGITAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311480399.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-08
Publication Date
2026-09-29
Estimated Expiration
2043-11-08

AI Technical Summary

Technical Problem

[0005]本发明提供一种物品检测识别方法、装置、电子设备及存储介质,用于解决现有技术中摄像头移动带来的画面偏移和旋转引起的检测识别精度损失,以及每张待测图需要检测并识别所有物品带来的算力损失的问题

Benefits of technology

[0018]1、本发明中,通过对待测货架的标准图像和待测图像的矫正图像分别进行图像预处理,并分别提取相似特征图对应生成第一特征图和第二特征图;其中,所述标准图像为所述待测货架的物品陈列变化前的图像,所述待测图像为所述待测货架的物品陈列变化后的图像;对所述第一特征图和所述第二特征图中的对应位置相减并取绝对值,生成第三特征图;基于所述第三特征图进行变化区域物品检测,得到变化区域物品的基础信息;基于所述变化区域物品的基础信息进行物品识别;能够有效的降低货架所需要识别的物品数量,从而降低系统计算量。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117935128B_ABST
    Figure CN117935128B_ABST
Patent Text Reader

Abstract

The application provides an article detection and identification method and device, electronic equipment and medium. The method comprises: performing image preprocessing on a standard image of a to-be-detected shelf and a rectified image of a to-be-detected image, respectively, and extracting similar feature maps to generate a first feature map and a second feature map; subtracting corresponding positions in the first feature map and the second feature map and taking an absolute value to generate a third feature map; performing change area article detection based on the third feature map to obtain basic information of the change area article; and performing article identification based on the basic information of the change area article. The application can effectively reduce the loss of detection and identification accuracy caused by picture deviation and rotation caused by camera movement, and the loss of computing power caused by the need to detect and identify all articles in each to-be-detected image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of image processing technology, and relates to an object detection and recognition method, and particularly to an object detection and recognition method, device, electronic device and storage medium. Background Technology

[0002] Store display identification plays a vital role in the retail industry. It not only assists in inventory management and timely responses to stockouts or surpluses by observing the merchandise displayed on shelves, but also effectively analyzes and optimizes merchandise displays to improve customer shopping experience and purchase intentions, thereby increasing sales.

[0003] Traditional in-store merchandise display inspection mainly relies on manual on-site checks or online manual checks based on acquired images. This method requires a significant investment of time and manpower, and is not only inefficient but also susceptible to human error leading to low accuracy.

[0004] Existing computer vision-based solutions primarily rely on image-based object detection and metric-based object recognition. After obtaining the bounding box location of an item, the specific name is identified. However, this method requires the camera angle and position to be as fixed as possible; rotation of the camera angle or shift in position affects the accuracy of detection and recognition. Furthermore, because each image is inspected and each detected item is identified, there is significant waste of computational resources when the item arrangement remains unchanged or changes only slightly. Summary of the Invention

[0005] This invention provides an object detection and recognition method, device, electronic device, and storage medium to solve the problems of loss of detection and recognition accuracy caused by image offset and rotation due to camera movement in the prior art, as well as the loss of computing power caused by the need to detect and recognize all items for each image to be tested.

[0006] In a first aspect, the present invention provides an item detection and recognition method, wherein the item detection and recognition method is used in a trained change region target detection model, the method comprising: performing image preprocessing on a standard image of a shelf to be tested and a corrected image of the image to be tested, and extracting similar feature maps to generate a first feature map and a second feature map respectively; wherein the standard image is an image of the shelf to be tested before the item arrangement changes, and the image to be tested is an image of the shelf to be tested after the item arrangement changes; subtracting corresponding positions in the first feature map and the second feature map and taking the absolute value to generate a third feature map; performing change region item detection based on the third feature map to obtain basic information of the items in the change region; and performing item recognition based on the basic information of the items in the change region.

[0007] In this invention, image preprocessing is performed on a standard image and a corrected image of the shelf to be tested, respectively, and similar feature maps are extracted to generate a first feature map and a second feature map. The standard image is the image of the shelf before the item arrangement changes, and the image to be tested is the image of the shelf after the item arrangement changes. A third feature map is generated by subtracting corresponding positions from the first and second feature maps and taking the absolute value. Based on the third feature map, item detection in the changed area is performed to obtain basic information about the items in the changed area. Item recognition is then performed based on this basic information. This effectively reduces the number of items that need to be recognized on the shelf, thereby reducing the system's computational load.

[0008] In one implementation of the first aspect, the step of detecting items in changed regions based on the third feature map to obtain basic information about the items in changed regions includes: performing changed region detection on the third feature map to obtain a change probability map; performing item detection on the third feature map to obtain basic information about the items; and obtaining basic information about the items in changed regions based on the change probability map and the basic information about the items.

[0009] In one implementation of the first aspect, the step of performing change region detection on the third feature map to obtain a change probability map includes: obtaining the change probability of the change probability map; if the change probability is close to 1, the item in the third feature map has changed; if the change probability is close to 0, the item in the third feature map has not changed.

[0010] In one implementation of the first aspect, obtaining the basic information of the item in the change region based on the change probability map and the item's basic information includes: obtaining the item's location based on the item's basic information; judging the item's location and the change region in the change probability map; if the item's location is consistent with the change region in the change probability map, then retaining the item's basic information; otherwise, not retaining it.

[0011] In one implementation of the first aspect, the method for obtaining the corrected image of the image to be tested includes: obtaining the image to be tested based on the shelf to be tested; obtaining the corresponding point coordinates and consistency probability of the image to be tested and the standard image of the shelf to be tested by calling an image matching model based on the image to be tested and the standard image of the shelf to be tested; selecting the N sets of corresponding point coordinates with the highest consistency probability as reference coordinates of the image to be tested and the standard image of the shelf to be tested; calculating a transformation matrix based on the image to be tested, the standard image of the shelf to be tested, and the respective reference coordinates; and performing a perspective transformation on the image to be tested according to the transformation matrix to obtain the corrected image of the image to be tested.

[0012] In one implementation of the first aspect, the method for acquiring the standard image of the shelf to be tested includes: establishing a standard shelf image library; taking pictures of the standard display shelf using a camera or a webcam to obtain a standard display shelf image; marking the location area to which the standard display shelf image belongs to obtain the area information of the standard display shelf image; recording the standard display shelf image and the area information of the standard display shelf image into the standard shelf image library; and obtaining a standard image of the same shelf as the image to be tested from the standard shelf image library.

[0013] In one implementation of the first aspect, the item identification based on the basic information of the items in the changing area includes: extracting features from the items in the changing area to obtain features of the item to be identified; matching and querying the features of the item to be identified with the features of items in a preset standard library one by one to obtain the best matching result; and obtaining item information based on the best matching result.

[0014] Secondly, the present invention provides an item detection and recognition device, comprising: a feature map generation module, used to preprocess a standard image of a shelf to be tested and a corrected image of the image to be tested, and extract similar feature maps to generate a first feature map and a second feature map respectively; wherein, the standard image is an image of the shelf to be tested before the item display changes, and the image to be tested is an image of the shelf to be tested after the item display changes; an image stitching module, used to subtract corresponding positions in the first feature map and the second feature map and take the absolute value to generate a third feature map; a change area item detection module, used to detect change area items based on the third feature map to obtain basic information of change area items; and a change area item recognition module, used to recognize items based on the basic information of change area items.

[0015] Thirdly, the present invention provides an electronic device, the electronic device comprising: a memory for storing a computer program; and a processor for executing the computer program stored in the memory to perform the above-described item detection and identification method.

[0016] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by an electronic device, implements the above-described item detection and identification method.

[0017] As described above, the article detection and identification method, device, electronic device, and storage medium of the present invention have the following beneficial effects:

[0018] 1. In this invention, image preprocessing is performed on the standard image and the corrected image of the shelf to be tested, respectively, and similar feature maps are extracted to generate a first feature map and a second feature map. The standard image is the image of the shelf before the item arrangement changes, and the image to be tested is the image of the shelf after the item arrangement changes. A third feature map is generated by subtracting the corresponding positions in the first and second feature maps and taking the absolute value. Based on the third feature map, item detection in the changed area is performed to obtain basic information about the items in the changed area. Item recognition is then performed based on the basic information about the items in the changed area. This effectively reduces the number of items that need to be recognized on the shelf, thereby reducing the system's computational load.

[0019] 2. This invention obtains a corrected image of the image under test by performing correction processing on the image under test, which can effectively eliminate or reduce errors such as camera offset or shooting angle rotation. Attached Figure Description

[0020] Figure 1 The diagram shown is a schematic representation of the item detection and identification system according to an embodiment of the present invention.

[0021] Figure 2 The diagram shown is a schematic flowchart of the item detection and identification method according to an embodiment of the present invention.

[0022] Figure 3 The diagram shown is a flowchart illustrating the method for acquiring a corrected image of the image to be tested according to an embodiment of the present invention.

[0023] Figure 4 The diagram shown is a schematic flowchart of the standard image acquisition method for the shelf under test as described in an embodiment of the present invention.

[0024] Figure 5A The diagram shows a flowchart of the method for obtaining basic information of items in the changing area as described in an embodiment of the present invention.

[0025] Figure 5B The diagram shown is a schematic flowchart of the change region detection method described in an embodiment of the present invention.

[0026] Figure 5C This is a schematic diagram of the method for obtaining basic information of items in a changing area according to another embodiment of the present invention.

[0027] Figure 6 The diagram shown is a schematic of the change area item detection function described in an embodiment of the present invention.

[0028] Figure 7A The image shown is a standard schematic diagram of the shelf to be tested according to an embodiment of the present invention.

[0029] Figure 7BThe image shown is a schematic diagram of the image to be tested according to an embodiment of the present invention.

[0030] Figure 7C The diagram shown is a schematic representation of a binary label map of the changed region of the image under test, as described in an embodiment of the present invention.

[0031] Figure 8 The diagram shown is a schematic representation of the item recognition process according to an embodiment of the present invention.

[0032] Figure 9 The diagram shown is a schematic representation of the object detection and identification device according to an embodiment of the present invention.

[0033] Figure 10 The diagram shown is a schematic representation of the electronic device structure according to an embodiment of the present invention.

[0034] Component designation explanation

[0035] 1. Item Detection and Identification System

[0036] 10 Image Acquisition Module

[0037] 20 Image Correction Module

[0038] 30. Change Area Detection Module

[0039] 40 Item Recognition Module

[0040] 100 Item Detection and Identification Device

[0041] 110 Feature Map Generation Module

[0042] 120 Image Stitching Module

[0043] 130 Item Detection Module for Changing Areas

[0044] 140 Item Recognition Module for Changing Areas

[0045] 200 electronic devices

[0046] 210 Memory

[0047] 220 processor

[0048] Steps S1 to Sn Detailed Implementation

[0049] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, unless otherwise specified, the following embodiments and features described therein can be combined with each other.

[0050] It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Therefore, the drawings only show the components related to the present invention and are not drawn according to the actual number, shape and size of the components in the actual implementation. In the actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.

[0051] The following embodiments of the present invention provide an item detection and recognition method, device, electronic device and storage medium, which solve the problems of loss of detection and recognition accuracy caused by image offset and rotation due to camera movement in the prior art, as well as the loss of computing power caused by the need to detect and recognize all items for each image to be tested.

[0052] like Figure 1 As shown, this embodiment provides an item detection and recognition system for executing the item detection and recognition method described in this embodiment of the invention. The item detection and recognition system 1 includes an image acquisition module 10, an image correction module 20, a change area detection module 30, and an item recognition module 40.

[0053] The image acquisition module 10 is used to acquire the image to be tested by taking a color image of the shelf display as the image to be tested.

[0054] The image correction module 20 is communicatively connected to the image acquisition module 10 and is used to perform correction and alignment processing on the image to be tested to obtain a corrected image of the image to be tested. The correction and alignment processing of the image to be tested is as follows:

[0055] 1. Take the image to be tested and the standard display image of the area as input, call the image matching model to establish the correspondence between multiple identical scenes or pixel areas in the two images, and the consistency probability.

[0056] 2. Sort the relationships from highest to lowest probability and select the top 15 relationships with the highest probability as the reference coordinates for the two images.

[0057] 3. Calculate the transformation matrix based on the coordinates of the associated regions of the image to be tested and the standard image.

[0058] 4. Based on the obtained transformation matrix, perform perspective transformation on the image to be tested, and crop it according to the width and height of the standard image with the center point as the origin to obtain the corrected and aligned image to be tested.

[0059] The change region detection module 30 detects change regions in the corrected image by calling a scene change region target detection algorithm, and detects objects in the change regions, outputting basic information about the objects in the change regions. The basic information about the objects in the change regions includes the location and category information of all objects in the change regions of the image, such as 1. background, 2. bottle, 3. jar, 4. box, etc.

[0060] The object recognition module 40 is used to perform object recognition on the basic information of objects in the changing area. Here, object recognition refers to the identification of specific items or displays, such as recognizing a bottle as a Coca-Cola can. The object recognition process is as follows:

[0061] 1. Call the object recognition model to extract features from the object images that have changed in the test image and obtain the feature vector of the object to be recognized.

[0062] 2. Using the feature vectors of items in the standard library, match and query with the feature vectors of each item to be identified to obtain the best matching result and acquire item information.

[0063] This invention provides a computer vision solution that effectively replaces the manual work of displaying and inventorying items on shelves. It not only saves labor but also avoids errors caused by human error.

[0064] Furthermore, compared to other computer vision-based solutions, this invention, through its image correction module, can effectively eliminate or reduce errors such as camera offset or shooting angle rotation. Moreover, through its change detection module, it effectively reduces the number of items that need to be identified on the shelf, thereby reducing the system's computational load.

[0065] The technical solutions of the present invention will now be described in detail with reference to the accompanying drawings.

[0066] like Figure 2 As shown, this embodiment of the invention provides an object detection and recognition method, which is used in a trained change region target detection model. The method includes the following steps.

[0067] Step S11: Perform image preprocessing on the standard image and the corrected image of the shelf to be tested, and extract similar feature maps to generate a first feature map and a second feature map respectively; wherein, the standard image is the image of the shelf to be tested before the change of the item display, and the image to be tested is the image of the shelf to be tested after the change of the item display.

[0068] Step S12: Subtract the corresponding positions in the first feature map and the second feature map and take the absolute value to generate the third feature map;

[0069] Step S13: Detect items in the changed area based on the third feature map to obtain basic information about the items in the changed area;

[0070] Step S14: Identify items based on the basic information of the items in the changed area.

[0071] Next will be Figure 3-4 The above process will be explained in detail using 5A-5C, 6, 7A-7C and 8 as examples.

[0072] Step S11: Perform image preprocessing on the standard image and the corrected image of the shelf to be tested, and extract similar feature maps to generate a first feature map and a second feature map respectively; wherein, the standard image is the image of the shelf to be tested before the change of the item display, and the image to be tested is the image of the shelf to be tested after the change of the item display.

[0073] Specifically, the standard image and the image to be tested of the shelf under test are images taken in the same shelf setting. The standard image of the shelf under test is a standardized image of the shelf under test before any changes to the item arrangement. The image to be tested is an image of the shelf under test after any changes to the item arrangement. This image can be taken using a camera or video camera, and the camera can be movable or fixed.

[0074] Before using the image to be tested, considering that the camera may be unbalanced when taking pictures, and the image may be rotated by a certain angle, it is necessary to correct the image to make the image better for subsequent use.

[0075] like Figure 3 As shown, in one embodiment of the present invention, the method for obtaining the corrected image of the image to be tested includes:

[0076] Step S31: Obtain the image to be tested based on the shelf to be tested;

[0077] Step S32: Based on the image to be tested and the standard image of the shelf to be tested, obtain the corresponding point coordinates and consistency probability of the image to be tested and the standard image of the shelf to be tested by calling the image matching model;

[0078] Step S33: Select the N sets of corresponding point coordinates with the highest consistency probability as the reference coordinates of the image to be tested and the standard image of the shelf to be tested;

[0079] Step S34: Based on the image to be tested, the standard image of the shelf to be tested, and their respective reference coordinates, calculate the transformation matrix;

[0080] Step S35: Perform perspective transformation on the image to be tested according to the transformation matrix to obtain the corrected image of the image to be tested.

[0081] Specifically, in the process of establishing a consistency relationship between the pixels of the image under test and the corresponding standard shelf image for the same scene or region, the image matching model invoked is not limited to any specific model or training data in this invention. Existing image matching models such as LoFTR, SuperGlue, and DISK can be used, or a custom-trained model can be used. Specific details are not provided here.

[0082] The input to the image matching model is an image of the shelf to be tested (e.g., the image to be tested) and a standard item display image of the shelf already entered into the annotation library. The output is the coordinate information and correspondence of the same pixel areas in the two images. A standard shelf image library needs to be established in advance. First, ensure the items are arranged in a standardized and compliant manner. Then, photograph each standard shelf, taking care to keep the camera or lens level and parallel to the shelf surface as much as possible. Mark the location area of ​​each standard shelf photograph. Enter the photographs and area information into the standard shelf reference library.

[0083] like Figure 4 As shown, in one embodiment of the present invention, the method for acquiring the standard image of the shelf to be tested includes:

[0084] Step S41: Establish a standard shelving diagram library;

[0085] Step S42: Take a picture of the standard display shelf using a camera or video camera to obtain an image of the standard display shelf;

[0086] Step S43: Mark the location area to which the standard display shelf image belongs, and obtain the area information of the standard display shelf image;

[0087] Step S44: Input the standard display shelf image and the area information of the standard display shelf image into the standard shelf image library;

[0088] Step S45: Obtain a standard image of the same shelf as the image to be tested from the standard shelf image library.

[0089] The image to be tested is aligned and calibrated according to the labeled shelf image. Based on the process of establishing the consistency relationship between pixels of the same scene or region in the image to be tested and the corresponding standard shelf image, the coordinates and scores of corresponding points in similar or consistent pixel regions of the image to be tested and the standard image are obtained. These coordinates are sorted from largest to smallest, and the top 10 sets of coordinates are selected as references to calculate homography, that is, the mapping relationship between points on one two-dimensional plane and points on another plane. Let the coordinate set P of the image to be tested be as follows:

[0090] P={(x1,y1),(x2,y2),(x3,y3),...,(x 10 ,y 10 )}

[0091] The standard map coordinate point set is P'

[0092] P'={(x'1,y'1),(x'2,y'2),(x'3,y'3),...,(x' 10 ,y' 10 )}

[0093] Use the findHomography function in OpenCV to calculate the homography matrix H;

[0094] H = cv2.findHomography(P,P')

[0095] Let im1 represent the image to be tested and im2 represent the standard image. width and height represent the width and height of the im2 image, respectively. The warpPerspective function in OpenCV can be used to obtain the calibrated and aligned image to be tested im1_reg for subsequent steps.

[0096] im1_reg=cv2.warpPerspective(im1,H,(width,height))

[0097] It should be noted that OpenCV's built-in warpPerspective() function performs a mapping transformation from one image to another based on perspective transformation parameters.

[0098] The standard image and the corrected image of the shelf to be tested are preprocessed separately, such as scaling the image to 512x512 resolution, or normalizing the image and subtracting the average pixel value; and similar feature maps are extracted to generate the first feature map and the second feature map respectively.

[0099] Step S12: Subtract the corresponding positions in the first feature map and the second feature map and take the absolute value to generate the third feature map;

[0100] Specifically, the first feature map and the second feature map are stitched together, that is, the corresponding positions in the first feature map and the absolute values ​​are subtracted and stitched together to generate the third feature map. This ensures that all the corresponding item information on the third feature map is preserved.

[0101] Step S13: Detect items in the changed area based on the third feature map to obtain basic information about the items in the changed area;

[0102] Specifically, the basic information of the items in the changed area is the location and category information of all items in the changed area of ​​the image to be tested, such as 1 background, 2 bottle, 3 jar, 4 box, etc.

[0103] like Figure 5A As shown, in one embodiment of the present invention, the step of detecting items in changing regions based on the third feature map to obtain basic information about the items in changing regions includes:

[0104] Step S131: Detect change regions in the third feature map to obtain a change probability map;

[0105] Step S132: Perform item detection on the third feature map to obtain the basic information of the item;

[0106] Step S133: Based on the change probability map and the basic information of the item, obtain the basic information of the item in the change area.

[0107] like Figure 5B As shown, in one embodiment of the present invention, the step of performing change region detection on the third feature map to obtain a change probability map includes:

[0108] Step S1311: Obtain the change probability of the change probability map;

[0109] Step S1322: If the change probability is close to 1, the item in the third feature map has changed;

[0110] Step S1323: If the change probability is close to 0, the items in the third feature map have not changed.

[0111] Specifically, the size of the change probability map is consistent with that of the input image (such as the corrected image of the image to be tested). The change probability of the change probability map is represented by pixel values. When the pixel value is closer to 1, it means that a change is more likely. When the pixel value is closer to 0, it means that no change has occurred.

[0112] like Figure 5C As shown, in one embodiment of the present invention, obtaining the basic information of the items in the change region based on the change probability map and the basic information of the items includes:

[0113] Step S1331: Obtain the location of the item based on its basic information;

[0114] Step S1332: Determine the position of the item and the change area in the change probability map;

[0115] Step S1333: If the position of the item is consistent with the change area in the change probability map, then retain the basic information of the item;

[0116] Step S1334: Otherwise, do not retain.

[0117] Specifically, the number of items in the standard image and the image to be tested of the shelf may be multiple, and the number of items in the corresponding generated third feature map may also be multiple. By judging the change areas in the position and change probability map of all items, only the basic information of the items in the change area is retained.

[0118] The trained change region object detection model possesses the functions of scene change detection and change region object detection. For example... Figure 6 The diagram shows the structure of the trained change region target detection model according to an embodiment of the present invention. This model is primarily used to identify changing products in image pairs of the same scene taken at different times, with the product as the primary target and other changes treated as noise. The recognition module is mainly divided into three parts: the first part is image preprocessing, which scales the image to 512x512 resolution, normalizes the image, and subtracts the pixel mean (optional). The second part is a bitemporal image transformer, which is a Siamese network. The network structure can be any mature visual transformer network structure, such as ViT-B, ViT-H, and ViT-L, etc. Its main function is to extract similar feature maps from the two preprocessed images. The third part is the prediction module, which has two independent prediction heads, both of which are FCN (Fully Convolutional Networks) network structures. The first prediction head subtracts the corresponding similar feature maps extracted from two images and takes the absolute value. This result is then processed through a three-layer fully convolutional network to obtain a change probability map. The size of the change probability map is the same as the input image, where pixel values ​​represent the probability of change; values ​​closer to 1 indicate a higher likelihood of change, while 0 indicates no change. The second item prediction and detection head primarily detects the basic information of items in the change region of the third feature map (including the item's location and category information).

[0119] FCN (Fully Convolutional Networks) replaces the fully connected layers at the back of traditional convolutional networks with convolutional layers, so that the network output is no longer a category but a heatmap; at the same time, in order to solve the impact of convolution and pooling on image size, it proposes to use upsampling to restore the image size.

[0120] It should be noted that a feature map is a three-dimensional number that contains a mapping of information such as the points, lines, surfaces, and textures of an image in a high-dimensional Hilbert space.

[0121] The detection head is a detection network, namely a convolutional neural network or a multilayer perceptron, which may involve one or more three-dimensional floating-point convolution operations.

[0122] The feature map C, after training, extracts a three-dimensional array of information that is essential for the tasks of scene change and object detection. This information is not specifically interpretable.

[0123] The so-called detection head is a multilayer perceptron for a specific task. Here, it is divided into two tasks: scene change detection and object detection. These two detection heads can be separated (as shown in Figure 5, the scene change detection head and the object detection head) because the output data sizes are too different. Of course, they can also be merged. The scene change detection head needs to be trained using the scene change detection loss function. Similarly, the object detection head also needs to be trained using the object detection loss function.

[0124] In addition to the main network structure, loss functions need to be designed. There are three loss functions here. The first loss function is the L1 norm change detection loss function, as follows:

[0125]

[0126] Among them, L ce This represents the change detection loss function, where H and W are the width and height of the input image (the standard image of the shelf to be tested mentioned above or the corrected image of the image to be tested, respectively); P h,w Y represents the probability value of change in row h and column w of the probability graph. h,w The value represents the change in row h and column w in the binary label graph. The change is 1 and the unchanged value is 0. h∈[1,H], w∈[1,W].

[0127] The binarized label map is generated based on the standard image and the corrected image of the shelf under test, depicting the changed regions. For each image pair, a binary label map of the changed regions with the same size as the original image is generated. Specifically, for each image pair, a change probability map is first obtained, then a threshold is set: values ​​greater than a threshold (th) are considered changed and assigned a value of 1, while values ​​less than the threshold are considered unchanged and assigned a value of 0. In this binary label map, regions with a pixel value of 1 represent changed regions in the image pair, and regions with a pixel value of 0 represent unchanged regions in the image pair.

[0128] The second loss function is the cross-entropy loss function for bounding box classification tasks, as follows:

[0129] L cls =-∑IOU mask ×(Y obj ×log(P obj )+(1-Y obj )×log(1-P obj ))

[0130] Among them, L cls IOU represents the loss function for bounding box classification tasks based on cross-entropy. mask Y is a mask vector indicating whether the predicted bounding box (the bounding box predicted by the change region object detection model on the corrected image of the test image) is located within a change region on the change probability map. It is 1 when the predicted bounding box is within a change region, and 0 otherwise. obj This represents the true class label for the object detection task (usually represented using one-hot encoding), for example: [1,0,0] means the first class is the object, and the other classes are not objects; while P obj This is the class probability distribution predicted by the model (change region object detection model). This is the output of the change region object detection model. When the model makes predictions, the output needs to include the probability of belonging to each class. For example, if the prediction is {bottle, box, jar}, then P... obj It refers to the probability of being a bottle, the probability of being a box, and the probability of being a jar. For example, [0.8, 0.1, 0.1] means that the model predicts a probability of 0.8 for the first category and 0.1 for each of the other categories.

[0131] The third loss function is the squared error loss function for the bounding box, which also needs to be applied only to the changing region. See below.

[0132] L box =∑IOU mask ×L2(box pred ,box ture )

[0133] L2 (box) pred ,box ture ) = L x +L y +L w +L h

[0134] L x =|x pred -x ture | 2

[0135] L y =|y pred -y ture | 2

[0136] L w =|w pred -w ture | 2

[0137] L h =|h pred -h ture | 2

[0138] Among them, box pred This represents the prediction box, with center point coordinates and width and height [x, y ... pred ,y pred ,w pred ,h pred ], box ture This represents a label box, with its center point coordinates and width and height as [x, y ... ture ,y ture ,w ture ,h ture ], L box L2 represents the squared error loss function for the varying region. pred ,box ture ) represents the specific loss function (the error between the predicted and actual values), L x L represents the specific loss of the center point's x-coordinate (the error between the predicted and actual values ​​of the center point's x-coordinate). y The specific loss of the center point ordinate (the error between the predicted and actual values ​​of the center point ordinate), L w The specific loss of the width (the error between the predicted and actual width), L h High specific losses (high error between predicted and actual values).

[0139] Finally, summing them proportionally, the total loss function is as follows:

[0140] L total =β1L ce +β2L cls +β3L box

[0141] Among them, L total The total loss function is represented by β1, β2, and β3, which represent the proportions of the loss function for each task, for example, β1 = 0.4, β2 = 0.3, and β3 = 0.3. The loss functions used for scene change detection and object detection are not limited to the methods mentioned above. Through this filtering mechanism and loss function design, object detection tasks will focus more on objects within changing regions. Interchange of Units (IOU) is used to meet functional requirements. mask The specific design is as follows:

[0142] First, generate the binary mask M:

[0143] M = threshold(y pred ,th)

[0144] Where M is a binary mask, y pred This is the probability map of the predicted change region. `threshold()` is a commonly used thresholding function, where `th` is the set threshold. Next, we need to generate a mask vector for the predicted bounding box:

[0145] IOU mask =IOU(box) pred ,M)>iou th

[0146] Among them, IOU mask It is a mask vector indicating whether the predicted bounding box is located within the change region. IOU(box) pred M) represents the predicted bounding box. pred The intersection-union ratio (IoU) between the binary mask M and the binary mask M th This is a threshold set for IOU (Intersection over Union) determination. When the intersection-union ratio (IoU) of the predicted bounding box and the changed region exceeds the threshold, it is used for IoU determination. th At that time, IOU mask The corresponding element is 1, indicating that the predicted bounding box is located within the change region; otherwise, it is 0. In this formula, box pred The predicted bounding box is represented by the object detection prediction head mentioned above (e.g., Figure 6 The output of the item detection head in the box. pred It contains the coordinate information of the bounding box (center point coordinates, width, and height).

[0147] Training data is required before training a change region object detection model. This invention uses some existing image annotation tools (such as VGG Image Annotator), which need to be able to manually or semi-automatically draw bounding boxes, segment regions, etc. on images, and assign category labels to these annotations.

[0148] 1. First, prepare image pairs. For each sample, at least two images taken at different times are needed. These two images are usually taken from the same location, but at different points in time. For example, one image could be taken before the change of merchandise display on the shelf, serving as the base image A (see [reference]). Figure 7A The other image is Figure B, which was created after changes were made to the product display on the shelf (see Figure B). Figure 7B (Excluding bounding boxes and object categories).

[0149] 2. Then, create a binary label map of the changed area of ​​the image to be tested (see...). Figure 7C For each image pair, a binary label map with the same size as the image needs to be generated. In this label map, the area with a pixel value of 1 represents the changed area in the image pair, and the area with a pixel value of 0 represents the unchanged area in the image pair.

[0150] 3. Finally, label the bounding boxes and category information of the changing objects in the image to be tested. For each sample, detection boxes need to be labeled for all objects that have changed in the image to be tested. This includes the bounding boxes of the objects (rectangles containing the object's location) and the categories to which the objects belong (e.g., bottles, jars, bags, boxes, etc.) (see...) Figure 7B ).

[0151] Training the changed region target detection model. Follow these steps to train the changed region target detection model:

[0152] 1. Data preprocessing: First, the training data is preprocessed, including image normalization and scaling to the same resolution.

[0153] 2. Data Augmentation (Optional): To improve the model's generalization ability, data augmentation can be performed on the training data. Data augmentation techniques include synchronous rotation, which keeps the rotation angle of the two images consistent; synchronous flipping, which flips the two images horizontally or vertically; and image brightness and contrast adjustment, which adjusts the brightness and contrast while keeping the differences between the two images unchanged.

[0154] 3. Split the dataset: Divide the dataset into a training set, a validation set, and a test set. The training set is used to train the model, the validation set is used to tune hyperparameters and check for overfitting, and the test set is used to evaluate the final performance of the model.

[0155] 4. Model Initialization: Initialize the model based on your chosen network architecture. You can train from scratch or use transfer learning with pre-trained weights based on ImageNet or other large datasets to accelerate training and improve performance.

[0156] 5. Set the optimizer and learning rate strategy: Select an optimizer (such as SGD, Adam, etc.) and set the learning rate and other optimizer parameters.

[0157] 6. Model Training: In each training iteration, a pair of images, their corresponding binary label maps of the changed regions, and the bounding boxes and category information of the changed objects are randomly selected from the dataset. The image pairs are input into the model, and the outputs of the scene change detection prediction head and the object detection prediction head are calculated. Based on the output results, the binary label maps of the changed regions, and the bounding boxes and category information of the changed objects, the value of the loss function is calculated. The model weights are updated using gradient descent methods (such as SGD, Adam, etc.) to minimize the loss function.

[0158] 7. Model Convergence and Validation: At the end of each training epoch, evaluate the model's performance using a validation set. Adjust hyperparameters or learning rate strategies based on the validation set performance. Stop training when the model's performance on the validation set reaches the expected level or no longer shows significant improvement. Save the trained model weights for evaluation on the test set.

[0159] 8. Model Testing: After training, evaluate the final performance of the model using a test set. Calculate the loss and other evaluation metrics on the test set to ensure the model performs well on unseen data. Select the best-performing model as the target detection module for changing regions.

[0160] like Figure 8 As shown, in one embodiment of the present invention, the item identification based on the basic information of the items in the changed area includes:

[0161] Step 81: Extract features from the items in the changed area to obtain the features of the items to be identified;

[0162] Step 82: Match and query the features of the item to be identified with the features of items in the preset standard library one by one to obtain the best matching result for the item;

[0163] Step 83: Obtain item information based on the best matching result of the item.

[0164] Specifically, item identification is performed based on the basic information of the items in the changed area. Here, item identification refers to the identification of specific items or displays, such as identifying a bottle as a Coca-Cola can.

[0165] The item recognition model can utilize existing metric-based product recognition models, such as Triplet Loss Network, ArcFace Network, and Baidu ppshitu. Before use, a product image library needs to be prepared, and features should be extracted using a product recognition network to establish a standard feature library for the items. During use, features are extracted from objects in changing areas using the product recognition model, and then similarity is calculated with the standard feature library. The item with the highest similarity is considered to match the target item, and the shelf display information is updated accordingly. Users can choose whether to update the floor plan and product information.

[0166] This invention provides a computer vision solution that effectively replaces manual inventory counting and shelf display, saving labor and avoiding errors caused by human negligence. The overall process for item detection and recognition described in this invention is as follows:

[0167] 1. Record the base map of the shelving and the information of the items into the warehouse;

[0168] 2. Acquisition of the image to be tested;

[0169] 3. Calculate the consistency relationship with the location annotation map and calibrate the image to be tested;

[0170] 4. Detect objects in changing areas of the image under test;

[0171] 5. Use a metric model to identify items in areas of change;

[0172] 6. Array information update;

[0173] 7. Update the base database map and update...

[0174] Compared with other computer vision-based solutions, this invention has the following advantages:

[0175] 1. In this invention, image preprocessing is performed on the standard image and the corrected image of the shelf to be tested, respectively, and similar feature maps are extracted to generate a first feature map and a second feature map. The standard image is the image of the shelf before the item arrangement changes, and the image to be tested is the image of the shelf after the item arrangement changes. A third feature map is generated by subtracting the corresponding positions in the first and second feature maps and taking the absolute value. Based on the third feature map, item detection in the changed area is performed to obtain basic information about the items in the changed area. Item recognition is then performed based on the basic information about the items in the changed area. This effectively reduces the number of items that need to be recognized on the shelf, thereby reducing the system's computational load.

[0176] 2. This invention obtains a corrected image of the image under test by performing correction processing on the image under test, which can effectively eliminate or reduce errors such as camera offset or shooting angle rotation.

[0177] The scope of protection of the item detection and identification method described in this embodiment is not limited to the execution order of the steps listed in this embodiment. Any solution implemented by adding, subtracting, or replacing steps in the prior art based on the principle of this invention is included within the scope of protection of this invention.

[0178] This invention also provides an item detection and identification device, which can implement the item detection and identification method described in this invention. However, the implementation device of the item detection and identification method described in this invention includes, but is not limited to, the structure of the item detection and identification device listed in this embodiment. All structural modifications and substitutions of the prior art made in accordance with the principles of this invention are included within the protection scope of this invention.

[0179] like Figure 9 As shown, this embodiment provides an item detection and recognition device. The item detection and recognition device 100 includes a feature map generation module 110, an image stitching module 120, a change area item detection module 130, and a change area item recognition module 140.

[0180] The feature map generation module 110 is used to perform image preprocessing on the standard image and the corrected image of the shelf to be tested, respectively, and extract similar feature maps to generate a first feature map and a second feature map; wherein, the standard image is the image of the shelf to be tested before the change of the item display, and the image to be tested is the image of the shelf to be tested after the change of the item display;

[0181] The image stitching module 120 is used to subtract corresponding positions from the first feature map and the second feature map and take the absolute value to generate a third feature map;

[0182] The change area item detection module 130 is used to detect change area items based on the third feature map to obtain basic information of change area items; wherein, the basic information of change area items is all the location and category information of the change area of ​​the image to be tested, such as 1 background, 2 bottle, 3 jar, 4 box, etc.

[0183] The change area item recognition module 140 is used to perform item recognition based on the basic information of the items in the change area. Here, item recognition refers to the identification of specific items or displays, such as identifying a bottle as a Coca-Cola can.

[0184] In the embodiments provided by this invention, it should be understood that the disclosed systems, apparatuses, or methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of modules / units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or units may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection of apparatuses or modules or units may be electrical, mechanical, or other forms.

[0185] The modules / units described as separate components may or may not be physically separate. The components shown as modules / units may or may not be physical modules; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules / units can be selected to achieve the objectives of the embodiments of the present invention, depending on actual needs. For example, the functional modules / units in the various embodiments of the present invention may be integrated into one processing module, or each module / unit may exist physically separately, or two or more modules / units may be integrated into one module / unit.

[0186] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0187] like Figure 10 As shown, this embodiment provides an electronic device 200, which includes a memory 210 and a processor 220; the memory 210 is used to store computer programs; the processor 220 is used to execute the computer programs stored in the memory 210 to perform the above-described item detection and identification method.

[0188] This invention also provides a computer-readable storage medium. Those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing a processor. The program can be stored in a computer-readable storage medium, which is a non-transitory medium, such as random access memory, read-only memory, flash memory, hard disk, solid-state drive, magnetic tape, floppy disk, optical disk, and any combination thereof. The storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. This available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., digital video disc (DVD)), or a semiconductor medium (e.g., solid-state drive (SSD)).

[0189] Embodiments of the present invention may also provide a computer program product comprising one or more computer instructions. When the computer instructions are loaded and executed on a computing device, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means.

[0190] When the computer program product is executed by a computer, the computer performs the method described in the foregoing method embodiments. The computer program product can be a software installation package; when the foregoing method is required, the computer program product can be downloaded and executed on the computer.

[0191] The descriptions of the processes or structures corresponding to the above figures each have their own emphasis. For parts of a process or structure that are not described in detail, please refer to the relevant descriptions of other processes or structures.

[0192] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the invention. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in the present invention should still be covered by the claims of the present invention.

Claims

1. A method for detecting and identifying objects, characterized in that, The method includes: Image preprocessing is performed on the standard image and the corrected image of the shelf under test, respectively, and then input into a Siamese network to extract similar feature maps to generate a first feature map and a second feature map. The standard image is the image of the shelf under test before the change in item arrangement, and the image under test is the image of the shelf under test after the change in item arrangement. The method for obtaining the corrected image of the image under test includes: obtaining the image under test based on the shelf under test; obtaining the corresponding point coordinates and consistency probability of the image under test and the standard image of the shelf under test by calling an image matching model; selecting the N sets of corresponding point coordinates with the highest consistency probability as reference coordinates for the image under test and the standard image of the shelf under test; calculating a transformation matrix based on the image under test, the standard image of the shelf under test, and their respective reference coordinates; and performing a perspective transformation on the image under test according to the transformation matrix to obtain the corrected image of the image under test. Subtract the corresponding positions from the first feature map and the second feature map and take the absolute value to generate the third feature map; Based on the third feature map, change region item detection is performed to obtain basic information of the change region items; change region detection is performed on the third feature map to obtain a change probability map; item detection is performed on the third feature map to obtain basic information of the items; based on the change probability map and the basic information of the items, basic information of the change region items is obtained; obtaining basic information of the change region items based on the change probability map and the basic information of the items includes: obtaining the location of the items based on the basic information of the items; judging the location of the items and the change region in the change probability map; if the location of the items is consistent with the change region in the change probability map, the basic information of the items is retained; otherwise, it is not retained; Item identification is performed based on the basic information of the items in the changed area; features are extracted from the items in the changed area to obtain the features of the items to be identified; the features of the items to be identified are matched and queried one by one with the features of items in the preset standard library to obtain the best matching result; and item information is obtained based on the best matching result.

2. The item detection and identification method according to claim 1, characterized in that, The step of performing change region detection on the third feature map to obtain a change probability map includes: Obtain the probability of change of the aforementioned probability map; If the probability of change is close to 1, the item in the third feature map has changed; If the probability of change is close to 0, the items in the third feature map have not changed.

3. The item detection and identification method according to claim 1, characterized in that, The methods for acquiring the standard image of the shelf to be tested include: Establish a standard shelving diagram library; Obtain images of standard display shelves by photographing them with a camera or webcam. Mark the location area to which the standard display shelf image belongs, and obtain the area information of the standard display shelf image; Enter the standard display shelf image and the area information of the standard display shelf image into the standard shelf image library; Obtain a standard image of the same shelf as the image to be tested from the standard shelf image library.

4. An item detection and identification device, characterized in that, The item detection and identification device includes: The feature map generation module is used to preprocess the standard image and the corrected image of the shelf under test, respectively, and input them into a Siamese network to extract similar feature maps to generate a first feature map and a second feature map. The standard image is the image of the shelf under test before the change in item arrangement, and the image under test is the image of the shelf under test after the change in item arrangement. The method for obtaining the corrected image of the image under test includes: obtaining the image under test based on the shelf under test; obtaining the corresponding point coordinates and consistency probability of the image under test and the standard image of the shelf under test by calling an image matching model; selecting the N sets of corresponding point coordinates with the highest consistency probability as reference coordinates for the image under test and the standard image of the shelf under test; calculating a transformation matrix based on the image under test, the standard image of the shelf under test, and their respective reference coordinates; and performing a perspective transformation on the image under test according to the transformation matrix to obtain the corrected image of the image under test. The image stitching module is used to subtract the corresponding positions in the first feature map and the second feature map and take the absolute value to generate a third feature map; The changing region item detection module is used to perform changing region item detection based on the third feature map to obtain basic information of the changing region items; perform changing region detection on the third feature map to obtain a changing probability map; perform item detection on the third feature map to obtain basic information of the items; and obtain basic information of the changing region items based on the changing probability map and the basic information of the items. The step of obtaining basic information of the changing region items based on the changing probability map and the basic information of the items includes: obtaining the location of the item based on the basic information of the item; judging the location of the item and the changing region in the changing probability map; if the location of the item is consistent with the changing region in the changing probability map, then the basic information of the item is retained; otherwise, it is not retained. The variable area item recognition module is used to identify items based on the basic information of the items in the variable area; extract features from the items in the variable area to obtain the features of the items to be identified; match and query the features of the items to be identified with the features of items in a preset standard library one by one to obtain the best matching result; and obtain item information based on the best matching result.

5. An electronic device, characterized in that, The electronic device includes: Memory, used to store computer programs; A processor is configured to execute a computer program stored in the memory to perform the article detection and identification method as described in any one of claims 1 to 3.

6. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by an electronic device, the program implements the item detection and identification method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Radar and vision fused vehicle identification method and device, equipment and medium

    CN113807471A

  • Center Net-based remote sensing image building change detection model and algorithm

    CN114241314A