A method and device for detecting similar products

By using deep neural network combined with linear discriminant analysis and clustering methods in shelf product detection, the problem of low detection accuracy of similar products is solved and higher detection accuracy is achieved.

CN115457320BActive Publication Date: 2025-06-13ZHEJIANG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211079491.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-05
Publication Date
2025-06-13
Estimated Expiration
2042-09-05

AI Technical Summary

Technical Problem

In shelf product inspection, similar products have low detection accuracy, especially products with the same name but different capacity are easily misidentified, resulting in a decrease in detection accuracy.

Method used

The shelf similar product detection method based on deep neural network is adopted. Through the combination of feature extraction network, regional candidate network and target detection network, and the linear discriminant analysis and clustering method is used to optimize the feature vector of the target detection network, reduce the distance of similar products, and increase the distance of different types of products, thereby improving the detection accuracy.

Benefits of technology

It effectively improves the detection accuracy of similar products, solves the problem of low detection accuracy of similar products on the shelf, and meets the accuracy requirements for shelf product inspection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115457320B_ABST
    Figure CN115457320B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for detecting similar products on a shelf based on a deep neural network. In the present invention, a product image is input into a trained product image detection model to obtain the specific positions and classification results of various products in the product image. When the product image detection model is trained, the linear discriminant analysis method is used to process the feature vectors of the intermediate layer of the target detection network to find a projection matrix that maximizes the distance between classes and minimizes the variance within classes. Then, the feature vectors of the intermediate layer are projected into the space corresponding to the projection matrix, and a clustering loss is constructed using a clustering method for the projected feature vectors to reduce the distance within the same class and increase the distance between different classes, thereby improving the detection accuracy of similar products on the shelf.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of object detection in computer vision, and specifically to a method and device for detecting similar products on a shelf based on a deep neural network. Background Art

[0002] In the retail industry, in order to understand the display situation of products on the shelf, better understand the market, and thus make decisions on marketing management, consumer goods manufacturers often analyze the types and placement positions of products placed on the shelf by taking pictures of store shelves, obtain information such as the product placement rate and the number of facings of different products in the store, and determine whether it meets the marketing strategies formulated by the manufacturer. Use the object detection method based on the deep neural network to quickly and accurately detect the products on the shelf in the picture.

[0003] In the actual application scenario of product detection, there is a problem of detecting similar products. For one or more pictures, there are often multiple products with the same name but different capacities. For example, Coca-Cola with 500 ml and Coca-Cola with 1.25 L have a very high similarity. When detecting products, it is easy to identify one product as another similar product, resulting in a significant decrease in the detection accuracy of similar products and unable to meet the accuracy requirements of shelf product detection. The high similarity between shelf product categories is currently a difficult point in the field of product detection and also the key to improving the detection accuracy of products. Summary of the Invention

[0004] In the actual application environment of shelf product detection, products with the same name but different capacities often have low detection accuracy due to high similarity. To solve this type of problem, the purpose of the present invention is to provide a method and device for detecting similar products on a shelf based on a deep neural network, and improve the detection accuracy of products with high similarity.

[0005] The technical solution adopted by the present invention to solve its technical problems is:

[0006] A method for detecting similar products on a shelf based on a deep neural network, which inputs a product image into a trained product image detection model to obtain the specific positions and classification results of various products in the product image; the product image detection model includes a feature extraction network, a region candidate network, and an object detection network, and is obtained through the following method:

[0007] (1) Input the samples of the constructed training dataset into the feature extraction network one by one to obtain feature maps; each sample of the constructed training dataset is a shelf product image, and the product image has a corresponding label and position coordinates;

[0008] (2) Then, the region proposal network gives candidate boxes that may contain the target for the feature map obtained in step (1);

[0009] (3) Map the candidate boxes obtained in step (2) to the feature map obtained in step (1) to obtain the corresponding feature matrix, input the feature matrix into the object detection network, construct the total loss function to train the commodity image detection model, and obtain the trained commodity image detection model; the total loss function is expressed as follows:

[0010] LOSS = λ 1 L 1 + λ 2 L 2 + λ dist L dist

[0011] Wherein, L 1 represents the loss between the classification result output by the object detection network and the ground truth, L 2 represents the loss between the position coordinates of the candidate boxes output by the region proposal network and the ground truth, λ 1 , λ 2 , λ dist are weights; L dist represents the clustering loss, which is established by the following method: use the linear discriminant analysis method to process the feature vectors in the middle layer of the object detection network to find the projection matrix that maximizes the inter-class distance and minimizes the intra-class variance, and then project the feature vectors in the middle layer into the space corresponding to the projection matrix to obtain the projected feature vectors; the projection matrix W n×d is composed of d eigenvectors of the matrix J, and n represents the dimension of the feature vectors; the matrix Where S b , S w are the between-class scatter matrix and within-class scatter matrix of the feature vectors obtained by passing all training data sets of the target commodity through the middle layer of the object detection network, respectively;

[0012]

[0013]

[0014] Wherein, N i (i = 0, 1,... K - 1) is the number of the i-th type of commodity in the training data set, μ i (i = 0, 1,... K - 1) is the mean vector of the feature vectors obtained by passing the i-th type of commodity through the middle layer of the object detection network, and μ is the mean vector of the feature vectors obtained by passing all commodities in the training data set through the middle layer of the object detection network; A i(i = 0, 1, … K-1) is the set of feature vectors obtained from the intermediate layer of the target detection network for the i-th category of goods, and a is the feature set A of the i-th (i = 0, 1, … K-1) category of goods i Elements in; K is the number of categories;

[0015] Then, use the clustering method on the projected feature vectors to construct and obtain the clustering loss:

[0016]

[0017] D(a′ j , p i ) represents the distance between the projected feature vector a′ j and the clustering center p of category i i , where d j,i is the Euclidean distance between the projected feature vector a′ of the j-th target good in step (3) j and the clustering center p of category i i , margin is the threshold of the distance, and R j,i represents the similarity between the j-th target good and category i, obtained based on the sample label. When R j,i = 1, it means the j-th target good belongs to category i. When R j,i = -1, it means the j-th target good does not belong to category i. K represents the number of categories.

[0018] Furthermore, in step (3), during the training process, the clustering center p i is updated. Represent the clustering center p i as the set P = p 0 , p 1 , …, p K-1 , and establish a memory S = q 0 , q 1 , …, q K-1 to store the temporary feature vectors generated during the training process. The way to update the clustering center is the weighted sum of the old clustering center and the new clustering center, so as to synchronously update the clustering loss L dist .

[0019] Furthermore, in step (3), at the beginning of network training, set λ dist = 0, and only calculate L 1 and L 2 ; when the training iteration number I of the network reaches the set number I s , set λ dist > 0, and calculate the clustering loss, L 1 and L 2 synchronously each time. At the same time, every time after It The clustering center p is updated once in the next iteration i until the total loss function converges or the total number of training times is reached

[0020] Further, the feature extraction network structure is a residual neural network or ResNeXt or SqueezeNet

[0021] Further, the target detection network structure is a region of interest pooling layer and a fully connected layer

[0022] Further, L 1 Adopt a multi-class cross-entropy loss function

[0023] Further, L 2 Adopt a Smooth L1 function

[0024] A detection device for similar commodities, comprising a memory, a processor and a computer program stored on the memory and operable on the processor, characterized in that when the processor executes the computer program, the detection method for similar commodities as described above is implemented

[0025] The beneficial effects of the present invention are mainly manifested in: it better solves the problem that the detection accuracy of the target detector is relatively low for similar commodities on the shelf. By adding linear discriminant analysis and clustering methods to the target detection network, the present invention further reduces the distance between the same categories and further increases the distance between different categories, thereby improving the detection accuracy of similar commodities Description of the Drawings

[0026] Figure 1 is the network training flow chart of the present invention

[0027] Figure 2 is the network structure diagram of the present invention

[0028] Figure 3 is the schematic diagram of classifying similar categories by the clustering method Detailed Embodiments

[0029] The present invention will be further described in detail below with reference to the drawings and specific embodiments

[0030] The present invention designs a method for detecting similar products on a shelf based on a deep neural network. This method inputs a product image into a trained product image detection model to obtain the specific positions and classification results of various products in the product image. When training the product image detection model, the linear discriminant analysis method is used to process the feature vectors of the intermediate layer of the object detection network to find a projection matrix that maximizes the inter-class distance and minimizes the intra-class variance. Then, the feature vectors of the intermediate layer are projected into the space corresponding to the projection matrix, and a clustering loss is constructed using the clustering method for the projected feature vectors to reduce the distance within the same category and increase the distance between different categories, thereby improving the detection accuracy of similar products on the shelf.

[0031] Reference Figure 2 , the product image detection model includes a feature extraction network, a region proposal network (RPN), and an object detection network. The feature extraction network can adopt a residual neural network, ResNeXt, SqueezeNet, etc. The object detection network can adopt a region of interest pooling layer and a fully connected layer. The training method of the product image detection model is as Figure 1 shown and includes the following steps:

[0032] (1) Input the samples (product images) of the constructed training dataset into the feature extraction network one by one to obtain feature maps; the feature extraction network used in the present invention is a 50-layer residual neural network; each sample of the constructed training dataset is a shelf product image, and the product image has a corresponding label and position coordinates;

[0033] (2) Then use the region proposal network to give candidate boxes that may contain the target for the feature maps obtained in step (1);

[0034] (3) Map the candidate boxes obtained in step (2) to the feature maps obtained in step (1) to obtain corresponding feature matrices, input the feature matrices into the object detection network, and use the linear discriminant analysis method to process the feature vectors of the intermediate layer of the object detection network to find a projection matrix that maximizes the inter-class distance and minimizes the intra-class variance. Then, project the feature vectors of the intermediate layer into the space corresponding to the projection matrix;

[0035] The specific steps of the linear discriminant analysis method are as follows:

[0036] (3a) Assume that there are K categories in the dataset, and the feature set is A = {(a 1 , y 1 ), (a 2 , y 2 ), …, (a m , y m )}, where any product target a j is an n-dimensional vector, and the label y corresponding to each targetj Belonging to {C 0 , C 1 , …, C K-1}; m is the target quantity of goods;

[0037] (3b) Define N i (i = 0, 1, … K - 1) as the number of goods of the i-th category, A i (i = 0, 1, … K - 1) as the feature set of the i-th category of goods, μ i (i = 0, 1, … K - 1) as the mean vector of the features of the i-th category of goods, where a is an element in the feature set A of the i-th (i = 0, 1, … K - 1) category of goods. Calculate the between-class scatter matrix i

[0038]

[0039] where μ is the mean vector of all goods features, calculate the within-class scatter matrix

[0040]

[0041] (3c) Calculate the matrix

[0042] (3d) Calculate the d largest eigenvalues of the matrix J and the corresponding d eigenvectors (w 1 , w 2 , … w d ), to obtain the projection matrix W n×d ; n represents the dimension of the eigenvector.

[0043] (3e) Project the feature set A into the space corresponding to the matrix W n×d a′ j = W T a j , to obtain the projected feature set A′ = {(a′ 1 , y 1 ), (a′ 2 , y 2 ), … (a′ m , y m )};

[0044] (4) Use the clustering method on the projected feature vectors in step (3) to construct a clustering loss to reduce the distance within the same category - that is, attract each other, and increase the distance between different categories - that is, repel each other, and jointly train the commodity image detection model with the loss between the classification result output by the target detection network and the ground truth and the loss between the predicted bounding box position coordinates output by the region candidate network and the ground truth to obtain a trained commodity image detection model. ​

[0045] The specific steps are as follows:

[0046] (4a) Refer to Figure 3 the schematic diagram of classifying similar categories by the clustering method; define the clustering loss, the number of data categories K, category i (i = 0, 1,..., K - 1), a' j as the feature vector of any target commodity after projection in step (3), p i as the mean of the feature vectors of category i after projection in step (3), that is, the clustering center. Define the clustering loss as

[0047]

[0048] , D(a' j , p i ) represents the distance between the feature vector a' j and the clustering center p i , where d j,i is the Euclidean distance between the feature vector a' j of the j-th target commodity after projection in step (3) and the clustering center p i of category i, margin is the threshold of the distance, R j,i represents the similarity between the j-th target commodity and category i, obtained according to the sample label. When R j,i = 1, it means that the j-th target commodity belongs to category i, that is, the two vectors are similar. When R j,i = -1, it means that the j-th target commodity does not belong to category i, that is, the two vectors are not similar;

[0049] (4b) Construct the total loss function to train the commodity image detection model. The total loss function is expressed as follows:

[0050] LOSS = λ 1 L 1 + λ 2 L 2 + λ dist L dist

[0051] Among them, L dist represents the clustering loss, L 1 represents the loss between the classification result output by the target detection network and the ground truth, and the multi-class cross-entropy loss function, etc. can be used. L 2 represents the loss between the predicted bounding box position coordinates output by the region candidate network and the ground truth, and the Smooth L1 function, etc. can be used. λ 1 , λ 2 , λ dist are weights.

[0052] As a preferred solution, during the training process, the clustering center p i is updated. The clustering center p i is represented as a set P = p 0 , p 1 , …, p K-1 . A memory S = q 0 , q 1 , …, q K-1 is established to store the temporary feature vectors generated during the training process. The clustering center is updated by the weighted sum of the old clustering center and the new clustering center, thereby synchronously updating the clustering loss L dist .

[0053] Furthermore, the network training is divided into multiple segments of operations. At the beginning of the network training, λ dist is set to 0, and only L 1 and L 2 are calculated; when the training iteration times I of the network reaches the set times I s (I s is a hyperparameter, and in the present invention, I s is set to 100), λ dist is set to be greater than 0, and clustering starts. When the iteration times is greater than I s , the loss of clustering, L 1 and L 2 are calculated each time. Every I t (I t is a hyperparameter, and in the present invention, I t is set to 200) iterations, the clustering center p i is updated once until the total loss function converges or reaches the total number of training times.

[0054] For a method for detecting similar goods on a shelf based on a deep neural network according to the present invention, the steps of inputting a commodity image into a trained commodity image detection model are as follows:

[0055] (1) Input the commodity image into a feature extraction network to obtain a feature map;

[0056] (2) Then use a region candidate network to give candidate boxes that may contain the target for the feature map obtained in step (1);

[0057] (3) Map the candidate boxes obtained in step (2) to the feature map obtained in step (1) to obtain a corresponding feature matrix. Input the feature matrix into a target detection network, and process the feature matrix using the weights of the target detection network obtained by training, and then the category of the predicted target and the position coordinates of the predicted bounding box can be output.

[0058] By adding the methods of linear discriminant analysis and clustering to the target detection network, the present invention further reduces the distance between the same categories and further increases the distance between different categories, thereby improving the detection accuracy of similar commodities.

[0059] Corresponding to the embodiment of a detection method for similar commodities described above, the present invention also provides an embodiment of a detection method device for similar commodities.

[0060] A detection method device for similar commodities provided by an embodiment of the present invention includes one or more processors for implementing a detection method for similar commodities in the above embodiment.

[0061] An embodiment of a detection device for similar commodities of the present invention can be applied to any device with data processing capabilities, and the any device with data processing capabilities can be a device or apparatus such as a computer.

[0062] The device embodiment can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a logically meaningful device, it is formed by the processor of any device with data processing capabilities reading the corresponding computer program instructions in the non-volatile memory into the memory for operation. From the hardware level, generally, a hardware structure of any device with data processing capabilities where a detection device for similar commodities of the present invention is located includes a processor, a memory, a network interface, and a non-volatile memory. In addition, depending on the actual functions of the any device with data processing capabilities where the device in the embodiment is located, other hardware may also be included, which will not be elaborated here.

[0063] For the specific implementation process of the functions and roles of each unit in the above device, please refer to the implementation process of the corresponding steps in the above method, which will not be elaborated here.

[0064] Obviously, the above embodiments are merely examples for clear illustration and are not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or variations can be made based on the above description. It is not necessary and impossible to list all the implementation manners here. And the obvious changes or variations derived therefrom are still within the protection scope of the present invention.

Claims

1. A detection method for similar products on a shelf based on a deep neural network, characterized in that, this method inputs a product image into a trained product image detection model to obtain the specific positions and classification results of various products in the product image; the product image detection model includes a feature extraction network, a region proposal network, and an object detection network, and is obtained by training through the following method: (1) Input the samples of the constructed training data set into the feature extraction network one by one to obtain feature maps; each sample of the constructed training data set is a shelf product image, and the product image has a corresponding label and position coordinates; (2) Then use the region proposal network to give candidate boxes that may contain targets for the feature maps obtained in step (1); (3) Map the candidate boxes obtained in step (2) to the feature maps obtained in step (1) to obtain corresponding feature matrices, input the feature matrices into the object detection network, construct a total loss function to train the product image detection model, and obtain a trained product image detection model; the total loss function is expressed as follows: LOSS = λ 1 L 1 + λ 2 L 2 + λ dist L dist Among them, L 1 represents the loss between the classification result output by the object detection network and the ground truth, and L 2 represents the loss between the position coordinates of the candidate box output by the region candidate network and the ground truth. λ 1 , λ 2 , and λ dist are weights; L dist represents the clustering loss, which is established by the following method: using the linear discriminant analysis method to process the feature vectors in the intermediate layer of the object detection network to find the projection matrix that maximizes the inter-class distance and minimizes the intra-class variance, and then projecting the feature vectors in the intermediate layer into the space corresponding to the projection matrix to obtain the projected feature vectors; the projection matrix W n×d is composed of d eigenvectors of the matrix J, and n represents the dimension of the eigenvector; the matrix where S b , S w are the between-class scatter matrix and the within-class scatter matrix of the feature vectors obtained by passing the target commodities in all training data sets through the intermediate layer of the object detection network, respectively; Then use a clustering method to construct a clustering loss for the projected feature vectors: D(a′ j , p i ) represents the distance between the projected feature vector a′ j and the cluster center p of class i i , where d j,i is the Euclidean distance between the projected feature vector a′ of the j-th target commodity in step (3) j and the cluster center p of class i i . Margin is the threshold of the distance, and R j,i represents the similarity between the j-th target commodity and class i, obtained according to the sample label. When R j,i = 1, it means that the j-th target commodity belongs to class i. When R j,i = -1, it means that the j-th target commodity does not belong to class i. K represents the number of classes.

2. The method according to claim 1, characterized in that, In step (3), during the training process, the clustering center p i is updated. The clustering center p i is represented as a set P = p 0 , p 1 , …, p K-1 . A memory S = q 0 , q 1 , …, q K-1 is established to store the temporary feature vectors generated during the training process. The clustering center is updated by the weighted sum of the old clustering center and the new clustering center, so as to synchronously update the clustering loss L dist .

3. The method according to claim 1, characterized in that, In step (3), at the beginning of network training, set λ dist = 0, and only calculate L 1 and L 2 ; when the training iteration number I of the network reaches the set number I s , set λ dist > 0, and synchronously calculate the clustering loss, L 1 and L 2 each time, and at the same time update the clustering center p i once every I t iterations until the total loss function converges or reaches the total number of training times.

4. The method according to claim 1, characterized in that, the structure of the feature extraction network is a residual neural network or ResNeXt or SqueezeNet.

5. The method according to claim 1, characterized in that, the structure of the object detection network is a region of interest pooling layer and a fully connected layer.

6. The method according to claim 1, characterized in that, L 1 The multi-class cross-entropy loss function is adopted.

7. The method according to claim 1, characterized in that, L 2 Adopt the Smooth L1 function.

8. A detection device for similar products, including a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, when the processor executes the computer program, it implements a detection method for similar products according to any one of claims 1-7.