Commodity Image Recognition Method and Device

By integrating local and overall features in product picture recognition, the problems of low recognition accuracy and slow efficiency in the prior art are solved, and efficient and accurate product picture recommendations are achieved.

CN113903026BActive Publication Date: 2025-07-18DIGITAL TRADING SCI & TECH (BEIJING) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111177667.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-09
Publication Date
2025-07-18
Estimated Expiration
2041-10-09

AI Technical Summary

Technical Problem

The prior art has problems in product image recognition with low recognition accuracy, slow execution speed, complex deployment and difficult updates, especially in large-scale recognition scenarios, which can easily lead to misidentification and waste of labor costs.

Method used

By intercepting the local pictures in the target product picture, the first picture feature vector is extracted, and the second picture feature vector is input to the designated intermediate layer of the pre-trained designated model for fusion processing, and combining local and overall features is determined whether to make recommendations.

Benefits of technology

It improves the accuracy and efficiency of product image recognition, reduces the cost of human review, and is suitable for large-scale product image recognition scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113903026B_ABST
    Figure CN113903026B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for identifying product pictures. The method includes: intercepting a partial picture from a target product picture and extracting a first picture feature vector; inputting the first picture feature vector into a pre-trained first specified model, and inputting a second picture feature vector into a specified intermediate layer of the first specified model for fusion processing to obtain a corresponding first output result; the second picture feature vector is extracted according to the target product picture; according to the first output result, it is determined whether to recommend the target product picture. The present invention fuses the features of the partial picture in the target product picture with the features of the overall picture of the target product picture, which not only retains the representativeness of the features of the partial picture, but also retains the relevance and connectivity of the overall picture, ensures the accuracy of the obtained result, and improves the recommendation effect of the product picture.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image recognition, and in particular to a commodity image recognition method and device. Background Art

[0002] Image recognition can be applied to different scenarios, such as identifying product images. Product images can help users understand the product in more detail, increase their favorability towards the product, and increase their willingness to buy the product. Therefore, choosing high-quality product images that meet display requirements and have user-attractive characteristics can better achieve the above effects.

[0003] When recognizing images, most existing technologies use deep learning image classification algorithms and transmission image algorithms. The recognition results are low in accuracy, slow in execution speed, complex in deployment in actual application, and troublesome to update the algorithm. They are not suitable for large-scale commodity image recognition scenarios. Most existing image classification algorithms use open source algorithms, and their recognition algorithms have limitations, which can easily lead to problems such as inaccurate recognition. For example, if a commodity image contains logo information, similar logo information will be directly recognized as the same type of commodity image when the existing recognition algorithm is used. For example, if a commodity image contains logo information dodo, which is similar to logo information diox, it is easy to be recognized as the same type of commodity image. If diox is a non-recommended commodity image, dodo cannot be recommended normally. The results obtained by the existing recognition algorithm cause "accidental damage" to the commodity image, and manual review is required, which wastes a lot of manpower costs. Summary of the invention

[0004] In view of the above problems, the present invention is proposed to provide a method and device for identifying product images that overcome the above problems or at least partially solve the above problems.

[0005] According to one aspect of the present invention, a method for identifying a product image is provided, which comprises:

[0006] Intercept a partial picture of the target product picture and extract a first picture feature vector;

[0007] Inputting the first image feature vector into a pre-trained first designated model, and inputting the second image feature vector into a designated intermediate layer of the first designated model to perform fusion processing to obtain a corresponding first output result; the second image feature vector is extracted according to the target product image;

[0008] According to the first output result, it is determined whether to recommend the target product image.

[0009] According to another aspect of the present invention, there is provided a commodity image recognition device, comprising:

[0010] An extraction module, adapted to intercept a partial image from a target product image and extract a first image feature vector;

[0011] A first prediction module, adapted to input the first image feature vector into a pre-trained first specified model and input a second image feature vector into a specified intermediate layer of the first specified model for fusion processing to obtain a corresponding first output result; the second image feature vector is extracted from the target product image;

[0012] A first recommendation module, adapted to determine whether to recommend the target product image according to the first output result.

[0013] According to another aspect of the present invention, an electronic device is provided, including: a processor, a memory, a communication interface, and a communication bus, and the processor, the memory, and the communication interface complete communication with each other through the communication bus;

[0014] The memory is used to store at least one executable instruction, and the executable instruction causes the processor to execute the operations corresponding to the above-mentioned product image recognition method.

[0015] According to still another aspect of the present invention, a computer storage medium is provided, and at least one executable instruction is stored in the storage medium, and the executable instruction causes the processor to execute the operations corresponding to the above-mentioned product image recognition method.

[0016] According to the product image recognition method and device of the present invention, the features of the partial image in the target product image are fused with the features of the overall image of the target product image, which not only retains the representativeness of the features of the partial image, but also retains the relevance and connectivity of the overall image, ensures the accuracy of the obtained result, and improves the recommendation effect of the product image.

[0017] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the following specifically illustrates the specific embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0019] Figure 1 The flowchart of the product image recognition method according to an embodiment of the present invention is shown;

[0020] Figure 2a Shows a schematic diagram of the convolutional block structure;

[0021] Figure 2b Shows a schematic diagram of the feature block structure;

[0022] Figure 2c Shows a schematic diagram of the first specified model structure;

[0023] Figure 3 Shows a flowchart of the commodity picture recognition method according to another embodiment of the present invention;

[0024] Figure 4 Shows a schematic diagram of the second specified model structure;

[0025] Figure 5 Shows a functional block diagram of the commodity picture recognition device according to an embodiment of the present invention;

[0026] Figure 6 Shows a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. Detailed implementation manners

[0027] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.

[0028] Figure 1 Shows a flowchart of the commodity picture recognition method according to an embodiment of the present invention. As Figure 1 shown, the commodity picture recognition method specifically includes the following steps:

[0029] Step S101, intercept a partial picture in the target commodity picture and extract a first picture feature vector.

[0030] The target commodity picture includes identification information of the commodity, such as identification information such as the logo of the commodity and the brand trademark of the commodity. The identification information can accurately help the user determine the classification of the commodity and improve the user's discrimination of the commodity.

[0031] In this embodiment, the target product image is recognized based on the representative identification information in the partial image. Specifically, the partial image contains the identification information of the product. First, the target product image is recognized to obtain the partial image representing the identification information of the product in the target product image, and the partial image is intercepted. The first image feature vector is extracted based on the intercepted partial image. The first image feature vector is a feature vector representing the identification information of the product, such as the coordinate position of the logo identification information of the target product image, the size of the partial image, the information in the partial image, etc., which is representative of the product identification. For example, the logo information in the target product image, that is, the partial image containing "dodo", is intercepted, and the corresponding first image feature vector is extracted.

[0032] Step S102: Input the first image feature vector into the first specified model pre-trained, and input the second image feature vector into the specified intermediate layer of the first specified model for fusion processing to obtain the corresponding first output result.

[0033] In this embodiment, the second image feature vector is a feature vector extracted based on the whole of the target product image, which contains the correlation of all features in the target product image and the connectivity between features. The first specified model includes at least one convolutional layer, at least one convolutional block, at least one feature block, etc. Among them, the convolutional block fuses the intermediate output results obtained by the same input object through different convolutional layers to obtain the output result of the convolutional block. As Figure 2a shown, the same input object is input into the 1*1 convolutional layer in the first row, and an intermediate output result a is obtained through the 3*3 convolutional layer and the 1*1 convolutional layer, and it is input into the 1*1 convolutional layer in the second row to obtain another intermediate output result b. The intermediate output result a and the intermediate output result b are fused to obtain the final output result of the convolutional block. The convolutional block can increase the receptive field, reduce feature loss, fuse feature vectors, extract them in multiple dimensions, and improve the accuracy of the overall model. The feature block fuses the intermediate output result obtained by the input object through the convolutional layer with the input object to obtain the output result of the feature block. As Figure 2b shown, the same input object is input into the 1*1 convolutional layer in the first row, and an intermediate output result c is obtained through the 3*3 convolutional layer and the 1*1 convolutional layer, and it is directly fused with the intermediate output result c to obtain the final output result of the feature block. The feature block fuses feature vectors, can extract features in multiple dimensions, and improve the precision rate of the overall model. Compared with the convolutional layer, the convolutional block and the feature block can improve the precision rate of the overall model and are more suitable for the recognition of product images.

[0034] The first specified model of this embodiment can be as Figure 2cAs shown, it includes convolutional layers, feature blocks, convolutional blocks, etc. of multiple different sizes to improve the accuracy of the overall model. This embodiment includes two stages of input. In addition to inputting the first picture feature vector into the first specified model pre-trained to obtain the first intermediate result, this embodiment also receives the second picture feature vector processed by the convolutional layer input at the specified intermediate layer of the first specified model and performs fusion processing with the first intermediate result to obtain the corresponding first output result. The first output result not only retains the representativeness of the first picture feature but also retains the overall relevance and connectivity in the target commodity picture, so that the obtained output result ensures the accuracy of identifying the local picture and the overall accuracy of the target commodity picture. The setting of the first specified model can be set as Figure 2c shown, or other convolutional layers plus convolutional block and feature block settings can be adopted, which are not limited here. The second picture feature can be input into the last layer of the first specified model or other intermediate layers of the first specified model to achieve the effect of fusion processing, which is not limited here.

[0035] This embodiment also includes the training of the preset first specified model. Specifically, the first input data and the first annotation information of the first training sample can be pre-constructed. For example, more than 100,000 first training samples are constructed. The first input data includes the first picture feature vector and the second picture feature vector of the sample commodity picture, such as the coordinate position, local picture size, local picture information, and overall information of the sample commodity picture of the logo identification information. The first annotation information includes positive sample annotation information and negative sample annotation information. The first annotation information is annotated according to the historical display data of the sample commodity picture, and can be annotated in combination with the display effect, user purchase volume, click-through rate, etc. of the sample commodity picture. For example, the positive sample annotation information is recommended, and the negative sample annotation information is not recommended. Input the first input data of the constructed first training sample into the first specified model to be trained for training, compare the obtained output result with the first annotation information, and adjust the training parameters of the first specified model according to the comparison result to obtain the trained first specified model.

[0036] In this embodiment, the training of the first specified model can be based on a GPU cluster server to improve the processing speed. The cross-validation training method can be adopted, and the ratio of the training set: validation set: test set can be 8:1:1. Training frameworks such as pytorh and TensorFlow can be used, and the Adam iterative optimizer is used. The initial learning rate is set to 0.0001. Every 10 epochs (one complete training of the training data is 1 epoch), the learning rate is multiplied by 0.1, that is, it is decreased by 10 times. It is expected to train for 100 epochs. When the loss function curve tends to be stable, the SGD optimizer is changed, and the learning rate remains unchanged. Continue to train for 20 - 30 epochs until the loss curve tends to be stable, then terminate the training and save the result file of the first specified model for subsequent recognition of target commodity pictures. The above is for illustrative purposes, and specifically, the framework, method, ratio, etc. during training can be adjusted according to the actual situation and are not limited here.

[0037] Step S103, according to the first output result, determine whether to recommend the target commodity picture.

[0038] According to the first output result of the first specified model, if the first output result is a recommendation, it is determined to use the target commodity picture for recommendation. If the first output result is not a recommendation, the target commodity picture is not used for commodity display.

[0039] Furthermore, this embodiment can also be used in the scenario of whether the commodity pictures on the commodity platform comply with the display rules. For example, the sample commodity pictures are labeled based on the display rules, trained based on the first specified model, and the target commodity pictures are input into the first specified model. It can be determined whether the target commodity pictures comply with the display rules according to the first output result to determine whether the target commodity pictures can be displayed, etc. By using the first specified model to identify the target commodity pictures, the human review cost can be greatly saved. At the same time, problems such as non-compliance with the display rules caused by the target commodity pictures can be effectively handled.

[0040] According to the commodity picture recognition method provided by the present invention, the features of the local pictures in the target commodity picture are fused with the features of the overall picture of the target commodity picture, which not only retains the representativeness of the features of the local pictures but also retains the relevance and connectivity of the overall picture, ensuring the accuracy of the obtained results and improving the recommendation effect of the commodity pictures.

[0041] Figure 3 The flowchart of the commodity picture recognition method according to another embodiment of the present invention is shown. As Figure 3 shown, the commodity picture recognition method specifically includes the following steps:

[0042] Step S301, intercept a partial image from the target product image, and extract the first image feature vector.

[0043] Step S302, input the first image feature vector into a pre-trained first specified model, and input a second image feature vector into a specified intermediate layer of the first specified model for fusion processing to obtain a corresponding first output result.

[0044] The above steps refer to Figure 1 the descriptions of steps S101 - S102 in the embodiment, which will not be elaborated here.

[0045] Step S303, determine whether the first output result is a recommendation for the target product image.

[0046] If the first output result is to recommend a product using the target product image, step S304 can be further executed to make a further judgment on the target product image.

[0047] Step S304, obtain the text information of the target product image, and convert the text information into a text feature vector.

[0048] Considering that in the prior art, only the features of product images are recognized, and other features of products such as clicks on product images, purchase volume based on product images, product prices and other factors are not considered, resulting in a single dimension considered when recommending products, only based on the product image dimension, without considering other dimension factors, and the overall recommendation effect of products is poor.

[0049] In this embodiment, based on the target product image, the corresponding text information is also obtained. The text information includes information related to the product such as the historical click-through rate of the product, historical purchase rate, product category, product price, etc. The text information is converted into a text feature vector for subsequent use. The conversion of the text feature vector is based on the embedding technology, and the text feature vector is extracted from the text information, that is, a low-dimensional vector is obtained after mapping through a neural network.

[0050] Step S305, input the second image feature vector into a pre-trained second specified model, and input the text feature vector into a specified intermediate layer of the second specified model for splicing processing and / or vector flattening processing to obtain a second output result.

[0051] This embodiment includes inputs in two stages. Among them, the second picture feature vector is input into a pre-trained second specified model to obtain a second intermediate result. Alternatively, the first picture feature vector can also be input into the pre-trained second specified model to obtain a second intermediate result. Then, at a specified intermediate layer of the second specified model, the input text feature vector processed by the deep neural network is received, and it is subjected to operations such as splicing and vector flattening with the second intermediate result to obtain the corresponding second output result. The splicing process is to splice the second intermediate result of the second picture feature vector with the text feature vector. The vector flattening process is to fold and flatten the spliced vector into a one-dimensional vector, thereby obtaining the second output result.

[0052] The second specified model can be, for example, Figure 4 as shown, and includes at least one convolutional layer, at least one convolutional block, at least one feature block, etc., and can also be set according to the actual situation, which is not limited here. The convolutional block and the feature block refer to the description in step S102 and will not be elaborated here. The second specified module can be pre-trained. Specifically, the second input data and second annotation information of the second training sample can be pre-constructed. For example, more than 100,000 second training samples are constructed. The second input data of the second training sample includes the second picture feature vector and text feature vector of the sample commodity picture. For example, the second picture feature vector extracted from the overall information of the sample commodity picture and the text feature vector transformed from the text information corresponding to the sample commodity picture. The second annotation information includes positive sample annotation information and negative sample annotation information. The second annotation information is annotated according to the historical display data of the sample commodity picture, and can be annotated in combination with the display effect, user purchase volume, click-through rate, etc. of the sample commodity picture. For example, the positive sample annotation information is "recommended", and the negative sample annotation information is "not recommended". The second input data of the constructed second training sample is input into the second specified model to be trained, and the obtained output result is compared with the second annotation information. According to the comparison result, the training parameters of the second specified model are adjusted to obtain the trained second specified model. The training process of the second specified model can refer to the training framework, method, and the proportions of each training set, validation set, and test set of the first specified model, which will not be elaborated here.

[0053] Step S306: Determine whether to recommend the target commodity picture according to the second output result.

[0054] Based on the first output result, through further identification, the second output result is obtained. If the second output result is "recommended", it is determined to use the target commodity picture for recommendation. If the second output result is "not recommended", the target commodity picture is not used for commodity display.

[0055] Furthermore, this embodiment can be used for product picture recommendation on a product platform, and can assist in improving the conversion rate of products by using product pictures and text information.

[0056] According to the product picture recognition method provided by the present invention, by combining the target product picture and text information, the text feature vector is spliced with the intermediate output result of the target product picture, and the second output result is integrally obtained, avoiding problems such as the single recognition accuracy of only product pictures.

[0057] Figure 5 The functional block diagram of a product picture recognition device according to an embodiment of the present invention is shown. As Figure 5 shown, the product picture recognition device includes the following modules:

[0058] An extraction module 510, adapted to intercept a partial picture in the target product picture and extract a first picture feature vector;

[0059] A first prediction module 520, adapted to input the first picture feature vector into a pre-trained first specified model, and input a second picture feature vector into a specified intermediate layer of the first specified model for fusion processing to obtain a corresponding first output result; the second picture feature vector is extracted according to the target product picture;

[0060] A first recommendation module 530, adapted to determine whether to recommend the target product picture according to the first output result.

[0061] Optionally, the partial picture includes the identification information of the product;

[0062] The extraction module 510 is further adapted to: identify the partial picture representing the identification information of the product in the target product picture and intercept the partial picture.

[0063] Optionally, the first specified model includes at least one convolutional layer, at least one convolutional block, and / or at least one feature block;

[0064] The first prediction module 520 is further adapted to:

[0065] Input the first picture feature vector into a pre-trained first specified model to obtain a first intermediate result;

[0066] Receive the input second picture feature vector processed by the convolutional layer in the specified intermediate layer of the first specified model, and perform fusion processing with the first intermediate result to obtain a corresponding first output result.

[0067] Optionally, the device further includes:

[0068] An acquisition module 540, adapted to acquire the text information of the target product picture and convert the text information into a text feature vector; the text information includes the historical click-through rate of the product, the historical purchase rate, the product category, and / or the product price;

[0069] A second prediction module 550, adapted to input the second picture feature vector into a pre-trained second specified model, and input the text feature vector at a specified intermediate layer of the second specified model for splicing processing and / or vector flattening processing to obtain a second output result;

[0070] A second recommendation module 560, adapted to determine whether to recommend the target product picture according to the second output result.

[0071] Optionally, the second specified model includes at least one convolutional layer with a specified number of layers, at least one convolutional block, and / or at least one feature block;

[0072] The second prediction module 550 is further adapted to:

[0073] Input the second picture feature vector into a pre-trained second specified model to obtain a second intermediate result;

[0074] Receive the input text feature vector processed by the deep neural network at a specified intermediate layer of the second specified model, and perform splicing processing and / or vector flattening processing with the second intermediate result to obtain the corresponding second output result.

[0075] Optionally, the convolutional block fuses the intermediate output results obtained by the same input object through different convolutional layers to obtain the output result of the convolutional block; the feature block fuses the intermediate output result obtained by the input object through the convolutional layer with the input object to obtain the output result of the feature block.

[0076] Optionally, the device further includes:

[0077] A training module 570, adapted to train and obtain a first specified model and / or a second specified model;

[0078] The training module 570 is further adapted to:

[0079] Construct the first input data and the first annotation information of the first training sample, input the first input data of the first training sample into the to-be-trained first specified model for training, compare the obtained output result with the first annotation information, and adjust the training parameters of the first specified model according to the comparison result to obtain the trained first specified model;

[0080] Among them, the first input data of the first training sample includes the first picture feature vector and the second picture feature vector of the sample commodity picture; the first annotation information includes positive sample annotation information and negative sample annotation information; the first annotation information is annotated according to the historical display data of the sample commodity picture.

[0081] And / or,

[0082] Construct the second input data and the second annotation information of the second training sample, input the second input data of the second training sample into the second specified model to be trained for training, compare the obtained output result with the second annotation information, and adjust the training parameters of the second specified model according to the comparison result to obtain the trained second specified model.

[0083] Among them, the second input data of the second training sample includes the second picture feature vector and the text feature vector of the sample commodity picture; the second annotation information includes positive sample annotation information and negative sample annotation information; the second annotation information is annotated according to the historical display data of the sample commodity picture.

[0084] The descriptions of the above modules refer to the corresponding descriptions in the method embodiments and will not be elaborated here.

[0085] This application also provides a non-volatile computer storage medium, and the computer storage medium stores at least one executable instruction, and this computer executable instruction can execute the commodity picture recognition method in any of the above method embodiments.

[0086] Figure 6 The structural schematic diagram of an electronic device according to an embodiment of the present invention is shown. The specific implementation of the electronic device in the specific embodiments of the present invention is not limited.

[0087] As Figure 6 shown, the electronic device may include: a processor 602, a communication interface 604, a memory 606, and a communication bus 608.

[0088] Among them:

[0089] The processor 602, the communication interface 604, and the memory 606 complete mutual communication through the communication bus 608.

[0090] The communication interface 604 is used to communicate with network elements of other devices such as clients or other servers.

[0091] The processor 602 is used to execute the program 610, and specifically can execute the relevant steps in the above commodity picture recognition method embodiments.

[0092] Specifically, program 610 may include program code that includes computer operation instructions.

[0093] Processor 602 may be a central processing unit (CPU), or a specific application integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention. One or more processors included in the electronic device may be of the same type of processor, such as one or more CPUs; or may be of different types of processors, such as one or more CPUs and one or more ASICs.

[0094] Memory 606 is used to store program 610. Memory 606 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk memory.

[0095] Program 610 may specifically be used to cause processor 602 to execute the product picture recognition method in any of the above method embodiments. For the specific implementation of each step in program 610, reference may be made to the corresponding steps and descriptions in the corresponding units in the above product picture recognition embodiments, which will not be elaborated here. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described devices and modules may refer to the corresponding process descriptions in the foregoing method embodiments, which will not be repeated here.

[0096] The algorithms and displays provided herein are not inherently related to any particular computer, virtual system, or other device. Various general-purpose systems may also be used in conjunction with the teachings provided herein. The structure required to construct such a system will be apparent from the above description. In addition, the present invention is not directed to any particular programming language. It should be understood that the content of the present invention described herein can be implemented using various programming languages, and the description of the specific language above is to disclose the best mode of the present invention.

[0097] In the specification provided herein, a large number of specific details are set forth. However, it can be understood that the embodiments of the present invention may be practiced without these specific details. In some instances, well-known methods, structures, and technologies have not been shown in detail so as not to obscure the understanding of this specification.

[0098] Similarly, it should be understood that, for the purpose of streamlining the present disclosure and facilitating the understanding of one or more of the various inventive aspects, in the foregoing description of the exemplary embodiments of the present invention, the various features of the present invention are sometimes grouped together into a single embodiment, figure, or description thereof. However, the disclosed method should not be construed as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as reflected in the claims, the inventive aspects lie in less than all the features of the single foregoing disclosed embodiment. Thus, the claims following the detailed description are hereby expressly incorporated into the detailed description, with each claim standing on its own as a separate embodiment of the present invention.

[0099] Those skilled in the art will appreciate that the modules in the devices in the embodiments can be adaptively changed and disposed in one or more devices different from the embodiments. The modules or units or components in the embodiments can be combined into one module or unit or component, and in addition, they can be divided into multiple sub-modules or sub-units or sub-components. Except that at least some of such features and / or processes or units are mutually exclusive, any combination can be adopted to combine all the features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all the processes or units of any method or device so disclosed. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) can be replaced by an alternative feature that provides the same, equivalent, or similar purpose.

[0100] In addition, those skilled in the art can understand that, although some of the embodiments described herein include certain features included in other embodiments rather than other features, the combination of the features of different embodiments means that it is within the scope of the present invention and forms different embodiments. For example, in the claims, any one of the claimed embodiments can be used in any combination.

[0101] The various component embodiments of the present invention can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that a microprocessor or a digital signal processor (DSP) can be used in practice to implement some or all of the functions of some or all of the components in the commodity picture recognition device according to the embodiments of the present invention. The present invention can also be implemented as a device or device program for executing part or all of the methods described herein (for example, a computer program and a computer program product). Such a program for implementing the present invention can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.

[0102] It should be noted that the above embodiments are illustrative of the present invention and not restrictive thereof, and alternative embodiments can be designed by those skilled in the art without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in the claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present invention can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In the unit claims listing several devices, several of these devices may be embodied by the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words may be interpreted as names.

Claims

1. A method for identifying product pictures, characterized in that the method include: Intercepting a partial picture of the target product picture and extracting a first picture feature vector; the partial picture includes identification information of the product; Inputting the first image feature vector into a pre-trained first designated model, and inputting the second image feature vector into a designated intermediate layer of the first designated model to perform fusion processing to obtain a corresponding first output result; the second image feature vector is obtained based on the overall extraction of the target product image, and includes the correlation of all features in the target product image and the connectivity between the features; According to the first output result, it is determined whether to recommend the target product image.

2. The method according to claim 1, wherein The capturing of a partial image in the target product image specifically includes: identifying a partial image representing identification information of the product in the target product image, and capturing the partial image.

3. The method according to claim 1, wherein The first specified model includes at least one convolutional layer, at least one convolutional block and / or at least one feature block; The step of inputting the first image feature vector into a pre-trained first designated model, and inputting the second image feature vector into a designated intermediate layer of the first designated model for fusion processing to obtain a corresponding first output result further includes: Inputting the first picture feature vector into a pre-trained first designated model to obtain a first intermediate result; The second image feature vector processed by the convolution layer is received as input at the designated intermediate layer of the first designated model, and is fused with the first intermediate result to obtain a corresponding first output result.

4. The method according to claim 1, characterized in that, If the target product image is recommended according to the first output result, the method further includes: Acquire text information of the target product image, and convert the text information into a text feature vector; the text information includes a historical click rate, a historical purchase rate, a product category and / or a product price; Inputting the second image feature vector into a pre-trained second designated model, and inputting the text feature vector into a designated intermediate layer of the second designated model to perform concatenation processing and / or vector flattening processing to obtain a second output result; According to the second output result, it is determined whether to recommend the target product image.

5. The method according to claim 4, wherein The second specified model includes at least one convolutional layer with a specified number of layers, at least one convolutional block and / or at least one feature block; The step of inputting the second image feature vector into a pre-trained second designated model, and inputting the text feature vector into a designated intermediate layer of the second designated model to perform concatenation processing and / or vector flattening processing to obtain a second output result further includes: Inputting the second image feature vector into a pre-trained second designated model to obtain a second intermediate result; The input text feature vector processed by the deep neural network is received at the designated intermediate layer of the second designated model, and is concatenated and / or vector flattened with the second intermediate result to obtain a corresponding second output result.

6. The method according to claim 3 or 5, characterized in that, The convolutional block fuses the intermediate output results obtained from the same input object through different convolutional layers to obtain the output result of the convolutional block; the feature block fuses the intermediate output result obtained from the input object through the convolutional layer with the input object to obtain the output result of the feature block.

7. The method according to claim 3 or 5, characterized in that, The method further includes: Training to obtain the first specified model and / or the second specified model; The training to obtain the first specified model and / or the second specified model specifically is: Construct the first input data and the first annotation information of the first training sample, input the first input data of the first training sample into the first specified model to be trained for training, compare the obtained output result with the first annotation information, and adjust the training parameters of the first specified model according to the comparison result to obtain the trained first specified model; Wherein, the first input data of the first training sample includes the first picture feature vector and the second picture feature vector of the sample commodity picture; the first annotation information includes positive sample annotation information and negative sample annotation information; the first annotation information is annotated according to the historical display data of the sample commodity picture; and / or, Construct the second input data and the second annotation information of the second training sample, input the second input data of the second training sample into the second specified model to be trained for training, compare the obtained output result with the second annotation information, and adjust the training parameters of the second specified model according to the comparison result to obtain the trained second specified model; Wherein, the second input data of the second training sample includes the second picture feature vector and the text feature vector of the sample commodity picture; the second annotation information includes positive sample annotation information and negative sample annotation information; the second annotation information is annotated according to the historical display data of the sample commodity picture.

8. A commodity picture recognition device, characterized in that, The apparatus includes: An extraction module, adapted to intercept a partial picture in the target commodity picture and extract the first picture feature vector; the partial picture contains the identification information of the commodity; A first prediction module, adapted to input the first picture feature vector into the first specified model pre-trained, and input the second picture feature vector into the specified intermediate layer of the first specified model for fusion processing to obtain the corresponding first output result; the second picture feature vector is extracted from the whole of the target commodity picture and contains the relevance of all features in the target commodity picture and the connectivity between features; A first recommendation module, adapted to determine whether to recommend the target commodity picture according to the first output result.

9. An electronic device, comprising: A processor, a memory, a communication interface and a communication bus, and the processor, the memory and the communication interface complete communication with each other through the communication bus; The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform the operations corresponding to the commodity picture recognition method according to any one of claims 1-7.

10. A computer storage medium, in which at least one executable instruction is stored, and the executable instruction enables the processor to perform the operations corresponding to the commodity picture recognition method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Commodity group recommendation method and device, electronic equipment and readable storage medium

    CN110490637A

  • Identifier identification method, device and system

    CN111241893A

  • Commodity recommendation method and device, terminal equipment and storage medium

    CN112529663A