Image retrieval method, device, equipment and medium based on multi-task learning
Through the image retrieval method based on multi-task learning, combined with the original image, product main cropping map and product attributes for model training, the problems of low image retrieval accuracy and time-consuming in the existing technology are solved, and more efficient image retrieval and recommendation effects are achieved.
Patent Information
- Application Number
- CN202111580851.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-22
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2041-12-22
AI Technical Summary
Existing image retrieval schemes are difficult to capture slight differences in local details of the image, it is difficult to accurately describe product attributes, only detect product subjects and lose full image information, and the model vector calculation similarity accuracy is low and time-consuming during inference.
The image retrieval method based on multi-task learning is adopted, and a twin neural network combines the original graph and the product main cropping graph, and combines the product attributes for model training to generate similarity feature vectors for related search and recommendation services.
It realizes the mining of image features while combining product attributes, generate similarity feature vectors, improve the accuracy and recall rate of the search model, meet user needs, and enhance the competitiveness of the platform.
Smart Images

Figure CN114282037B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to an image retrieval method, apparatus, device and medium based on multi-task learning. Background Art
[0002] With the rapid development of e-commerce, product search plays a vital role in order to accurately meet user needs, improve product click-through rate and conversion rate, and reduce operating costs.
[0003] With the vigorous development of computer vision technology in the field of deep learning, more and more e-commerce platforms have introduced this technology to provide a powerful supplement to traditional product searches, among which retrieval technology based on similarity calculation is particularly prominent.
[0004] However, there are still some problems that need to be solved in the current image retrieval scheme:
[0005] 1. It can only calculate the similarity of image content, and it is difficult to capture the slight differences in local details;
[0006] 2. The image content cannot accurately describe the attributes of the product (such as the brand, name, applicable population, applicable season, etc.);
[0007] 3. Only the main body of the product is detected, but other information of the entire image is lost;
[0008] 4. Directly combining multiple model vectors to calculate similarity during inference has low accuracy and is time-consuming. Summary of the invention
[0009] In order to solve at least one of the problems mentioned in the above background technology, the present application provides an image retrieval method, device, equipment and medium based on multi-task learning, which can mine image features while combining product attributes to generate similarity feature vectors for related retrieval and recommendation services, better meet user needs and enhance platform competitiveness.
[0010] The specific technical solutions provided by the embodiments of this application are as follows:
[0011] In a first aspect, a multi-task learning-based image retrieval method is provided, comprising:
[0012] Get the input product image;
[0013] Performing image preprocessing on the product image to obtain an original image and a cropped image of the product body;
[0014] Importing the original image and the cropped image of the product body into a twin neural network;
[0015] Performing model training according to the original image, the cropped image of the commodity body, and the commodity attributes of the commodity image to obtain a trained image model;
[0016] The trained image model is deployed to a server to output a final vector of the product image for image vector similarity retrieval.
[0017] Furthermore, the model training is performed according to the original image, the cropped image of the commodity body and the commodity attributes of the commodity image to obtain a trained image model, which includes the following steps:
[0018] Vector generation step: generating a full image vector and a cropped image vector according to the original image and the cropped image of the commodity body;
[0019] Feature interaction step: performing feature interaction on the full image vector and the cropped image vector to obtain a feature vector;
[0020] Loss calculation step: according to the feature vector and the commodity attributes of the commodity image, the loss is calculated by a loss function, and back propagation is performed according to the calculated loss to optimize and update the weights of each node in the twin neural network;
[0021] Repeat the vector generation step, feature interaction step, and loss calculation step until the image model converges to obtain the trained image model.
[0022] Furthermore, the commodity attributes include at least one of the commodity name, the commodity brand, the applicable population of the commodity, and the applicable season of the commodity;
[0023] The loss calculation step specifically includes:
[0024] Calculating a multi-label classification loss by at least one first loss function according to the commodity attributes of the commodity image to constrain the vector space of the feature vector;
[0025] The metric distance loss is calculated according to the feature vector through one or more second loss functions to further expand the inter-class distance and reduce the intra-class distance.
[0026] Furthermore, the feature interaction step specifically includes:
[0027] Feature interaction is performed on the full image vector and the cropped image vector, and addition, subtraction, dot product and variance operations are performed bit by bit, and then a normalization operation is performed to obtain a feature vector.
[0028] Furthermore, the step of deploying the trained image model to a server to output a final vector of the product image for image vector similarity retrieval further includes:
[0029] extracting the final vector of the product image from the trained image model at regular intervals, and saving it to a vector similarity search engine for real-time image vector similarity retrieval;
[0030] The vector similarity search engine includes a Milvus vector search engine.
[0031] Furthermore, the image preprocessing of the product image to obtain the original image and the cropped image of the product body specifically includes:
[0032] Performing an enhancement operation on the entire image of the product to obtain an original image of the product image;
[0033] Detecting the product image and cropping the product body, and then performing an enhancement operation to obtain a cropped image of the product body of the product image;
[0034] The enhancement operation includes at least one of enlarging, reducing, cropping and rotating.
[0035] Furthermore, the detecting the commodity image and cropping the commodity body includes:
[0036] Detecting a commodity main frame of the commodity image by a one-stage detection algorithm or a two-stage detection algorithm;
[0037] The commodity body is cut out according to the commodity body frame.
[0038] In a second aspect, an image retrieval device based on multi-task learning is provided, the device comprising:
[0039] A receiving module, used to obtain an input product image;
[0040] An image preprocessing module, used to perform image preprocessing on the product image to obtain an original image and a cropped image of the product body;
[0041] A control module, used for importing the original image and the cropped image of the commodity body into a twin neural network;
[0042] A model training module, used to perform model training according to the original image, the cropped image of the commodity body and the commodity attributes of the commodity image to obtain a trained image model;
[0043] A management module is used to deploy the trained image model to a server to output a final vector of the product image for image vector similarity retrieval.
[0044] According to a third aspect, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the multi-task learning-based image retrieval method when executing the computer program.
[0045] In a fourth aspect, a computer-readable storage medium is provided, storing computer-executable instructions, wherein the computer-executable instructions are used to execute the image retrieval method based on multi-task learning.
[0046] The embodiments of the present application have the following beneficial effects:
[0047] The embodiments of the present application provide an image retrieval method, apparatus, device and medium based on multi-task learning. During model training, the product attributes can be used as the vector space of a special branch constraint model, so that the model can not only recognize the image content, but also aggregate image vectors of the same attributes as much as possible, and increase the similarity distance between different vectors, thereby mining image features while combining product attributes to generate similarity feature vectors for related retrieval and recommendation services, thereby improving the accuracy and recall rate of the retrieval model. Different loss function combinations can be used to achieve end-to-end model training of product attributes and image content, and good compatibility between the input parameter differences and model output of training and reasoning, thereby better meeting user needs and improving platform competitiveness. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0049] Figure 1 The overall flow chart of the image retrieval method based on multi-task learning provided by the embodiment of the present application is shown;
[0050] Figure 2 A specific flow chart of an image retrieval method based on multi-task learning according to an embodiment of the present application is shown;
[0051] Figure 3 A schematic diagram showing the structure of an image retrieval device based on multi-task learning provided in an embodiment of the present application is shown;
[0052] Figure 4 An exemplary system is shown that can be used to implement the various embodiments described in this application. DETAILED DESCRIPTION
[0053] In order to make the purpose, technical solutions and advantages of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0054] It should be understood that in the description of the present application, unless the context clearly requires otherwise, words such as "include", "comprises", and the like in the entire specification and claims should be interpreted as having an inclusive meaning rather than an exclusive or exhaustive meaning; that is, the meaning of "including but not limited to".
[0055] It should also be understood that the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. In addition, in the description of this application, unless otherwise specified, "plurality" means two or more.
[0056] It should be noted that the terms "S1", "S2", etc. are only used for the purpose of describing the steps, and do not specifically refer to the order or sequence, nor are they used to limit the present application. They are only for the convenience of describing the method of the present application, and cannot be understood as indicating the order of the steps. In addition, the technical solutions between the various embodiments can be combined with each other, but they must be based on the ability of ordinary technicians in this field to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present application.
[0057] Embodiment 1
[0058] This application provides an image retrieval method based on multi-task learning. Figure 1 ,include:
[0059] S1. Obtain an input product image;
[0060] S2, performing image preprocessing on the product image to obtain an original image and a cropped image of the product body;
[0061] S3, importing the original image and the cropped image of the commodity body into the twin neural network;
[0062] S4, performing model training according to the original image, the cropped image of the commodity body, and the commodity attributes of the commodity image to obtain a trained image model;
[0063] S5. Deploy the trained image model to a server to output a final vector of the product image for image vector similarity retrieval.
[0064] Specifically, the Siamese network refers to a "Siamese neural network". In a neural network, "Siamese" is achieved by sharing weights. The Siamese neural network includes two neural networks, and the two neural networks share weights / weights. The two neural networks can be ViT (Vision Transformer) neural networks or CNN (convolutional neural networks). The two neural networks can be neural networks of the same type, or they can be different neural networks (such as one is ViT and the other is CNN). If the two neural networks are different, the entire neural network is called a pseudo-siamese network, a pseudo-Siamese neural network. Among them, the two neural networks are used to extract the features of the original image and the cropped image of the main body of the product, respectively, and fuse them using feature interaction. The following is a further explanation of the example of two neural networks sharing weights. It should be noted that the backbone (backbone neural network) can be flexibly replaced according to actual needs. For example, different backbones are used for the original image and the cropped image of the main body of the product to further enhance the feature expression ability of the model.
[0065] For example, the backbone can select a suitable neural network according to the requirements (speed or accuracy). For example, if the inference speed and deployment performance are very high, you can choose a smaller network such as MobileNet (lightweight CNN) or resnet50 (residual network); if the inference accuracy is very high and the performance is guaranteed by a powerful backend server, you can use a larger network such as ViT (Vision Transformer), MLP-mixer, seResnet, etc.
[0066] Specifically, the input product image is received and can be read locally or downloaded remotely.
[0067] In some embodiments, S2 specifically includes:
[0068] S21. Perform an enhancement operation on the entire image of the product to obtain an original image of the product image.
[0069] S22: Detect the product image and crop the product body, and then perform an enhancement operation to obtain a cropped image of the product body of the product image.
[0070] The enhancement operation includes at least one of enlargement, reduction, cropping and rotation.
[0071] In some embodiments, S22 further includes: detecting a commodity body frame of the commodity image by using a one-stage detection algorithm or a two-stage detection algorithm, and cropping the commodity body according to the commodity body frame.
[0072] Specifically, during image preprocessing, one branch only performs enhancement operations (such as zooming in, zooming out, cropping, etc.) on the entire product image; the other branch first detects the product main frame through a detection algorithm, then crops according to the detected product main frame, and then enhances the crop sub-image, and finally obtains the product main cropped image. Among them, the above-mentioned one-stage detection algorithm can include the YOLO algorithm series (such as yolov5); the above-mentioned two-stage detection algorithm can include the RCNN (Region with CNN feature) series.
[0073] The following will be combined Figure 2 To further elaborate:
[0074] In some embodiments, S4 specifically includes the following steps:
[0075] The vector generation step, the feature interaction step, the loss calculation step and the vector generation step, the feature interaction step and the loss calculation step are repeated until the image model converges to obtain the trained image model.
[0076] Specifically, the vector generation step refers to generating a full-image vector and a cropped image vector according to the original image and the cropped image of the product body; the feature interaction step refers to performing feature interaction on the full-image vector and the cropped image vector to obtain a feature vector; the loss calculation step refers to calculating the loss through a loss function based on the feature vector and the product attributes of the product image, performing backpropagation based on the calculated loss, and optimizing and updating the weights of each node in the twin neural network.
[0077] For example, during the model training process, hundreds of thousands to millions of images are needed for training. Each iterative step starts with the input of the image, generates the full image vector and the cropped image vector according to the input image, and then performs feature interaction to obtain the feature vector. Then, based on the feature vector and the corresponding product attributes, the loss is calculated, and the gradient of each node is calculated in reverse based on the loss, and then the weights and bias of each node are optimized and updated. Finally, the loss and learning metrics are used to determine whether the model has converged.
[0078] In some embodiments, the feature interaction step specifically includes:
[0079] The feature interaction is performed on the full image vector and the cropped image vector, and addition, subtraction, dot product and variance operations are performed bit by bit, and then normalization is performed to obtain the feature vector.
[0080] For example, normalize(concat(vec_a+vec_b,(vec_a-vec_b).abs(),(vec_a-vec_b).sqrt(),vec_a*vec_b)) can be used to perform addition, subtraction, variance, and dot product operations on vector a and vector b to obtain a vector group, and then perform normalization operations to obtain a eigenvector.
[0081] In some embodiments, the commodity attributes include at least one of the name of the commodity, the brand of the commodity, the applicable population of the commodity, and the applicable season of the commodity. Based on this, the loss calculation step specifically includes: calculating the multi-label classification loss through at least one first loss function according to the commodity attributes of the commodity image to constrain the vector space of the feature vector; calculating the metric distance loss through one or more second loss functions according to the feature vector to further expand the inter-class distance and reduce the intra-class distance.
[0082] Specifically, different loss functions are needed to constrain the vector space of feature vectors. The attribute information of the product can be used to perform multi-label classification through the first loss function, and images with the same attributes can be grouped together as much as possible to increase the similarity distance between different vectors. Exemplarily, features should have intra-class compactness and inter-class separability. Intra-class compactness means that features of the same class should be grouped together, and inter-class separability means that features can be easily divided into different categories. Features of the same type of samples are grouped together, and the intervals between features of different classes are larger, making it easier to classify features, improving the classification ability of the network model, and improving robustness. This can be achieved through loss functions such as center loss.
[0083] For example, the anchor-positive-negative triplet can be constructed by using a metric learning method, where loss = (dist(anchor, positive) - dist(anchor, negative) + margin), and the vector space of the feature vector can be constrained by continuously optimizing / reducing the loss, increasing the Euclidean or cosine distance between the anchor-positive vector and the anchor-negative vector, and using the attribute information of the product for multi-label classification.
[0084] Specifically, the first loss function and the second loss function mentioned above should be implemented differently. It should be noted that the first loss function can realize multi-label classification; the second loss function can further expand the distance between classes and reduce the distance within classes. It is also possible to flexibly replace and use multiple losses, such as focal loss, center loss, contrastive loss, etc.
[0085] Specifically, after the model is determined to have converged, the trained image model is obtained. The final_vector can be used as the output of inference. It should be noted that the training input includes the product image and the corresponding product attributes, while the inference input can only include the product image. There is no need to calculate loss or perform back propagation when inferring the output.
[0086] In some embodiments, S5 further includes:
[0087] The final vector of the product image is extracted from the trained image model at regular intervals and saved to the vector similarity search engine for real-time image vector similarity retrieval.
[0088] Among them, the vector similarity search engine includes the Milvus vector search engine.
[0089] Specifically, Milvus vector search engine is designed for approximate nearest neighbor search (ANNS) of massive feature vectors. Compared with operator libraries such as Faiss and SPTAG, Milvus provides a complete vector data update, indexing and query framework. Milvus uses GPU (Nvidia) for indexing and query acceleration, which can greatly improve the performance of a single machine.
[0090] In this embodiment, during model training, product attributes can be used as the vector space of a special branch constraint model, so that the model can not only recognize image content, but also aggregate image vectors of the same attributes as much as possible, and widen the similarity distance between different vectors, thereby mining image features while combining product attributes to generate similarity feature vectors for related retrieval and recommendation services, thereby improving the accuracy and recall rate of the retrieval model; different loss function combinations can be used to achieve end-to-end model training of product attributes and image content, and good compatibility between the input parameter differences of training and reasoning as well as model output, thereby better meeting user needs and improving platform competitiveness.
[0091] Embodiment 2
[0092] Corresponding to the above embodiment, the present application also provides an image retrieval device based on multi-task learning, referring to Figure 3 The device includes: a receiving module, an image preprocessing module, a control module, a model training module and a management module.
[0093] Among them, the receiving module is used to obtain the input product image; the image preprocessing module is used to perform image preprocessing on the product image to obtain the original image and the product body cropped image; the control module is used to import the original image and the product body cropped image into the twin neural network; the model training module is used to perform model training according to the product attributes of the original image, the product body cropped image and the product image to obtain the trained image model; the management module is used to deploy the trained image model to the server to output the final vector of the product image for image vector similarity retrieval.
[0094] Furthermore, the model training module is also used for the vector generation step, the feature interaction step, the loss calculation step, and repeating the vector generation step, the feature interaction step, and the loss calculation step until the image model converges to obtain a trained image model. The specific contents of the vector generation step, the feature interaction step, and the loss calculation step have been fully introduced in the method embodiment, so they will not be repeated here.
[0095] Furthermore, the commodity attributes include at least one of the name of the commodity, the brand of the commodity, the applicable population of the commodity, and the applicable season of the commodity. Based on this, the model training module is also used to calculate the multi-label classification loss through at least one first loss function according to the commodity attributes of the commodity image to constrain the vector space of the feature vector; and to calculate the metric distance loss through one or more second loss functions according to the feature vector to further expand the inter-class distance and reduce the intra-class distance.
[0096] Furthermore, the model training module is also used to perform feature interaction on the full image vector and the cropped image vector, perform bitwise addition, subtraction, dot product and variance operations, and then perform normalization operations to obtain a feature vector.
[0097] Furthermore, the management module is also used to periodically extract the final vector of the product image from the trained image model and save it to a vector similarity search engine for real-time image vector similarity retrieval. The vector similarity search engine includes a Milvus vector search engine.
[0098] Furthermore, the image preprocessing module is further used to perform an enhancement operation on the entire image of the product image to obtain an original image of the product image; and to detect and crop the product body of the product image, and then perform an enhancement operation to obtain a cropped image of the product body of the product image. The enhancement operation includes at least one of zooming in, zooming out, cropping, and rotating.
[0099] Furthermore, the image preprocessing module is further configured to detect a commodity body frame of the commodity image by a one-stage detection algorithm or a two-stage detection algorithm; and to crop the commodity body according to the commodity body frame.
[0100] Embodiment 3
[0101] Corresponding to the above embodiment, the present application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the above-mentioned image retrieval method based on multi-task learning can be implemented.
[0102] like Figure 4 As shown, in some embodiments, the system can be used as any of the above-mentioned electronic devices for the image retrieval method based on multi-task learning in each of the described embodiments. In some embodiments, the system may include one or more computer-readable media (e.g., system memory or NVM / storage device) having instructions and one or more processors (e.g., (one or more) processors) coupled to the one or more computer-readable media and configured to execute instructions to implement the module to perform the actions described in this application.
[0103] For one embodiment, the system control module may include any suitable interface controller to provide any suitable interface to at least one of the processor(s) and / or any suitable device or component in communication with the system control module.
[0104] The system control module may include a memory controller module to provide an interface to the system memory. The memory controller module may be a hardware module, a software module, and / or a firmware module.
[0105] The system memory may be used, for example, to load and store data and / or instructions for the system. For one embodiment, the system memory may include any suitable volatile memory, such as a suitable DRAM. In some embodiments, the system memory may include double data rate type four synchronous dynamic random access memory (DDR4 SDRAM).
[0106] For one embodiment, the system control module may include one or more input / output (I / O) controllers to provide interfaces to the NVM / storage devices and communication interface(s).
[0107] For example, the NVM / storage device may be used to store data and / or instructions. The NVM / storage device may include any suitable non-volatile memory (e.g., flash memory) and / or may include any suitable non-volatile storage device(s) (e.g., one or more hard disk drives (HDDs), one or more compact disk (CD) drives, and / or one or more digital versatile disk (DVD) drives).
[0108] The NVM / storage device may include storage resources that are physically part of the device on which the system is installed, or it may be accessible to the device without being part of the device. For example, the NVM / storage device may be accessed over a network via (one or more) communication interfaces.
[0109] The communication interface(s) may provide an interface for the system to communicate over one or more networks and / or with any other suitable device. The system may wirelessly communicate with one or more components of a wireless network in accordance with any of one or more wireless network standards and / or protocols.
[0110] For one embodiment, at least one of the processor(s) may be packaged together with the logic of one or more controllers of a system control module (e.g., a memory controller module). For one embodiment, at least one of the processor(s) may be packaged together with the logic of one or more controllers of a system control module to form a system-in-package (SiP). For one embodiment, at least one of the processor(s) may be integrated on the same die with the logic of one or more controllers of a system control module. For one embodiment, at least one of the processor(s) may be integrated on the same die with the logic of one or more controllers of a system control module to form a system on chip (SoC).
[0111] In various embodiments, the system may be, but is not limited to: a server, a workstation, a desktop computing device, or a mobile computing device (e.g., a laptop computing device, a handheld computing device, a tablet computer, a netbook, etc.). In various embodiments, the system may have more or fewer components and / or a different architecture. For example, in some embodiments, the system includes one or more cameras, a keyboard, a liquid crystal display (LCD) screen (including a touch screen display), a non-volatile memory port, multiple antennas, a graphics chip, an application specific integrated circuit (ASIC), and a speaker.
[0112] It should be noted that the present application can be implemented in software and / or a combination of software and hardware, for example, can be implemented using an application specific integrated circuit (ASIC), a general purpose computer or any other similar hardware device. In one embodiment, the software program of the present application can be executed by a processor to implement the steps or functions described above. Similarly, the software program of the present application (including relevant data structures) can be stored in a computer-readable recording medium, for example, a RAM memory, a magnetic or optical drive or a floppy disk and similar devices. In addition, some steps or functions of the present application can be implemented using hardware, for example, as a circuit that cooperates with a processor to perform each step or function.
[0113] In addition, a part of the present application may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present application through the operation of the computer. Those skilled in the art should understand that the existence of computer program instructions in computer-readable media includes but is not limited to source files, executable files, installation package files, etc., and accordingly, the way in which computer program instructions are executed by a computer includes but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to the computer.
[0114] Communication media include media by which communication signals containing, for example, computer readable instructions, data structures, program modules, or other data are transmitted from one system to another. Communication media may include guided transmission media such as cables and wires (e.g., fiber optic, coaxial, etc.) and wireless (unguided transmission) media that can propagate energy waves, such as acoustic, electromagnetic, RF, microwave, and infrared. Computer readable instructions, data structures, program modules, or other data may be embodied as a modulated data signal in, for example, a wireless medium such as a carrier wave or similar mechanism such as embodied as part of spread spectrum technology. The term "modulated data signal" refers to a signal whose one or more characteristics are changed or set in such a manner as to encode information in the signal. Modulation may be analog, digital, or a hybrid modulation technique.
[0115] Here, according to one embodiment of the present application, a device is included, which includes a memory for storing computer program instructions and a processor for executing the program instructions, wherein, when the computer program instructions are executed by the processor, the device is triggered to run the methods and / or technical solutions based on the aforementioned multiple embodiments of the present application.
[0116] Embodiment 4
[0117] Corresponding to the above embodiment, the present application also provides a computer-readable storage medium storing computer-executable instructions, where the computer-executable instructions are used to execute an image retrieval method based on multi-task learning.
[0118] In this embodiment, the computer-readable storage medium may include volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules or other data. For example, computer-readable storage media include, but are not limited to, volatile memory, such as random access memory (RAM, DRAM, SRAM); and non-volatile memory, such as flash memory, various read-only memories (ROM, PROM, EPROM, EEPROM), magnetic and ferromagnetic / ferroelectric memories (MRAM, FeRAM); and magnetic and optical storage devices (hard disks, magnetic tapes, CDs, DVDs); or other media now known or developed in the future that can store computer-readable information / data for use by computer systems.
[0119] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present application.
[0120] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.
Claims
1. An image retrieval method based on multi-task learning, characterized in that: include: Get the input product image; Performing image preprocessing on the product image to obtain an original image and a cropped image of the product body; Importing the original image and the cropped image of the product body into a twin neural network; Performing model training according to the original image, the cropped image of the product body, and the product attributes of the product image to obtain a trained image model, wherein the product attributes include at least one of a product name, a product brand, a product applicable population, and a product applicable season; Deploy the trained image model to a server to output a final vector of the product image for image vector similarity retrieval; The step of performing model training according to the original image, the cropped image of the commodity body and the commodity attributes of the commodity image to obtain a trained image model comprises the following steps: Vector generation step: generating a full image vector and a cropped image vector according to the original image and the cropped image of the commodity body; Feature interaction step: performing feature interaction on the full image vector and the cropped image vector to obtain a feature vector; Loss calculation step: according to the feature vector and the commodity attributes of the commodity image, the loss is calculated by a loss function, and back propagation is performed according to the calculated loss to optimize and update the weights of each node in the twin neural network; Repeat the vector generation step, feature interaction step, and loss calculation step until the image model converges to obtain a trained image model; The loss calculation step specifically includes: Calculating a multi-label classification loss by at least one first loss function according to the commodity attributes of the commodity image to constrain the vector space of the feature vector; The metric distance loss is calculated according to the feature vector through one or more second loss functions to further expand the inter-class distance and reduce the intra-class distance.
2. The image retrieval method based on multi-task learning according to claim 1, characterized in that: The feature interaction step specifically includes: Feature interaction is performed on the full image vector and the cropped image vector, and addition, subtraction, dot product and variance operations are performed bit by bit, and then a normalization operation is performed to obtain a feature vector.
3. The image retrieval method based on multi-task learning according to claim 1, characterized in that: The step of deploying the trained image model to a server to output a final vector of the product image for image vector similarity retrieval further includes: extracting the final vector of the product image from the trained image model at regular intervals, and saving it to a vector similarity search engine for real-time image vector similarity retrieval; The vector similarity search engine includes a Mil vus vector search engine.
4. The image retrieval method based on multi-task learning according to claim 1, characterized in that: The image preprocessing of the product image to obtain the original image and the product main body cropped image specifically includes: Performing an enhancement operation on the entire image of the product to obtain an original image of the product image; Detecting the product image and cropping the product body, and then performing an enhancement operation to obtain a cropped image of the product body of the product image; The enhancement operation includes at least one of enlarging, reducing, cropping and rotating.
5. The image retrieval method based on multi-task learning according to claim 4, characterized in that: The detecting the commodity image and cropping the commodity body includes: Detecting a commodity main frame of the commodity image by a one-stage detection algorithm or a two-stage detection algorithm; The commodity body is cut out according to the commodity body frame.
6. An image retrieval device for implementing the image retrieval method based on multi-task learning as claimed in claim 1, characterized in that: The device comprises: A receiving module, used to obtain an input product image; An image preprocessing module, used to perform image preprocessing on the product image to obtain an original image and a cropped image of the product body; A control module, used for importing the original image and the cropped image of the commodity body into a twin neural network; A model training module, used to perform model training according to the original image, the cropped image of the commodity body and the commodity attributes of the commodity image to obtain a trained image model; A management module is used to deploy the trained image model to a server to output a final vector of the product image for image vector similarity retrieval.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the image retrieval method based on multi-task learning as described in any one of claims 1 to 5 is implemented.
8. A computer-readable storage medium storing computer-executable instructions, characterized in that: The computer executable instructions are used to execute the image retrieval method based on multi-task learning as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Image feature extraction model training method, image search method and computer device
CN110866140A
Training method, generating method, searching method and system of commodity attribute generating model
CN110968775A