Similarity detection model training method and device and image similarity detection method and device
By preprocessing and image enhancement processing of the original footwear image, an image similarity detection model is generated, which solves the problem of low detection accuracy due to the lack of data sets in the prior art, and achieves higher detection accuracy and faster training speed.
Patent Information
- Application Number
- CN202510256736.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-24
AI Technical Summary
Due to the lack of large-scale public data sets in the prior art, the accuracy of footwear appearance infringement detection models is low.
By preprocessing the original footwear image, standard image pairs are constructed and image-enhanced, inputting them into the preset deep neural network model, and parameter adjustments are performed to generate an image similarity detection model.
The accuracy of the similarity detection model is improved, the dependence on large-scale training data is reduced, and the generalization ability and training speed of the model are improved.
Smart Images

Figure CN120196956A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of image processing technology. Specifically, the present disclosure relates to a method for training a similarity detection model, an apparatus for training a similarity detection model, and a method for detecting image similarity. Background Art
[0002] During the training process of existing image detection models for detecting the similarity of footwear products, there are extremely high requirements for the quality and quantity of training data. However, since footwear appearance infringement is a relatively vertical and niche field, there is currently a lack of large-scale publicly available datasets to support it. Therefore, the accuracy of the obtained similarity detection model is relatively low.
[0003] It should be noted that the information disclosed in the above background art is only used to enhance the understanding of the background of the present disclosure. Therefore, it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0004] The purpose of the present disclosure is to provide a method for training a similarity detection model, an apparatus for training a similarity detection model, and a method for detecting image similarity, so as to at least overcome to a certain extent the problem of relatively low accuracy of the similarity detection model caused by the limitations and defects of related technologies.
[0005] According to one aspect of the present disclosure, a method for training a similarity detection model is provided, including:
[0006] Preprocessing the original footwear image to obtain a standard footwear image, and constructing a standard image pair according to the standard footwear image;
[0007] Performing image enhancement processing on the standard image pair to obtain an enhanced image pair, and inputting the enhanced image pair into a preset deep neural network model to obtain a first similarity detection result;
[0008] Constructing a target loss function according to the first similarity detection result and the first actual image label of the enhanced image pair;
[0009] Adjusting the parameters of the preset deep neural network model based on the target loss function to obtain an image similarity detection model.
[0010] In an exemplary embodiment of the present disclosure, preprocessing the original footwear image to obtain a standard footwear image includes: performing image cropping processing on the original footwear image based on a preset image cropping model to obtain an image cropping result; performing denoising processing on the image cropping result based on a preset image denoising model to obtain the standard footwear image; wherein the standard footwear image is a footwear image with a single style.
[0011] In an exemplary embodiment of the present disclosure, constructing a standard image pair according to the standard footwear image includes: classifying the standard footwear image based on a preset style classification model to obtain a style classification result, and determining a positive image pair according to the standard footwear images having the same style in the style classification result; determining a negative image pair according to the standard footwear images having different styles in the style classification result, and constructing the standard image pair according to the positive image pair and the negative image pair.
[0012] In an exemplary embodiment of the present disclosure, determining a positive image pair according to the standard footwear images having the same style in the style classification result includes: calculating a first image similarity between the standard footwear images having the same style in the style classification result, and determining whether the first image similarity is less than a first preset similarity threshold and greater than a second preset similarity threshold; when it is determined that the first image similarity is less than the first preset similarity threshold and greater than the second preset similarity threshold, splicing the standard footwear images having the same style to obtain a first positive image pair, and when it is determined that the first image similarity is less than the second preset similarity threshold, splicing the standard footwear images having the same style to obtain a second positive image pair; splicing the standard footwear images having the same style but different colors in the style classification result to obtain a third positive image pair, and splicing the standard footwear images having the same style but different carriers in the style classification result to obtain a fourth positive image pair; generating the positive image pair according to the first positive image pair and / or the second positive image pair and / or the third positive image pair and / or the fourth positive image pair.
[0013] In an exemplary embodiment of the present disclosure, determining a negative image pair according to the standard footwear images having different styles in the style classification result includes: randomly selecting standard footwear images having different styles from different style classification results, and splicing the standard footwear images having different styles to obtain a first negative image pair; calculating a second image similarity between the standard footwear images having different styles, and determining whether the second image similarity is greater than a third preset similarity threshold; when it is determined that the second image similarity is greater than the third preset similarity threshold, splicing the standard footwear images having different styles to obtain a second negative image pair, and generating the negative image pair according to the first negative image pair and / or the second negative image pair.
[0014] In an exemplary embodiment of the present disclosure, performing image enhancement processing on the standard image pair to obtain an enhanced image pair, including: transforming the original image color of the standard image pair, and / or adding image noise to the standard image pair, and / or performing geometric transformation on the standard image pair, and / or performing occlusion processing and / or blurring processing and / or affine transformation on the standard image pair to obtain the enhanced image pair.
[0015] In an exemplary embodiment of the present disclosure, the preset deep neural network model includes an angle discriminator, a first encoder, a first embedding mapping layer, and a multimodal large model; wherein, inputting the enhanced image pair into the preset deep neural network model to obtain a first similarity detection result, including: determining the image angle information of the enhanced image pair based on the angle discriminator, and generating first context information to be predicted according to the image angle information and the preset model hint information; encoding the first context information to be predicted based on the text encoder in the first encoder to obtain a first context flag sequence; encoding the enhanced image pair based on the image encoder in the first encoder to obtain a first image pair feature, and performing embedding mapping processing on the first context flag sequence and the first image pair feature based on the first embedding mapping layer to obtain a first overall context representation; performing similarity detection on the first overall context representation based on the multimodal large model to obtain the first similarity detection result.
[0016] In an exemplary embodiment of the present disclosure, the multimodal large model includes a second encoder, a multi-layer perceptron, a second embedding mapping layer, and a language model; wherein, performing similarity detection on the first overall context representation based on the multimodal large model to obtain the first similarity detection result, including: performing visual encoding processing on the image part in the first overall context representation based on the second encoder to obtain a visual encoding result, and performing classification processing on the visual encoding result based on the multi-layer perceptron to obtain an image classification result; performing embedding mapping processing on the text part in the first overall context representation based on the second embedding mapping layer to obtain a second context flag sequence; performing root mean square normalization processing on the image classification result and the second context flag sequence based on the language model to obtain the first similarity detection result.
[0017] In an exemplary embodiment of the present disclosure, parameter adjustment is performed on the preset deep neural network model based on the target loss function to obtain an image similarity detection model, including: determining low-rank matrix parameters of the multi-modal large model in the preset deep neural network model in the shoe image similarity comparison scenario based on the target loss function, and fine-tuning the multi-modal large model based on the low-rank matrix parameters to obtain a fine-tuned large model; adjusting the parameters in the angle discriminator and the parameters in the first embedding mapping layer in the preset deep neural network model based on the target loss function; generating the image similarity detection model according to the fine-tuned large model, the angle discriminator with adjusted parameters, the first embedding mapping layer with adjusted parameters, and the first encoder.
[0018] According to one aspect of the present disclosure, there is provided a method for detecting image similarity, including:
[0019] Obtaining a first shoe product image to be detected and a second shoe product image to be detected, and generating a pair of images to be detected according to the first shoe product image to be detected and the second shoe product image to be detected;
[0020] Inputting the pair of images to be detected into a preset image similarity detection model to obtain a shoe product similarity detection result; wherein, the image similarity detection model is obtained according to the training method of the similarity detection model described in any one of the above.
[0021] Determining whether there is an appearance infringement behavior of shoe products between the first shoe product image to be detected and the second shoe product image to be detected according to the shoe product similarity detection result.
[0022] In an exemplary embodiment of the present disclosure, the shoe product similarity detection result includes at least one of a sole similarity detection result, a tongue similarity detection result, a vamp similarity detection result, a heel similarity detection result, and a strap similarity detection result.
[0023] In an exemplary embodiment of the present disclosure, determining whether there is an infringement on the appearance of a footwear product between the first footwear product image to be detected and the second footwear product image according to the footwear product similarity detection result includes: If the sole similarity detection result, tongue similarity detection result, upper similarity detection result, heel similarity detection result, and accessory similarity detection result in the footwear product similarity detection result are all similar, it is determined that there is an infringement on the appearance of the footwear product between the first footwear product image to be detected and the second footwear product image; If any one of the sole similarity detection result, tongue similarity detection result, upper similarity detection result, heel similarity detection result, and accessory similarity detection result in the footwear product similarity detection result is dissimilar, it is determined that there is no infringement on the appearance of the footwear product between the first footwear product image to be detected and the second footwear product image.
[0024] According to one aspect of the present disclosure, there is provided a training device for a similarity detection model, including:
[0025] An image pair construction module, configured to preprocess an original footwear image to obtain a standard footwear image, and construct a standard image pair according to the standard footwear image;
[0026] An image enhancement module, configured to perform image enhancement processing on the standard image pair to obtain an enhanced image pair, and input the enhanced image pair into a preset deep neural network model to obtain a first similarity detection result;
[0027] A loss function construction module, configured to construct a target loss function according to the first similarity detection result and the first actual image label of the enhanced image pair;
[0028] A parameter adjustment module, configured to adjust the parameters of the preset deep neural network model based on the target loss function to obtain an image similarity detection model.
[0029] According to one aspect of the present disclosure, there is provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the training method of the similarity detection model described in any one of the above exemplary embodiments, and the detection method of the image similarity described in any one of the above exemplary embodiments.
[0030] According to one aspect of the present disclosure, there is provided an electronic device, including:
[0031] A processor; and
[0032] A memory, configured to store executable instructions of the processor;
[0033] Wherein, the processor is configured to execute the training method of the similarity detection model described in any one of the above exemplary embodiments and the detection method of image similarity described in any one of the above exemplary embodiments by executing the executable instructions.
[0034] A training method for a similarity detection model provided by an embodiment of the present disclosure. On the one hand, by preprocessing the original shoe image to obtain a standard shoe image and constructing a standard image pair according to the standard shoe image; then performing image enhancement processing on the standard image pair to obtain an enhanced image pair, and inputting the enhanced image pair into a preset deep neural network model to obtain a first similarity detection result; further constructing a target loss function according to the first similarity detection result and the first actual image label of the enhanced image pair; finally, adjusting the parameters of the preset deep neural network model based on the target loss function to obtain an image similarity detection model; since during the model training process, the corresponding number of enhanced image pairs can be constructed according to actual needs, thus solving the problem in the prior art that the accuracy of the similarity detection model obtained by training is relatively low due to the lack of corresponding training data, and improving the accuracy of the similarity detection model; on the other hand, since the corresponding enhanced image pairs can be constructed by itself for model training, the generalization ability of the model is improved, the dependence on large-scale training data during the model training process is reduced, and the training speed of the model is increased.
[0035] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present disclosure and used together with the specification to explain the principles of the present disclosure. Obviously, the following described drawings are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0037] Figure 1 Schematically showing a flowchart of a training method for a similarity detection model according to an exemplary embodiment of the present disclosure.
[0038] Figure 2 Schematically showing a structural example diagram of a preset deep neural network model according to an exemplary embodiment of the present disclosure.
[0039] Figure 3 Schematically showing a structural example diagram of a multimodal large model in a preset deep neural network model according to an exemplary embodiment of the present disclosure.
[0040] Figure 4A scenario diagram schematically showing a type of footwear data according to an exemplary embodiment of the present disclosure.
[0041] Figure 5 A scenario diagram schematically showing footwear data of the same style and different styles according to an exemplary embodiment of the present disclosure.
[0042] Figure 6 A scenario example diagram of a spliced sample pair schematically shown according to an exemplary embodiment of the present disclosure.
[0043] Figure 7 A flowchart of a method for detecting image similarity schematically shown according to an exemplary embodiment of the present disclosure.
[0044] Figure 8 A structural example diagram of a training device for a similarity detection model schematically shown according to an exemplary embodiment of the present disclosure.
[0045] Figure 9 A structural example diagram of a device for detecting image similarity schematically shown according to an exemplary embodiment of the present disclosure.
[0046] Figure 10 An electronic device for implementing a training method for a similarity detection model and a method for detecting image similarity schematically shown according to an exemplary embodiment of the present disclosure. Detailed implementation manners
[0047] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, the exemplary embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be more complete and comprehensive, and will fully convey the concept of the exemplary embodiments to those skilled in the art. The described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art will realize that the technical solutions of the present disclosure can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, well-known technical solutions are not shown or described in detail to avoid obscuring the various aspects of the present disclosure.
[0048] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0049] Multimodal Large Models are deep learning models that combine multiple data types (such as text, images, audio, and video, etc.), aiming to improve the performance and understanding ability of tasks by integrating information from different modalities; in the process of actual application, Multimodal Large Models can not only process input data of a single type, but also perform cross-modal reasoning and generation, so as to achieve the purpose of showing stronger capabilities in complex tasks.
[0050] The identification of appearance infringement of footwear products is a key technology in the field of product compliance; traditional infringement identification methods usually rely on image retrieval technology, and the specific steps include: First, establish an image library of the footwear appearance to be guarded against; then, train a special footwear appearance retrieval model; finally, use this footwear appearance retrieval model to perform appearance matching between the picture to be detected and all the pictures in the image library to find out potentially similar footwear pictures. However, this method has the following defects: On the one hand, it has extremely high requirements for the retrieval model, and the obtained infringement identification results almost entirely depend on the performance of the model, thus resulting in a relatively low accuracy of the obtained identification results; on the other hand, such retrieval models often need to be trained from scratch and have extremely high requirements for the quality and quantity of training data; however, the field of footwear appearance infringement is a relatively vertical and niche field, and currently lacks large-scale public datasets to support it, so the accuracy of the obtained model is relatively low; moreover, the cost of collecting sufficient high-quality data by oneself is extremely high, which directly limits the improvement of footwear infringement retrieval ability.
[0051] Under this premise, the exemplary embodiments of the present disclosure propose a method for identifying appearance infringement based on a multimodal large model; in the actual application process, since the multimodal large model can be pre-trained by using a large amount of unlabeled or multi-label data, it can learn rich general knowledge and patterns; therefore, these pre-trained models can usually exhibit better generalization ability and higher performance in downstream tasks, especially in the few-shot learning scenario; further, this advantage greatly reduces the dependence on large-scale training data when the multimodal large model technology is fine-tuned in vertical fields such as footwear; that is, the accuracy of the model can be improved on the basis of reducing the amount of data in the required dataset; at the same time, by adjusting some parameters of the model or adding specific task layers, the new task can be quickly adapted while maintaining the original capabilities, so as to achieve the purpose of identifying the appearance infringement of footwear products.
[0052] In this exemplary embodiment, a training method for a similarity detection model is first provided, and this method can run on a terminal device, a server, a server cluster, a cloud server, etc.; of course, those skilled in the art can also run the method of the present disclosure on other platforms according to requirements, and no special limitation is made in this exemplary embodiment. Specifically, referring to Figure 1 as shown, the training method of the similarity detection model may include the following steps:
[0053] Step S110. Preprocess the original footwear image to obtain a standard footwear image, and construct a standard image pair according to the standard footwear image;
[0054] Step S120. Perform image enhancement processing on the standard image pair to obtain an enhanced image pair, and input the enhanced image pair into a preset deep neural network model to obtain a first similarity detection result;
[0055] Step S130. Construct a target loss function according to the first similarity detection result and the first actual image label of the enhanced image pair;
[0056] Step S140. Adjust the parameters of the preset deep neural network model based on the target loss function to obtain an image similarity detection model.
[0057] In the training method of the above similarity detection model, on the one hand, by preprocessing the original footwear image, a standard footwear image is obtained, and a standard image pair is constructed based on the standard footwear image; then the standard image pair is subjected to image enhancement processing to obtain an enhanced image pair, and the enhanced image pair is input into a preset deep neural network model to obtain a first similarity detection result; furthermore, a target loss function is constructed based on the first similarity detection result and the first actual image label of the enhanced image pair; finally, the parameters of the preset deep neural network model are adjusted based on the target loss function to obtain an image similarity detection model; since during the model training process, the corresponding number of enhanced image pairs can be constructed according to actual needs, the problem of low accuracy of the similarity detection model obtained by training in the prior art due to the lack of corresponding training data is solved, and the accuracy of the similarity detection model is improved; on the other hand, since the corresponding enhanced image pairs can be constructed by oneself for model training, the generalization ability of the model is improved, the dependence on large-scale training data during the model training process is reduced, and the training speed of the model is increased.
[0058] Hereinafter, the training method of the similarity detection model recorded in the exemplary embodiments of the present disclosure will be further explained and described with reference to the accompanying drawings.
[0059] First, the preset deep neural network model involved in the exemplary embodiments of the present disclosure will be explained and described. Specifically, referring to Figure 2 As shown, the preset deep neural network model may include a first input layer 210, an angle discriminator 220, a first encoder 230, a first embedding mapping layer 240, a multimodal large model 250, and a first output layer 260. Among them, the functions that each model layer included in the preset deep neural network model needs to perform in the specific similarity comparison process will be described in detail later, and will not be further elaborated here; at the same time, the first encoder described here may include a text encoder and an image encoder.
[0060] In one exemplary embodiment, referring to Figure 3 As shown, the multimodal large model described here may include a second input layer 301, a second encoder 302, a multi-layer perceptron 303, a second embedding mapping layer 304, a language model 305, and a second output layer 306. Among them, the functions that each model layer included in the multimodal large model needs to perform in the specific similarity comparison process will be described in detail later, and will not be further elaborated here.
[0061] In an exemplary embodiment, the language model described herein is LlaMA; wherein, the language model uses the decoder part of Transformer and adopts a decoder-only structure; and this structure is now adopted by most generative language models. In terms of the specific model structure, compared with the Transformer model, the main changes in LLaMA2 are that the LayerNorm in it is replaced by RMSNorm, the Multi-Head Attention is replaced by Grouped Query Attention (GQA, which is Multi-Query Attention MQA in LLaMA), and the Positional Encoding is replaced by Rotary Position Embedding (RoPE). In the actual application process, replacing LayerNorm with RMSNorm can enhance the stability during training; at the same time, using Rotary Position Embedding to introduce positional information can realize introducing positional encoding into Query and Key through geometric operations of complex numbers, enabling the model to understand the order of words in the sequence; replacing Multi-Head Attention with Grouped Query Attention can group queries and share KV pairs within the group, thereby reducing the computational amount and improving the efficiency.
[0062] Next, in conjunction with Figure 2 and Figure 3 to Figure 1 the training method of the similarity detection model shown in
[0063] In step S110, the original footwear image is preprocessed to obtain a standard footwear image, and a standard image pair is constructed based on the standard footwear image.
[0064] In this exemplary embodiment, first, the original footwear image is preprocessed to obtain a standard footwear image; specifically, it can be achieved in the following way: the original footwear image is subjected to image cropping processing based on a preset image cropping model to obtain an image cropping result; the image cropping result is subjected to denoising processing based on a preset image denoising model to obtain the standard footwear image; wherein, the standard footwear image is a footwear image with a single style.
[0065] Secondly, construct a standard image pair based on the standard footwear images; specifically, it can be achieved in the following manner: classify the standard footwear images based on a preset style classification model to obtain a style classification result, and determine a positive image pair according to the standard footwear images with the same style in the style classification result; determine a negative image pair according to the standard footwear images with different styles in the style classification result, and construct the standard image pair based on the positive image pair and the negative image pair.
[0066] In one exemplary embodiment, determining a positive image pair according to the standard footwear images with the same style in the style classification result can be achieved in the following manner: calculate the first image similarity between the standard footwear images with the same style in the style classification result, and determine whether the first image similarity is less than a first preset similarity threshold and greater than a second preset similarity threshold; when it is determined that the first image similarity is less than the first preset similarity threshold and greater than the second preset similarity threshold, splice the standard footwear images with the same style to obtain a first positive image pair, and when it is determined that the first image similarity is less than the second preset similarity threshold, splice the standard footwear images with the same style to obtain a second positive image pair; splice the standard footwear images with the same style but different colors in the style classification result to obtain a third positive image pair, and splice the standard footwear images with the same style but different carriers in the style classification result to obtain a fourth positive image pair; generate the positive image pair according to the first positive image pair and / or the second positive image pair and / or the third positive image pair and / or the fourth positive image pair.
[0067] In one exemplary embodiment, determining a negative image pair according to the standard footwear images with different styles in the style classification result can be achieved in the following manner: randomly select standard footwear images with different styles from different style classification results, and splice the standard footwear images with different styles to obtain a first negative image pair; calculate the second image similarity between the standard footwear images with different styles, and determine whether the second image similarity is greater than a third preset similarity threshold; when it is determined that the second image similarity is greater than the third preset similarity threshold, splice the standard footwear images with different styles to obtain a second negative image pair, and generate the negative image pair according to the first negative image pair and / or the second negative image pair.
[0068] Next, the specific processing process of the standard footwear images and the specific construction process of the standard image pair will be further explained and described. Specifically, in the actual application process, due to the small amount of publicly available data in the current footwear scenario, and the data publicly available for the footwear detection scenario includes pictures of shoes and other backgrounds; for example, it can be referred toFigure 4 as shown at 401 in; however, since the training data for footwear recognition needs to be sorted into close-up images of single-style shoes first, for example, reference can be made to Figure 4 as shown at 402 in; therefore, in order to obtain close-up images of single-style shoes, it is necessary to use existing footwear detection models (such as image detection models, image denoising models, etc.) for matte painting conversion; in addition, in the publicly available data of footwear, since the styles of shoes in different pictures are often different, and the training data requirements for footwear infringement recognition are, one is that the styles are as many as possible, and the other is that there are as many shoe pictures of different angles of the same style as possible, but the publicly available data obviously cannot meet the requirement of having as many shoe pictures of different angles of the same style as possible; based on this, in order to match the footwear infringement recognition framework based on the multi-modal large model, the format of the training data involved in this exemplary embodiment is the splicing of two close-up images of shoes, one on the left and one on the right; further, all the training data can be divided into two parts: positive samples and negative samples, where the close-up images of the left and right shoes in the positive samples are of the same style, specifically, reference can be made to Figure 5 as shown at 501 in; the close-up images of the left and right shoes in the negative samples are of different styles, specifically, reference can be made to Figure 5 as shown at 502 in.
[0069] Further, the same style recorded in this publicly disclosed exemplary embodiment means that the appearance designs of the shoes are the same, regardless of whether the colors and materials are the same; on this premise, for the positive sample data, in order to enhance the accuracy of the model, it can be divided into the following four categories: one is the general positive sample; another is the positive sample with a large difference in lighting angles; still another is the positive sample with different colors; and there is also the positive sample with or without a human foot (that is, the positive sample with or without a carrier); at the same time, through experiments, it can be known that the configuration ratio of the above four types of data is set to 6:1:2:1. On this premise, the specific construction process of the positive sample pair can be realized in the following way: First, collect a batch of footwear pictures with style labels, and use the existing footwear detection models to crop the footwear pictures into close-up images of single shoes; then, preliminarily screen these close-up images through the basic model to remove noise (other-style shoes that may be contained in the pictures), and then store them according to style classification. That is to say, first, perform image cropping processing on the original footwear image based on a preset image cropping model to obtain an image cropping result; perform denoising processing on the image cropping result based on a preset image denoising model to obtain the standard footwear image; where the image cropping models recorded here may include, but are not limited to, SegNet, DeepLab, Mask R-CNN, U-Net, Gated SCNN, etc.; the image denoising models recorded here may include, but are not limited to, Gaussian filtering models, median filtering models, bilateral filtering models, and deep neural network models, etc., and this example does not make special restrictions on this.
[0070] In an exemplary embodiment, during the process of constructing the first positive image pair (i.e., the general positive sample pair) in the positive image pairs, theoretically, any two close-up images in the same style can be randomly selected for splicing; however, due to data limitations, if the random selection method is used for splicing, it is very easy to splice two almost identical close-up images; among them, the obtained splicing result can be referred to Figure 6 as shown in 601, 602, and 603 in
[0071] ; the splicing result obtained based on this method has very limited training gain for the subsequent model and also consumes a lot of training resources; therefore, during the process of constructing the first positive image pair, it is necessary to use the basic model for screening and comparison, and only the close-up images with a similarity score below a certain threshold (<0.85) in the same style are spliced; that is, when constructing the first positive image pair, first, it is necessary to calculate the first image similarity between the standard shoe images with the same style in the style classification result, and determine whether the first image similarity is less than the first preset similarity threshold (for example, it can be 0.85, and of course, other values can also be taken, and this example does not make special restrictions on this) and greater than the second preset similarity threshold (for example, it can be 0.5, and of course, other values can also be taken, and this example does not make special restrictions on this); when it is determined that the first image similarity is less than the first preset similarity threshold and greater than the second preset similarity threshold, the standard shoe images with the same style are spliced to obtain the first positive image pair; at the same time, during the process of determining the first image similarity, it can be determined based on the cosine value between the two images, or directly based on the similarity calculation model, and this example does not make special restrictions on this; the similarity calculation model described here can be a large model based on Bert or a deep neural network model, and this example does not make special restrictions on this.
[0072] In an exemplary embodiment, during the process of constructing the third positive image pair (i.e., the positive sample pair with different colors) in the positive image pairs, the color classification of each close-up image can be obtained from the same-style close-up images through a large multimodal model, and then matching is performed to obtain it; that is, standard footwear images with the same style but different colors in the style classification results are spliced to obtain the third positive image pair; among them, the large multimodal model (LargeMultimodal Models, LMM) described here has a similar model structure to the large multimodal model described above, and no further elaboration will be made here.
[0073] In an exemplary embodiment, during the process of constructing the fourth positive image pair (i.e., the positive sample pair with or without a human foot) in the positive image pairs, the classification of whether there is a human foot in each close-up image can be obtained from the same-style close-up images through a large multimodal model, and then a close-up image with a human foot and a close-up image without a human foot are selected for matching to obtain it; that is, standard footwear images with the same style but different carriers in the style classification results can be spliced to obtain the fourth positive image pair.
[0074] In an exemplary embodiment, during the process of constructing the negative image pairs, to enhance the model effect, it can be divided into the first negative image pair (i.e., the general negative sample pair) and the second negative image pair (i.e., the difficult negative sample); during the actual application process, through a large number of experiments, it can be known that the configuration ratio of the above two types of data is set to 5:1. Further, during the process of constructing the first negative image pair, two randomly selected close-up images with different styles can be spliced to obtain the first negative image pair; during the process of constructing the second negative image pair, all close-up images with different styles need to be paired in pairs, and then close-up images with a similarity above a certain level (>0.6) are selected for splicing; that is, during the process of constructing the second negative image pair, first, the second image similarity between standard footwear images with different styles needs to be calculated, and it is determined whether the second image similarity is greater than the third preset similarity threshold (for example, it can be 0.6, and of course other values can also be taken, and this example does not make special restrictions on this); when it is determined that the second image similarity is greater than the third preset similarity threshold, the standard footwear images with different styles are spliced to obtain the second negative image pair.
[0075] In an exemplary embodiment, for better results, the exemplary embodiment of the present disclosure makes a targeted design for the image stitching method. Specifically, in the actual application process, since the close-up images to be stitched are directly cropped from the original image, there is a problem of different sizes. If directly resized to a fixed resolution for stitching, it will cause distortion in the shape of the shoes (for example, the original image is 267*650 and is directly resized to 510*510), which invisibly increases the difficulty for the model, especially when the size difference between the two close-up images is relatively large. Under this premise, in the process of image stitching in the exemplary embodiment of the present disclosure, first, all the close-up images to be stitched are padded into squares according to the longest side, then the obtained squares are resized to a fixed size, and finally the two close-up images of the fixed size are stitched together.
[0076] In step S120, perform image enhancement processing on the standard image pair to obtain an enhanced image pair, and input the enhanced image pair into a preset deep neural network model to obtain a first similarity detection result.
[0077] In the exemplary embodiment of the present disclosure, first, perform image enhancement processing on the standard image pair to obtain an enhanced image pair. Specifically, it can be achieved in the following ways: transform the original image color of the standard image pair, and / or add image noise to the standard image pair, and / or perform geometric transformation on the standard image pair, and / or perform occlusion processing and / or blur processing and / or affine transformation on the standard image pair to obtain the enhanced image pair. That is, in the actual application process, for better results, the exemplary embodiment of the present disclosure specifically designs a data enhancement module, and the data enhancement methods can include but are not limited to: changing the color of the standard image pair (for example, adjusting parameters such as brightness, contrast, saturation, and hue to simulate images under different lighting conditions), adding noise to the standard image pair (for example, adding Gaussian noise and salt-and-pepper noise to the image to improve the robustness of the model to noise), performing geometric transformation on the standard image pair (for example, including but not limited to operations such as rotation, flipping, cropping, and scaling to increase the diversity of the image), occluding some regions in the standard image pair (for example, randomly occluding some image regions), performing blur processing on the standard image pair (for example, applying Gaussian blur or other types of blur filters to simulate images taken with low quality or from a long distance), performing affine transformation on the standard image pair (for example, including but not limited to translation, rotation, scaling, and translation combination on the standard image pair to simulate different perspective changes), and so on. And, in the specific model training process, all training data can randomly select one or more of the above-mentioned data enhancement methods.
[0078] Secondly, input the enhanced image pair into a preset deep neural network model to obtain a first similarity detection result. Specifically, it can be implemented in the following way: determine the image angle information of the enhanced image pair based on the angle discriminator, and generate first context information to be predicted according to the image angle information and preset model prompt information; perform encoding processing on the first context information to be predicted based on the text encoder in the first encoder to obtain a first context token sequence; perform encoding processing on the enhanced image pair based on the image encoder in the first encoder to obtain a first image pair feature, and perform embedding mapping processing on the first context token sequence and the first image pair feature based on the first embedding mapping layer to obtain a first overall context representation; perform similarity detection on the first overall context representation based on the multi-modal large model to obtain the first similarity detection result.
[0079] In an example embodiment, performing similarity detection on the first overall context representation based on the multi-modal large model to obtain the first similarity detection result can be implemented in the following way: perform visual encoding processing on the image part in the first overall context representation based on the second encoder to obtain a visual encoding result, and perform classification processing on the visual encoding result based on the multi-layer perceptron to obtain an image classification result; perform embedding mapping processing on the text part in the first overall context representation based on the second embedding mapping layer to obtain a second context token sequence; perform root mean square normalization processing on the image classification result and the second context token sequence based on the language model to obtain the first similarity detection result.
[0080] Next, the specific determination process of the first similarity detection result will be further explained and described. Specifically, first, design a corresponding model prompt information Prompt. Among them, after detailed design and multiple deliberations, the model prompt information Prompt used in the example embodiment of the present disclosure is as follows:
[0081] Prompt = "Please determine whether the styledesigns of the left andright shoes are consistent.Pay attention to the details of the sole,upper,laces,tongue,and heel.Output the result in JSON format,where'yes'representsconsistency and'no'represents inconsistency.If a specific detail is notvisible in the image,output'unsure'.For example:{'sole':'yes','upper':'no','laces':'unsure','tongue':'yes','heel':'no'}. (Model prompt information = \"Please determine whether the styledesigns of the left andright shoes are consistent.Pay attention to the details of the sole,upper,laces,tongue,and heel.Output the result in JSON format,where'yes'representsconsistency and'no'represents inconsistency.If a specific detail is notvisible in the image,output'unsure'.For example:{'sole':'yes','upper':'no','laces':'unsure','tongue':'yes','heel':'no'})
[0082] Note (Note):
[0083] 1) Focus only on the shape and design when making the judgment,anddisregard color and material.Failure to do so will result in a penalty. (When making the judgment, only focus on the shape and design, and ignore the color and material. Otherwise, a penalty will be imposed)
[0084] 2) If the left or right image is not a shoe (e.g., socks, pants,isolatedshoe soles), output 'no' for all details. (If the left or right image is not a shoe (such as socks, pants, isolated shoe soles), output 'no' for all details)
[0085] 3) Thick and thin heels for women's shoes are considered inconsistent heel designs.
[0086] 4) The left and right shoes to be compared are usually captured from different angles (i.e., each shoe may only reveal certain details, and the visible details may differ between the two shoes). You only need to consider the details that are visible in both shoes. Note that you should only make a judgment of consistency (yes or no) if the details in both shoes are clearly visible and sufficient. If any detail is not clearly visible in either shoe, output 'unsure' for that detail. Do not attempt to make judgments on details that are unclear, as doing so will result in a penalty. (For example, if one shoe shows a clear view of the heel, while the other shoe's heel is only partially visible, the judgment for the heel should be 'unsure'.)
[0087] Furthermore, the specific Chinese corresponding to the above content is as follows: The left and right shoes to be compared are usually photographed from different angles (i.e., each shoe may only show certain details, and the visible details of the two shoes may be different). You only need to consider the details visible in the pair of shoes. Note that you should only make a judgment (yes or no) on consistency when the details of both pairs of shoes are clearly visible and sufficient. If any detail is not clearly visible in either shoe, output "uncertain" for that detail. Do not attempt to make a judgment on unclear details, as doing so will result in penalties. For example, if the heel of one shoe can be clearly seen while the heel of the other shoe can only be partially seen, the judgment on the heel should be "uncertain".
[0088] Secondly, after obtaining the model prompt information Prompt, the model prompt information Prompt and the standard image pair can be input into a preset deep neural network model to obtain a first similarity detection result; among them, the output of the model is not to directly judge whether the two spliced close-up pictures are of the same style in a conventional way, but to match and compare the details of the two close-up pictures; in the actual application process, if the output is to match and compare the details of the two close-up pictures, the effect is better than the direct output result of the model; for example, when the detail designs of the'sole' and 'tongue' parts of the shoes in the two close-up pictures are the same, while the detail designs of the 'upper' and 'heel' parts are different, and the 'laces' part cannot be clearly seen in one or more of the close-up pictures, the model will output: {'sole': 'yes', 'upper': 'no', 'laces': 'unsure', 'tongue': 'yes', 'heel': 'no'}.
[0089] It should also be supplemented and explained here that the angle discriminator described in the exemplary embodiments of the present disclosure is specifically designed for the footwear scenario, which is also an important feature distinguishing it from other multi-modal large model training frameworks; in the exemplary embodiments of the present disclosure, the main function of the angle discriminator is to judge the input pictures; that is, the angle discriminator can be used to judge whether the shoe angles of the left and right two close-up pictures are very different; when the angle difference is large, the probability of hallucinations appearing in the output of the model is relatively high. For example, because the angle difference is large, one shoe can see the'sole' part while the other shoe cannot see the'sole' part, the model should output 'unsure'; however, during actual testing, it is found that the model is very likely to output hallucinations for the invisible parts and get the results of 'yes' or 'no'; although this phenomenon can be optimized by adding certain hints in the prompt; however, the obtained effect is relatively limited. Therefore, the exemplary embodiments of the present disclosure add an angle discriminator to the model, combine the result of the angle discriminator with certain descriptions, and input them into the multi-modal large model for auxiliary guidance; at the same time, using this method can greatly reduce the hallucinations that occur in the model, so as to achieve the purpose of improving the accuracy of the obtained first similarity detection result.
[0090] In step S130, according to the first similarity detection result and the first actual image label of the enhanced image pair, a target loss function is constructed.
[0091] Specifically, the target loss function described here can be a cross-entropy loss function or a Sigmoid function, and this example does not make special restrictions on this. Further, taking the cross-entropy loss function as an example, the specific target loss function can be shown as the following formula (1):
[0092]
[0093] where L(y, f(x; Θ)) is the target loss function, y i is the first actual image label of the i-th enhanced image pair, and f(x i ; Θ) is the first similarity detection result of the i-th enhanced image pair; N is the total number of samples of the enhanced image pair.
[0094] In step S140, based on the target loss function, the parameters of the preset deep neural network model are adjusted to obtain an image similarity detection model.
[0095] Specifically, the parameter adjustment of the preset deep neural network model based on the target loss function to obtain the image similarity detection model can be achieved in the following way: determining the low-rank matrix parameters of the multi-modal large model in the preset deep neural network model in the shoe image similarity comparison scenario based on the target loss function, and fine-tuning the multi-modal large model based on the low-rank matrix parameters to obtain the fine-tuned large model; adjusting the parameters in the angle discriminator and the parameters in the first embedding mapping layer in the preset deep neural network model based on the target loss function; generating the image similarity detection model according to the fine-tuned large model, the angle discriminator with adjusted parameters, the first embedding mapping layer with adjusted parameters, and the first encoder. That is, in the actual application process, the angle discriminator is trained from scratch with all parameters, the multi-modal large model can be fine-tuned using Lora, while the text encoder and image encoder can freeze their parameters and do not participate in the training.
[0096] So far, the training method of the similarity detection model described in the exemplary embodiments of the present disclosure has been fully implemented. Further, after obtaining the image similarity detection model, the infringement behavior of shoe products can be determined based on this image similarity detection model. Specifically, as shown in Figure 7 The specific infringement judgment process can include the following steps:
[0097] Step S710, obtaining the first shoe product image to be detected and the second shoe product image to be detected, and generating a pair of images to be detected according to the first shoe product image to be detected and the second shoe product image to be detected;
[0098] Step S720, inputting the pair of images to be detected into the preset image similarity detection model to obtain the shoe product similarity detection result; wherein, the image similarity detection model is obtained according to the training method of the similarity detection model described in any one of the above.
[0099] Step S730, determining whether there is an appearance infringement behavior of shoe products between the first shoe product image to be detected and the second shoe product image to be detected according to the shoe product similarity detection result. The shoe product similarity detection result includes the sole similarity detection result, the tongue similarity detection result, the upper similarity detection result, the heel similarity detection result, and the carrying similarity detection result, etc.
[0100] In an exemplary embodiment, based on the detection result of the similarity of the footwear products, to determine whether there is an infringement of the appearance of the footwear products between the first footwear product image to be detected and the second footwear product image to be detected can be achieved in the following manner: If the detection results of the sole similarity, tongue similarity, upper similarity, heel similarity, and lacing similarity in the detection result of the similarity of the footwear products are all similar, it is determined that there is an infringement of the appearance of the footwear products between the first footwear product image to be detected and the second footwear product image; If any one of the detection results of the sole similarity, tongue similarity, upper similarity, heel similarity, and lacing similarity in the detection result of the similarity of the footwear products is dissimilar, it is determined that there is no infringement of the appearance of the footwear products between the first footwear product image to be detected and the second footwear product image.
[0101] Hereinafter, the specific infringement determination process will be further explained and described. Specifically, in the actual application process, the general infringement recognition method directly uses the result of the large model as the final result; However, the image similarity detection model described in the exemplary embodiment of the present disclosure changes the output mode of the model to improve the effect; For example, the output result is whether the appearances of the first footwear product image to be detected and the second footwear product image are consistent in the five parts of'sole', 'upper', 'laces', 'tongue', and 'heel'; Then, based on the output result of the model, it is determined whether the overall appearances are consistent finally. Among them, the specific discrimination logic is: Only when the output results of the five parts do not contain 'no', that is, no clear part has a different appearance, it is determined that the shoes in the two close-up pictures on the left and right have the same appearance, that is, there is a risk of infringement. For example, when the input of the comprehensive evaluation module is {'sole': 'yes', 'upper': 'no', 'laces': 'unsure', 'tongue': 'yes', 'heel': 'no'}, the output is non-infringement; Another example is that when the input of the comprehensive evaluation module is {'sole': 'yes', 'upper': 'unsure', 'laces': 'unsure', 'tongue': 'yes', 'heel': 'unsure'}, the output is infringement.
[0102] Hereinafter, the infringement determination process will be explained and described in combination with specific embodiments. Specifically, in the actual application process, for any footwear picture that needs to be detected for appearance infringement, all the close-up pictures of the shoes in the picture are cropped using the existing footwear detection model; Then, the cropped close-up Figure 11. Perform infringement detection. For example, first obtain the close-up image a for infringement detection, and then splice the close-up image a pairwise with all the close-up images of shoes in the infringement database to obtain the spliced images to be detected. Suppose there are m images in the existing infringement database, such as {b1, b2, b3,..., bm} (the infringement database is generally collected and established by each enterprise itself). Then, the close-up image a can form m spliced images {ab1, ab2,..., abm}. Further, input each spliced image into the shoe infringement recognition model to obtain the corresponding preliminary result t. For example, the preliminary result of ab1 is t1. Furthermore, input the obtained preliminary result t into the comprehensive evaluation module to obtain the final result out. Among them, each spliced image has a discrimination result, and the same close-up image may be infringed by the appearance of multiple shoes in the database. For example, if the discrimination results of the spliced images ab1 and ab10 are both infringement, it means that the close-up image a has an infringement risk with b1 and b10 in the database. Finally, repeat the above process to detect all the close-up images and use the results as the infringement results of the entire shoe image.
[0103] So far, the specific judgment process has been fully implemented. Based on the foregoing content, it can be known that the exemplary embodiment of the present disclosure uses a multi-modal large model to detect shoe appearance infringement, greatly reducing the requirements for training data. At the same time, the proposed shoe data construction scheme can construct a high-quality data set in a low-cost manner. In the actual application process, by carefully designing the construction and proportion of positive and negative samples, such a small-scale data set can train a superior model. The targeted data enhancement method can effectively improve the model effect. And the shoe infringement recognition framework based on the multi-modal large model designed in the exemplary embodiment of the present disclosure greatly improves the detection effect through the auxiliary recognition of the angle discriminator and the design of effective prompts. Further, the exemplary embodiment of the present disclosure fine-tunes the large model in the way of Lora, and only the angle discriminator with a small number of parameters needs to be trained from scratch, greatly reducing the training difficulty and training cost. Using the comprehensive evaluation module to obtain the final result instead of directly outputting the result by the multi-modal large model enables the model to increase the thinking chain thinking mode, obtain more detailed and reliable preliminary results, and then obtain more accurate final results through a certain discrimination method. In summary, the exemplary embodiment of the present disclosure has multiple advantages in different dimensions, such as fast training speed, low training cost, high robustness, and high model accuracy.
[0104] The following is the device embodiment of the present disclosure, which can be used to execute the method embodiment of the present disclosure. For the details not disclosed in the device embodiment of the present disclosure, please refer to the method embodiment of the present disclosure.
[0105] The present disclosure also provides a training device for a similarity detection model. Specifically, refer to Figure 8As shown in the figure, the training device of the similarity detection model may include an image pair construction module 810, an image enhancement module 820, a loss function construction module 830, and a parameter adjustment module 840. Among them:
[0106] The image pair construction module 810 can be used to preprocess the original footwear image to obtain a standard footwear image, and construct a standard image pair according to the standard footwear image;
[0107] The image enhancement module 820 can be used to perform image enhancement processing on the standard image pair to obtain an enhanced image pair, and input the enhanced image pair into a preset deep neural network model to obtain a first similarity detection result;
[0108] The loss function construction module 830 can be used to construct a target loss function according to the first similarity detection result and the first actual image label of the enhanced image pair;
[0109] The parameter adjustment module 840 can be used to adjust the parameters of the preset deep neural network model based on the target loss function to obtain an image similarity detection model.
[0110] In an exemplary embodiment of the present disclosure, preprocessing the original footwear image to obtain a standard footwear image includes: performing image cropping processing on the original footwear image based on a preset image cropping model to obtain an image cropping result; performing denoising processing on the image cropping result based on a preset image denoising model to obtain the standard footwear image; wherein, the standard footwear image is a footwear image with a single style.
[0111] In an exemplary embodiment of the present disclosure, constructing a standard image pair according to the standard footwear image includes: classifying the standard footwear image based on a preset style classification model to obtain a style classification result, and determining a positive image pair according to the standard footwear images with the same style in the style classification result; determining a negative image pair according to the standard footwear images with different styles in the style classification result, and constructing the standard image pair according to the positive image pair and the negative image pair.
[0112] In an exemplary embodiment of the present disclosure, determining a positive image pair according to standard footwear images with the same style in the style classification result includes: calculating a first image similarity between standard footwear images with the same style in the style classification result, and determining whether the first image similarity is less than a first preset similarity threshold and greater than a second preset similarity threshold; when it is determined that the first image similarity is less than the first preset similarity threshold and greater than the second preset similarity threshold, splicing the standard footwear images with the same style to obtain a first positive image pair, and when it is determined that the first image similarity is less than the second preset similarity threshold, splicing the standard footwear images with the same style to obtain a second positive image pair; splicing standard footwear images with the same style but different colors in the style classification result to obtain a third positive image pair, and splicing standard footwear images with the same style but different carriers in the style classification result to obtain a fourth positive image pair; generating the positive image pair according to the first positive image pair and / or the second positive image pair and / or the third positive image pair and / or the fourth positive image pair.
[0113] In an exemplary embodiment of the present disclosure, determining a negative image pair according to standard footwear images with different styles in the style classification result includes: randomly selecting standard footwear images with different styles from different style classification results, and splicing the standard footwear images with different styles to obtain a first negative image pair; calculating a second image similarity between the standard footwear images with different styles, and determining whether the second image similarity is greater than a third preset similarity threshold; when it is determined that the second image similarity is greater than the third preset similarity threshold, splicing the standard footwear images with different styles to obtain a second negative image pair, and generating the negative image pair according to the first negative image pair and / or the second negative image pair.
[0114] In an exemplary embodiment of the present disclosure, performing image enhancement processing on the standard image pair to obtain an enhanced image pair includes: transforming the original image color of the standard image pair, and / or adding image noise to the standard image pair, and / or performing geometric transformation on the standard image pair, and / or performing occlusion processing and / or blurring processing and / or affine transformation on the standard image pair to obtain the enhanced image pair.
[0115] In an exemplary embodiment of the present disclosure, the preset deep neural network model includes an angle discriminator, a first encoder, a first embedding mapping layer, and a multimodal large model; wherein, inputting the enhanced image pair into the preset deep neural network model to obtain a first similarity detection result includes: determining the image angle information of the enhanced image pair based on the angle discriminator, and generating first context information to be predicted according to the image angle information and preset model hint information;
[0116] encoding the first context information to be predicted by the text encoder in the first encoder to obtain a first context token sequence; encoding the enhanced image pair by the image encoder in the first encoder to obtain a first image pair feature, and performing embedding mapping processing on the first context token sequence and the first image pair feature based on the first embedding mapping layer to obtain a first overall context representation; performing similarity detection on the first overall context representation based on the multimodal large model to obtain the first similarity detection result.
[0117] In an exemplary embodiment of the present disclosure, the multimodal large model includes a second encoder, a multi-layer perceptron, a second embedding mapping layer, and a language model; wherein, performing similarity detection on the first overall context representation based on the multimodal large model to obtain the first similarity detection result includes: performing visual encoding processing on the image part in the first overall context representation based on the second encoder to obtain a visual encoding result, and performing classification processing on the visual encoding result based on the multi-layer perceptron to obtain an image classification result; performing embedding mapping processing on the text part in the first overall context representation based on the second embedding mapping layer to obtain a second context token sequence; performing root mean square normalization processing on the image classification result and the second context token sequence based on the language model to obtain the first similarity detection result.
[0118] In an exemplary embodiment of the present disclosure, adjusting the parameters of the preset deep neural network model based on the target loss function to obtain an image similarity detection model includes: determining the low-rank matrix parameters of the multimodal large model in the preset deep neural network model in the shoe image similarity comparison scenario based on the target loss function, and fine-tuning the multimodal large model based on the low-rank matrix parameters to obtain a fine-tuned large model; adjusting the parameters in the angle discriminator and the first embedding mapping layer in the preset deep neural network model based on the target loss function; generating the image similarity detection model according to the fine-tuned large model, the angle discriminator with adjusted parameters, the first embedding mapping layer with adjusted parameters, and the first encoder.
[0119] The exemplary embodiments of the present disclosure also provide a device for detecting image similarity. Specifically, referring to Figure 9 As shown, the device for detecting image similarity may include a pair of images to be detected generation module 910, a shoe product similarity detection result determination module 920, and an appearance infringement determination module 930. Among them:
[0120] The pair of images to be detected generation module 910 can be used to obtain a first shoe product image to be detected and a second shoe product image to be detected, and generate a pair of images to be detected based on the first shoe product image to be detected and the second shoe product image to be detected;
[0121] The shoe product similarity detection result determination module 920 can be used to input the pair of images to be detected into a preset image similarity detection model to obtain a shoe product similarity detection result; wherein, the image similarity detection model is obtained according to the training method of the similarity detection model described in any one of the above.
[0122] The appearance infringement determination module 930 can be used to determine whether there is an appearance infringement of shoe products between the first shoe product image to be detected and the second shoe product image to be detected according to the shoe product similarity detection result.
[0123] In an exemplary embodiment of the present disclosure, the shoe product similarity detection result includes at least one of a sole similarity detection result, a tongue similarity detection result, a vamp similarity detection result, a heel similarity detection result, and a carrying similarity detection result.
[0124] In an exemplary embodiment of the present disclosure, determining whether there is an appearance infringement of shoe products between the first shoe product image to be detected and the second shoe product image to be detected according to the shoe product similarity detection result includes: if the sole similarity detection result, the tongue similarity detection result, the vamp similarity detection result, the heel similarity detection result, and the carrying similarity detection result in the shoe product similarity detection result are all similar, it is determined that there is an appearance infringement of shoe products between the first shoe product image to be detected and the second shoe product image to be detected; if any one of the sole similarity detection result, the tongue similarity detection result, the vamp similarity detection result, the heel similarity detection result, and the carrying similarity detection result in the shoe product similarity detection result is dissimilar, it is determined that there is no appearance infringement of shoe products between the first shoe product image to be detected and the second shoe product image to be detected.
[0125] The specific details of each module in the above training device for the similarity detection model and the detection device for image similarity have been described in detail in the corresponding training method for the similarity detection model and the detection method for image similarity, and thus will not be elaborated here.
[0126] It should be noted that although several modules or units of a device for action execution are mentioned in the above detailed description, such a division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more of the above-described modules or units can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0127] In addition, although the steps of the methods in the present disclosure are described in a specific order in the drawings, this does not require or imply that these steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution, etc.
[0128] In an exemplary embodiment of the present disclosure, there is also provided an electronic device capable of implementing the above method. Those skilled in the art can understand that various aspects of the present disclosure can be implemented as a system, a method, or a program product. Therefore, various aspects of the present disclosure can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as a circuit, a module, or a system here.
[0129] Next, refer to Figure 10 to describe the electronic device 1000 according to this embodiment of the present disclosure. Figure 10 The electronic device 1000 shown is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.
[0130] As Figure 10 shown, the electronic device 1000 is presented in the form of a general-purpose computing device. The components of the electronic device 1000 may include, but are not limited to: at least one of the above-described processing units 1010, at least one of the above-described storage units 1020, a bus 1030 connecting different system components (including the storage unit 1020 and the processing unit 1010), and a display unit 1040.
[0131] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 1010, so that the processing unit 1010 executes the steps according to various exemplary embodiments of the present disclosure described in the "Exemplary Method" section above of this specification. For example, the processing unit 1010 can execute as Figure 1 the steps S110 shown in: preprocess the original footwear image to obtain a standard footwear image, and construct a standard image pair according to the standard footwear image; step S120: perform image enhancement processing on the standard image pair to obtain an enhanced image pair, and input the enhanced image pair into a preset deep neural network model to obtain a first similarity detection result; step S130: construct a target loss function according to the first similarity detection result and the first actual image label of the enhanced image pair; step S140: adjust the parameters of the preset deep neural network model based on the target loss function to obtain an image similarity detection model.
[0132] For another example, the processing unit 1010 can execute as Figure 9 the steps S910 shown in: obtain a first footwear product image to be detected and a second footwear product image to be detected, and generate an image pair to be detected according to the first footwear product image to be detected and the second footwear product image to be detected; step S920: input the image pair to be detected into a preset image similarity detection model to obtain a footwear product similarity detection result; wherein, the image similarity detection model is obtained according to the training method of the similarity detection model described in any one of the above; step S930: determine whether there is an appearance infringement behavior of footwear products between the first footwear product image to be detected and the second footwear product image to be detected according to the footwear product similarity detection result.
[0133] The storage unit 1020 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 10201 and / or a cache storage unit 10202, and may further include a read-only storage unit (ROM) 10203.
[0134] The storage unit 1020 may further include a program / utilities 10204 having a set (at least one) of program modules 10205. Such program modules 10205 include but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.
[0135] The bus 1030 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the various bus structures.
[0136] The electronic device 1000 can also communicate with one or more external devices 1100 (such as a keyboard, a pointing device, a Bluetooth device, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 1000, and / or communicate with any device that enables the electronic device 1000 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication can be carried out through the input / output (I / O) interface 1050. Moreover, the electronic device 1000 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 1060. As shown in the figure, the network adapter 1060 communicates with other modules of the electronic device 1000 through the bus 1030. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 1000, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0137] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software, or can be implemented by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, and the software product can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure. In an exemplary embodiment of the present disclosure, a computer-readable storage medium is also provided, on which a program product capable of implementing the above method of this specification is stored. In some possible implementation manners, various aspects of the present disclosure can also be implemented in the form of a program product, which includes program code, and when the program product runs on a terminal device, the program code is used to enable the terminal device to execute the steps according to various exemplary embodiments of the present disclosure described in the above "exemplary method" section of this specification.
[0138] The program product for implementing the above method according to the embodiments of the present disclosure can adopt a portable compact disc read-only memory (CD-ROM) and include program code, and can run on a terminal device, such as a personal computer. However, the program product of the present disclosure is not limited thereto. In this document, the readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device.
[0139] The program product may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the foregoing. More specific examples of the readable storage medium (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0140] The computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The readable signal medium may also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the readable medium may be transmitted with any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0141] The program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on the remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., by using an Internet service provider to connect through the Internet).
[0142] In addition, the above-mentioned drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present disclosure, and are not for limiting purposes. It is easy to understand that the processes shown in the above-mentioned drawings do not indicate or limit the chronological order of these processes. Additionally, it is also easy to understand that these processes may be executed, for example, synchronously or asynchronously in multiple modules.
[0143] Other embodiments of the present disclosure will be readily apparent to those skilled in the art upon consideration of the specification and practice of the invention herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common general knowledge or conventional technical means in the technical field not invented by the present disclosure. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims.
Claims
1. A training method for a similarity detection model, characterized in that: include: Preprocessing the original footwear image to obtain a standard footwear image, and constructing a standard image pair based on the standard footwear image; Performing image enhancement processing on the standard image pair to obtain an enhanced image pair, and inputting the enhanced image pair into a preset deep neural network model to obtain a first similarity detection result; constructing a target loss function according to the first similarity detection result and the first actual image label of the enhanced image pair; Based on the target loss function, the parameters of the preset deep neural network model are adjusted to obtain an image similarity detection model.
2. The training method of the similarity detection model according to claim 1, characterized in that: The original footwear image is preprocessed to obtain a standard footwear image, including: Performing image cropping processing on the original footwear image based on a preset image cropping model to obtain an image cropping result; The image cropping result is denoised based on a preset image denoising model to obtain the standard footwear image; wherein the standard footwear image is a footwear image with a single style.
3. The training method of the similarity detection model according to claim 1, characterized in that: Constructing a standard image pair according to the standard shoe image, including: Classify the standard footwear images based on a preset style classification model to obtain a style classification result, and determine a positive image pair based on standard footwear images of the same style in the style classification result; A reverse image pair is determined according to the standard footwear images with different styles in the style classification result, and the standard image pair is constructed according to the forward image pair and the reverse image pair.
4. The training method of the similarity detection model according to claim 3 is characterized in that: The positive image pairs are determined based on the standard footwear images of the same style in the style classification results, including: Calculating a first image similarity between standard footwear images having the same style in the style classification results, and determining whether the first image similarity is less than a first preset similarity threshold and greater than a second preset similarity threshold; When it is determined that the first image similarity is less than a first preset similarity threshold and greater than a second preset similarity threshold, the standard footwear images with the same style are spliced to obtain a first positive image pair, and when it is determined that the first image similarity is less than a second preset similarity threshold, the standard footwear images with the same style are spliced to obtain a second positive image pair; splicing the standard footwear images with the same style but different colors in the style classification results to obtain a third positive image pair, and splicing the standard footwear images with the same style but different carriers in the style classification results to obtain a fourth positive image pair; The forward image pair is generated according to the first forward image pair and / or the second forward image pair and / or the third forward image pair and / or the fourth forward image pair.
5. The training method of the similarity detection model according to claim 3, characterized in that: Determine reverse image pairs based on standard footwear images with different styles in the style classification results, including: Randomly selecting standard footwear images with different styles from different style classification results, and splicing the standard footwear images with different styles to obtain a first reverse image pair; calculating a second image similarity between standard footwear images having different styles, and determining whether the second image similarity is greater than a third preset similarity threshold; When it is determined that the second image similarity is greater than a third preset similarity threshold, the standard footwear images with different styles are spliced to obtain a second reverse image pair, and the reverse image pair is generated according to the first reverse image pair and / or the second reverse image pair.
6. The training method of the similarity detection model according to claim 1, characterized in that: Performing image enhancement processing on the standard image pair to obtain an enhanced image pair includes: The original image color of the standard image pair is transformed, and / or image noise is added to the standard image pair, and / or the standard image pair is geometrically transformed, and / or the standard image pair is occluded and / or blurred and / or affine transformed to obtain the enhanced image pair.
7. The training method of the similarity detection model according to claim 1, characterized in that: The preset deep neural network model includes an angle discriminator, a first encoder, a first embedding mapping layer and a multimodal large model; The enhanced image pair is input into a preset deep neural network model to obtain a first similarity detection result, including: Determining image angle information of the enhanced image pair based on the angle discriminator, and generating first context information to be predicted according to the image angle information and preset model prompt information; encoding the first to-be-predicted context information based on a text encoder in the first encoder to obtain a first context flag sequence; encoding the enhanced image pair based on the image encoder in the first encoder to obtain a first image pair feature, and embedding and mapping the first context marker sequence and the first image pair feature based on the first embedding mapping layer to obtain a first context overall representation; Based on the multimodal large model, a similarity detection is performed on the first context overall representation to obtain the first similarity detection result.
8. The training method of the similarity detection model according to claim 7, characterized in that: The multimodal large model includes a second encoder, a multi-layer perceptron, a second embedding mapping layer, and a language model; The similarity detection is performed on the first context overall representation based on the multimodal large model to obtain the first similarity detection result, including: Performing visual encoding processing on the image part in the first context overall representation based on the second encoder to obtain a visual encoding result, and performing classification processing on the visual encoding result based on the multi-layer perceptron to obtain an image classification result; Performing embedding mapping processing on the text portion in the first context overall representation based on the second embedding mapping layer to obtain a second context marker sequence; The image classification result and the second context marker sequence are subjected to root mean square normalization processing based on the language model to obtain the first similarity detection result.
9. The training method of the similarity detection model according to claim 1, characterized in that: The preset deep neural network model is parameterized based on the target loss function to obtain an image similarity detection model, including: Determining low-rank matrix parameters of a multimodal large model in a preset deep neural network model in a shoe image similarity comparison scenario based on the target loss function, and fine-tuning the multimodal large model based on the low-rank matrix parameters to obtain a fine-tuned large model; Adjusting parameters in the angle discriminator and the first embedding mapping layer in the preset deep neural network model based on the target loss function; The image similarity detection model is generated according to the fine-tuned large model, the angle discriminator after parameter adjustment, the first embedding mapping layer after parameter adjustment, and the first encoder.
10. A method for detecting image similarity, characterized in that: include: Acquire a first image of a shoe product to be detected and a second image of a shoe product to be detected, and generate a pair of images to be detected according to the first image of the shoe product to be detected and the second image of the shoe product to be detected; Inputting the image pair to be detected into a preset image similarity detection model to obtain a footwear product similarity detection result; wherein the image similarity detection model is obtained according to the training method of the similarity detection model according to any one of claims 1 to 9; According to the shoe product similarity detection result, it is determined whether there is any infringement on the appearance of the shoe products between the first shoe product image to be detected and the second shoe product image to be detected.
11. The method for detecting image similarity according to claim 10, characterized in that: The footwear product similarity detection result includes at least one of a sole similarity detection result, a tongue similarity detection result, an upper similarity detection result, a heel similarity detection result and a carrying similarity detection result.
12. The method for detecting image similarity according to claim 11, characterized in that: Determining whether there is an infringement of the appearance of the footwear products between the first to-be-detected footwear product image and the second to-be-detected footwear product image according to the footwear product similarity detection result includes: If the sole similarity detection result, the tongue similarity detection result, the upper similarity detection result, the heel similarity detection result and the carrying similarity detection result in the shoe product similarity detection results are all similar, it is determined that there is an infringement of the appearance of the shoe product between the first shoe product image to be detected and the second shoe product image to be detected; If any of the sole similarity detection result, tongue similarity detection result, upper similarity detection result, heel similarity detection result and carrying similarity detection result in the shoe product similarity detection results is dissimilar, it is determined that there is no infringement of the appearance of the shoe products between the first shoe product image to be detected and the second shoe product image to be detected.
13. A training device for a similarity detection model, characterized in that: include: An image pair construction module, used for preprocessing the original footwear image to obtain a standard footwear image, and constructing a standard image pair according to the standard footwear image; An image enhancement module, used for performing image enhancement processing on the standard image pair to obtain an enhanced image pair, and inputting the enhanced image pair into a preset deep neural network model to obtain a first similarity detection result; A loss function construction module, used to construct a target loss function according to the first similarity detection result and the first actual image label of the enhanced image pair; A parameter adjustment module is used to adjust the parameters of the preset deep neural network model based on the target loss function to obtain an image similarity detection model.
14. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the training method of the similarity detection model described in any one of claims 1 to 9 and the method for detecting image similarity described in any one of claims 10 to 12 are implemented.
15. An electronic device, characterized in that: include: processor; as well as A memory, configured to store executable instructions of the processor; The processor is configured to execute the similarity detection model training method described in any one of claims 1 to 9 and the image similarity detection method described in any one of claims 10 to 12 by executing the executable instructions.