A training method for an image detection model, an image detection method, and related devices

By generating and extracting text features and training network models with image features, the problems of small number of real training samples and poor characteristics in the prior art are solved, which improves the detection accuracy of the image detection model and reduces labor costs.

CN117173501BActive Publication Date: 2025-06-10CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310974673.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-03
Publication Date
2025-06-10
Estimated Expiration
2043-08-03

AI Technical Summary

Technical Problem

In the prior art, after using pre-trained samples to train the network model, when fine-tuning the real training samples, the detection accuracy improvement is limited, mainly because the number of real training samples is small and the image characteristics are poor.

Method used

By obtaining a collection of training samples, multiple texts are generated and text features are extracted, and network models are trained in combination with image features. The specific steps include: obtaining sample images of multiple categories, generating texts of each category, extracting text features, determining the loss parameters of each image, and adjusting the network model parameters based on the loss parameters.

Benefits of technology

It improves the detection accuracy of the image detection model, reduces the cost of manual labeling, increases the amount of data during the training process, and improves the robustness of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117173501B_ABST
    Figure CN117173501B_ABST
Patent Text Reader

Abstract

The present application provides a training method for an image detection model, an image detection method, and related devices, which are used to improve the detection accuracy of the image detection model and reduce the labor cost. The method includes: obtaining a training sample set, where the training sample set includes sample images of multiple categories; generating multiple texts according to the multiple categories, and extracting text features of each text; performing the following operations on each image in the training sample set to determine the loss parameter of each image: inputting each image into the network model to be trained, and extracting the first image features of the regions containing the target objects in each image; determining the similarity between the first image features and the text features of each text in the multiple texts; and determining the loss parameter of each image based on the multiple similarities corresponding to the determined first image features; adjusting the parameters of the network model to be trained based on the loss parameters of all the images in the training sample set and a preset loss function.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of object detection, and discloses a training method for an image detection model, an image detection method, and related devices. Background Art

[0002] In related technologies, a pre-trained network model is obtained by training a network using pre-training samples. Then, real training samples are input into the pre-trained network model for fine-tuning training to obtain a final model. The pre-training samples usually have a large quantity and rich image features, such as color, shape, texture and other features. While the real training samples usually have a small quantity and poor image features. When using the real training samples to fine-tune the training network model, the effect is poor, resulting in limited improvement in detection accuracy. Summary of the Invention

[0003] The present application provides a training method for an image detection model, an image detection method, and related devices, which are used to improve the detection accuracy of the image detection model and reduce the labor cost.

[0004] In a first aspect, an embodiment of the present application provides a training method for an image detection model, including:

[0005] Obtain a training sample set, where the training sample set includes sample images of multiple categories, one sample image corresponds to one category, and the label of the sample image is the category corresponding to the sample image;

[0006] Generate multiple texts according to the multiple categories, and extract text features of each text, where the multiple texts correspond to the multiple categories one by one, and each text includes the category corresponding to each text;

[0007] For each image in the training sample set, perform the following operations to determine the loss parameter of each image: input each image into the network model to be trained, and extract the first image feature of the region containing the target object in each image; determine the similarity between the first image feature and the text features of each text in the multiple texts; and determine the loss parameter of each image based on the multiple similarities corresponding to the determined first image feature;

[0008] Adjust the parameters of the network model to be trained based on the loss parameters of all images in the training sample set and a preset loss function.

[0009] In a possible implementation manner, in the training method for an image detection model provided by an embodiment of the present application, the generating multiple texts according to the multiple categories includes:

[0010] Generate each text according to at least one extended field and the category corresponding to each text, where the extended field includes one or more of the following:

[0011] The acquisition method of the images in the training sample set and the scenes to which the images in the training sample set belong.

[0012] In a possible implementation manner, in the training method of the image detection model provided by the embodiments of the present application, the training sample set includes: a plurality of sample subsets; the plurality of sample subsets include a first sample subset and at least one second sample subset;

[0013] The acquisition method of the sample images in the first sample subset is different from the acquisition method of the sample images in the second sample subset;

[0014] And, the total number of sample images in the first sample subset is greater than the total number of sample images in each of the at least one second sample subsets.

[0015] In a possible implementation manner, in the training method of the image detection model provided by the embodiments of the present application, in the loss function, the weight of the loss sub-function corresponding to the first sample subset is less than the weight of the loss sub-function corresponding to the at least one second sample subset.

[0016] For example, the loss function can be L = α M *Lpre + α Q *Lreal, where Lpre represents the loss sub-function corresponding to the first sample subset, and α M represents the weight of the loss sub-function corresponding to the first sample subset; Lreal represents the loss sub-function corresponding to the at least one second sample subset, and α Q represents the weight of the loss sub-function corresponding to the at least one second sample subset.

[0017] In some examples, the The Wherein, the represents the image feature of the a-th sample image in the first sample subset pre; real_N represents the N-th second sample subset among the at least one second sample subset, NT represents the total number of the at least one second sample subset, and the represents the image feature of the b-th sample image in the N-th second sample subset.

[0018] The Characterize the target text feature corresponding to the image feature of the a-th sample image in the first sample subset pre. Among the similarities between the image feature of the a-th sample image in the first sample subset pre and the text features of the multiple texts, the similarity between the target text feature corresponding to the image feature of the a-th sample image in the first sample subset pre and the image feature of the a-th sample image in the first sample subset pre is the largest;

[0019] The Characterize the target text feature corresponding to the image feature of the b-th sample image in the N-th second sample subset. Among the similarities between the image feature of the b-th sample image in the N-th second sample subset and the text features of the multiple texts, the similarity between the target text feature corresponding to the b-th sample image in the N-th second sample subset and the image feature of the b-th sample image in the N-th second sample subset is the largest.

[0020] Characterize the Loss parameter of Characterize the Loss parameter of; where The F te_k Characterize the k-th text feature among the text features of the multiple texts; Characterize the And the Similarity of Characterize the And the Similarity of, where τ is a constant.

[0021] In a possible implementation manner, in the training method of the image detection model provided by the embodiments of the present application, the multiple sample subsets include multiple second sample subsets, where the scene to which the image in any second sample subset belongs is different from the scenes to which the images in the multiple second sample subsets other than the image in the any second sample subset belong.

[0022] In a possible implementation manner, in the training method of the image detection model provided by the embodiments of the present application, the loss function includes a loss sub-function corresponding to the first sample subset and a loss sub-function corresponding to each second sample subset in the second sample subsets;

[0023] Among them, the weight of the loss sub-function corresponding to the first sample subset is less than the minimum value of the weights of the loss sub-functions corresponding to the multiple second sample subsets.

[0024] For example, in the training method of the image detection model provided by the embodiments of the present application, the loss function is:

[0025]

[0026] Among them, the Lpre represents the loss sub-function corresponding to the first sample subset, and the α M represents the weight of the loss sub-function corresponding to the first sample subset; the NT represents the total number of the multiple second sample subsets, and the represents the loss sub-function corresponding to the Nth second sample subset among the multiple second sample subsets, and the α real_N represents the loss sub-function corresponding to the Nth second sample subset.

[0027] In some examples, the Sohu Among them, the represents the image feature of the ath sample image in the first sample subset pre; the represents the image feature of the bth sample image in the Nth second sample subset;

[0028] The represents the target text feature corresponding to the image feature of the ath sample image in the first sample subset pre. Among the similarities between the image feature of the ath sample image in the first sample subset pre and the text features of the multiple texts, the similarity between the target text feature corresponding to the image feature of the ath sample image in the first sample subset pre and the image feature of the ath sample image in the first sample subset pre is the largest;

[0029] The represents the target text feature corresponding to the image feature of the bth sample image in the Nth second sample subset. Among the similarities between the image feature of the bth sample image in the Nth second sample subset and the text features of the multiple texts, the similarity between the target text feature corresponding to the bth sample image in the Nth second sample subset and the image feature of the bth sample image in the Nth second sample subset is the largest;

[0030] represents the loss parameter, represents the loss parameter; among them, The F te_k represents the kth text feature among the text features of the multiple texts; represents the similarity between and represents the similarity between The similarity, where τ is a constant.

[0031] In a second aspect, an image detection method provided by an embodiment of the present application may include:

[0032] Obtain an image to be detected;

[0033] If the image to be detected includes a target object, extract the image features of the region including the target object;

[0034] Determine the similarity between the image features and the text features of each text in a plurality of texts, where the plurality of texts are pre-generated, and the plurality of texts correspond one-to-one to a plurality of object categories, and each text includes the corresponding object category;

[0035] Determine the object category included in the target text as the object category of the target object, where the similarity between the image features and the text features of the target text is the largest among the plurality of similarities corresponding to the determined image features.

[0036] In a possible implementation manner, in the image detection method provided by an embodiment of the present application, the number of characters in each text is greater than a preset character number threshold.

[0037] In a possible implementation manner, in the image detection method provided by an embodiment of the present application, the image to be detected is collected by an image acquisition device in a kitchen scene;

[0038] The plurality of object categories include:

[0039] Wearing a chef's hat and a mask, wearing a chef's hat and not wearing a mask, not wearing a chef's hat and wearing a mask, not wearing a chef's hat and not wearing a mask.

[0040] In a third aspect, an electronic device provided by an embodiment of the present application may include a memory and a processor;

[0041] The memory is used to store program instructions;

[0042] The processor is used to execute the program instructions to implement the method as described in the first aspect and any of its possible implementation manners, or execute the method as described in the second aspect and any of its possible implementation manners.

[0043] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, which may include computer program instructions. When the computer program instructions are executed by a computer, the computer executes the method as described in the first aspect and any of its possible implementation manners, or executes the method as described in the second aspect and any of its possible implementation manners.

[0044] In a fifth aspect, an embodiment of the present application further provides a training device, including:

[0045] A training sample set obtaining module, configured to obtain a training sample set, where the training sample set includes sample images of multiple categories, one sample image corresponds to one category, and the label of the sample image is the category corresponding to the sample image;

[0046] A text feature generation module, configured to generate multiple texts according to the multiple categories, and extract text features of each text, where the multiple texts correspond to the multiple categories one by one, and each text includes the category corresponding to each text;

[0047] A model training module, configured to perform the following operations on each image in the training sample set: determine a loss parameter of each image: input each image into a network model to be trained, and extract a first image feature of a region including a target object in each image; determine similarities between the first image feature and text features of each text in the multiple texts; and determine the loss parameter of each image based on the determined multiple similarities corresponding to the first image feature; and adjust parameters of the network model to be trained based on loss parameters of all images in the training sample set and a preset loss function.

[0048] In a sixth aspect, an embodiment of the present application further provides an image detection device, including:

[0049] An image acquisition module, configured to acquire an image to be detected;

[0050] An image detection module, configured to, if the image to be detected includes a target object, extract an image feature of a region including the target object; determine similarities between the image feature and text features of each text in multiple texts, where the multiple texts are pre-generated, and the multiple texts correspond to multiple object categories one by one, and each text includes a corresponding object category; and determine the object category included in the target text as the object category of the target object, where the similarity between the image feature and the text feature of the target text is the largest among the determined multiple similarities corresponding to the image feature.

[0051] The beneficial effects of the embodiments of the present application are as follows:

[0052] The present application provides a training method for an image detection model, an image detection method, and related devices. The method can generate extended texts for each category by using the categories of sample images in a training sample set. By using the text features of the extended texts and the image features of the sample images to train a network model, this approach can increase the amount of data during the training process, improve the robustness of the model, enrich the features of the sample images, have better detection accuracy, and do not require a large number of real training samples. Therefore, it can also reduce the manual annotation cost. There is no situation where the difference in the feature distributions between the pre-training samples and the real training samples leads to a poor effect in training the network model.

[0053] Other features and advantages of the present application will be described in the following specification. And, in part, they will be obvious from the specification or understood by implementing the present application. The objectives and other advantages of the present application can be realized and obtained by the structures specifically pointed out in the written specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] To more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0055] Figure 1 For the training process of the related art;

[0056] Figure 2 Schematic flowchart of the training method for the image detection model provided by the embodiment of the present application;

[0057] Figure 3 Schematic diagram of the training process provided by the embodiment of the present application;

[0058] Figure 4 Schematic diagram of a pre-training sample image;

[0059] Figure 5 Schematic diagram of a real-scene sample image;

[0060] Figure 6 Schematic diagram of the training process provided by the embodiment of the present application;

[0061] Figure 7 Schematic flowchart of the image detection method provided by the embodiment of the present application;

[0062] Figure 8 Schematic diagram of the structure of a training device provided by the embodiment of the present application;

[0063] Figure 9Schematic structural diagram of a detection device provided by an embodiment of the present application;

[0064] Figure 10 Schematic structural diagram of an electronic device provided by an embodiment of the present application;

[0065] Figure 11 Schematic structural diagram of the electronic device provided by an embodiment of the present application. Detailed implementation manners

[0066] To make the objectives, technical solutions, and advantages of the present application clearer and more understandable, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part rather than all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application. Without conflict, the embodiments in the present application and the features in the embodiments may be arbitrarily combined with each other. And although the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a sequence different from that here.

[0067] In the description and claims of the present application and the above-mentioned accompanying drawings, the terms "first" and "second" are used to distinguish different objects rather than to describe a specific order. In addition, the term "comprising" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally further include steps or units not listed, or may optionally further include other steps or units inherent to these processes, methods, products, or devices. "Multiple" in the present application may represent at least two, for example, it may be two, three, or more, and the embodiments of the present application do not make limitations.

[0068] In the technical solution of the present application, the acquisition, transmission, use, etc. of image data all comply with the requirements of relevant national laws and regulations.

[0069] Before introducing a training method for an image detection model provided by an embodiment of the present application, for the convenience of understanding, first, the technical background of the embodiments of the present application will be introduced in detail below.

[0070] In the related art, such as Figure 1As shown in the figure, a network model is trained using a pre-training sample set to obtain a pre-trained network model. Each pre-training sample is input into the network, and the network is trained with the label of the output sample as the target to obtain a pre-trained network model. The pre-trained network model is fine-tuned using real training samples. Each real training sample is input into the pre-trained network model, and the pre-trained network model is adjusted with the label of the output sample as the target to obtain a final detection model.

[0071] Generally, the pre-training sample set is pre-acquired labeled images, and the image features in the pre-training sample set are relatively rich. The real training sample set consists of images in the actual model application scenario, and the labor cost required for image annotation is relatively high, resulting in a small number of samples in the real training sample set. In addition, the image features in the real training sample set are poor. All these make the effect of fine-tuning the pre-trained network model on real sample images not ideal, and the improvement of detection accuracy is limited.

[0072] In view of this, the embodiments of the present application provide a training method for an image detection model, an image detection method, and related devices. The label of the sample image is expanded into text, and the text features are extracted. The network model is trained by combining the text features and the image features, so that the training of the image detection model combines semantic information and image information, and the detection accuracy can be improved.

[0073] Figure 2 According to an exemplary embodiment, a training method for an image detection model is shown. This method can be executed by an electronic device, and this method may include the following steps:

[0074] S201, obtain a training sample set, the training sample set includes sample images of multiple categories, one sample image corresponds to one category, and the label of the sample image is the category corresponding to the sample image.

[0075] Specifically, as Figure 3 shown, when the electronic device performs image detection model training, the training sample set can be input into the network model and the result is output. According to the output result, the network model parameters are adjusted to implement the process of training the network model. In the training method for the image detection model provided by the embodiments of the present application, no pre-trained network model is generated.

[0076] Among them, the training sample set obtained by the electronic device may include sample images from multiple sources, and may include the aforementioned pre-training sample images and real-scene sample images. To facilitate the introduction of the method provided by the present application, taking the application scenario as the kitchen and the detection purpose as detecting whether the chef in the kitchen wears a chef's hat and a mask as an example for introduction. It should be noted that the training method provided by the present application can also be applied to other scenarios and other detection purposes.

[0077] As shown Figure 4 in the figure, the pre-trained sample images are usually processed images with relatively rich image features. The real-scene sample images are generally collected in the application scenario, as shown Figure 5 in the figure, with poor image features.

[0078] The training sample set may include sample images of multiple categories. Among them, one sample image corresponds to one category, and the category corresponding to the sample image is used as the label of the sample image. For example, the multiple categories can be respectively recorded as: wearing a chef's hat and wearing a mask, wearing a chef's hat and not wearing a mask, not wearing a chef's hat and wearing a mask, not wearing a chef's hat and not wearing a mask. Another example, the multiple categories can be respectively recorded as: chef‘s hat+mask, chef‘s hat+nomask, no chef’s hat+mask, nochef’s hat+nomask. The embodiments of the present application do not specifically limit the language and form of the image labels, but can respectively represent the meanings corresponding to each category recorded in Chinese.

[0079] S202. Generate multiple texts according to the multiple categories, and extract the text features of each text, where the multiple texts correspond to the multiple categories one by one, and each text includes the category corresponding to each text.

[0080] Specifically, the electronic device can expand each category to generate corresponding texts. Optionally, the electronic device can generate each text according to at least one expansion field and the category corresponding to each text, where the expansion field includes one or more of the following: the acquisition method of the images in the training sample set, the scene to which the images in the training sample set belong. The acquisition method of the images can represent the source of the images. The scene to which the images belong can represent the application scenario of the network model.

[0081] In some examples, the texts generated by the electronic device expanding each category may include the acquisition method of the sample images. For example, when expanding each category, the generated texts are respectively:

[0082] Text 1: An online image of a person wearing a chef's hat and wearing amask;

[0083] Text 2: An online image of a person not wearing a chef's hat andwearing a mask;

[0084] Text 3: A real photo of a person wearing a chef's hat and not wearing a mask;

[0085] Text 4: A real photo of a person not wearing a chef's hat and wearing a mask.

[0086] Among them, "online" represents a way of obtaining an image, and "real" represents a way of obtaining an image.

[0087] In some other examples, the text generated by the electronic device for each category expansion may include the acquisition method of the sample image and the scene to which the image belongs. For example, for each category expansion, the generated texts are respectively:

[0088] Text 1: An online image of a person wearing a chef's hat and wearing a mask in scene X;

[0089] Text 2: An online image of a person not wearing a chef's hat and wearing a mask in scene X;

[0090] Text 3: A real photo of a person wearing a chef's hat and not wearing a mask in scene X;

[0091] Text 4: A real photo of a person not wearing a chef's hat and wearing a mask in scene X.

[0092] Among them, "online" represents a way of obtaining an image, and "real" represents a way of obtaining an image. "scene X" represents the scene X to which the image belongs, that is, the image is collected (or taken) in scene X. For example, "scene1" can represent the scene 1 to which the image belongs, and "scene2" can represent the scene 2 to which the image belongs. In this application, the acquisition method of the image may include but is not limited to: online acquisition, generation, collection, etc. In the foregoing examples, "online" can represent online acquisition. "real" can represent collection.

[0093] The electronic device can extract the text features of the generated text, and can obtain the text feature 1 of text 1, the text feature 2 of text 2, the text feature 3 of text 3, and the text feature 4 of text 4.

[0094] Optionally, the number of characters of the text generated by the electronic device is greater than or equal to a preset character threshold. For example, the preset character threshold can be 16 characters. Optionally, the number of characters of the text generated by the electronic device can be greater than or equal to 128 characters.

[0095] S203. For each image in the training sample set, perform the following operations to determine the loss parameter of each image: input each image into the network model to be trained, and extract the first image feature of the region containing the target object in each image; determine the similarity between the first image feature and the text feature of each text in the multiple texts; and based on the multiple similarities corresponding to the determined first image feature, determine the loss parameter of each image.

[0096] S204. Based on the loss parameters of all the images in the training sample set and a preset loss function, adjust the parameters of the network model to be trained.

[0097] In a possible implementation manner, during steps 203 and 204, the loss function configured in the electronic device may not distinguish the acquisition methods of each sample image. A first loss function may be configured in the electronic device: L = ∑ i∈训练样本集合 -logp(F im_i , F tem_i );

[0098] wherein, the F im_i represents the image feature of the i-th sample image in the training sample set, and the F te_k represents the k-th text feature among the text features of the multiple texts; the F tem_i represents the target text feature corresponding to the i-th sample image. Among the similarities between the image feature of the i-th sample image and the text features of the multiple texts, the similarity between the target text feature corresponding to the i-th sample image and the image feature of the i-th sample image is the largest;

[0099] the p(F im_i , F tem_i ) represents the loss parameter of the F im_i , S(F im_i , F tem_i ) represents the similarity between the F im_i and the F tem_i , S(Fim_i , F te_k ) representing the F im_i similarity with the F te_k , where τ is a constant.

[0100] For each sample image, the electronic device can determine the similarity between the sample image and each text feature. For example, where F im represents any one image feature, and F te represents any one text feature.

[0101] For each sample image, the electronic device can determine the text feature corresponding to the maximum value among the similarities between the image feature of the sample image and each text feature as the target text feature corresponding to the image feature of the sample image. And based on the similarity between the image feature of the sample image and each text feature, and the target text feature corresponding to the image feature of the sample image, the electronic device determines the loss parameter of the sample image. For example, the loss parameter of the F im_i is:

[0102]

[0103] Based on the above first loss function and the loss parameter of each sample image, the electronic device can determine the loss value of this training, and after adjusting the model based on this loss value, perform the foregoing training process again. Optionally, the electronic device can end the training after determining that the loss value is less than the preset loss value threshold, and use the trained model as the trained image detection model.

[0104] In another possible implementation, during steps 203 and 204, the loss function configured in the electronic device can distinguish the acquisition methods of each sample image and the scene to which the image belongs. For example, the training sample set may include: multiple sample subsets. The multiple sample subsets may include a first sample subset and at least one second sample subset. The acquisition method of the sample images in the first sample subset is different from the acquisition method of the sample images in the second sample subset. As an example, the acquisition method of the sample images in the first sample subset may be obtained or generated online, and the acquisition method of the sample images in the second sample subset may be collected. The total number of sample images in the first sample subset is greater than the total number of sample images in each of the at least one second sample subsets.

[0105] In some examples, the multiple sample subsets include at least one second sample subset. In the case where the at least one second sample subset is a plurality of second sample subsets, for any second sample subset, the scene to which the image in the second sample subset belongs is different from the scene to which the image in the at least one second sample subset other than the image in the any second sample subset belongs. For example, the second sample subset may include Second Sample Subset 1 and Second Sample Subset 2. The scene to which the image in Second Sample Subset 1 belongs may be Scene 1. The scene to which the image in Second Sample Subset 2 belongs may be Scene 2.

[0106] A second loss function may be configured in the electronic device:

[0107] L = α M *Lpre + α Q *Lreal

[0108] Wherein, the Lpre represents the loss sub-function corresponding to the first sample subset, and the α M represents the weight of the loss sub-function corresponding to the first sample subset; the Lreal represents the loss sub-function corresponding to the at least one second sample subset, and the α Q represents the weight of the loss sub-function corresponding to the at least one second sample subset.

[0109] Optionally, the the Wherein, the represents the image feature of the a-th sample image in the first sample subset pre; the real_N represents the N-th second sample subset in the at least one second sample subset, the NT represents the total number of the at least one second sample subset, and the represents the image feature of the b-th sample image in the N-th second sample subset.

[0110] the represents the target text feature corresponding to the image feature of the a-th sample image in the first sample subset pre. Among the similarities between the image feature of the a-th sample image in the first sample subset pre and the text features of the multiple texts, the similarity between the target text feature corresponding to the image feature of the a-th sample image in the first sample subset pre and the image feature of the a-th sample image in the first sample subset pre is the largest;

[0111] the Characterize the target text feature corresponding to the image feature of the b-th sample image in the N-th second sample subset. Among the similarities between the image feature of the b-th sample image in the N-th second sample subset and the text features of the multiple texts, the similarity between the target text feature corresponding to the b-th sample image in the N-th second sample subset and the image feature of the b-th sample image in the N-th second sample subset is the largest.

[0112] Characterize the Loss parameter of Characterize the Loss parameter of; where The F te_k Characterize the k-th text feature among the text features of the multiple texts; Characterize the And the Similarity of Characterize the And the Similarity of, where τ is a constant.

[0113] For each sample image in each sample subset, the electronic device can determine the text feature corresponding to the maximum value among the similarities between the image feature of the sample image and each text feature as the target text feature corresponding to the image feature of the sample image. And based on the similarity between the image feature of the sample image and each text feature, and the target text feature corresponding to the image feature of the sample image, determine the loss parameter of the sample image.

[0114] For example Loss parameter of For example Loss parameter

[0115] Based on the above second loss function and the loss parameters of each sample image in each sample subset, the electronic device can determine the loss value of this training, and based on this loss value, adjust the model and then perform the foregoing training process again. Optionally, the electronic device can end the training after determining that the loss value is less than the preset loss value threshold, and use the trained model as the trained image detection model.

[0116] In a possible design, in the above second loss function, the weight α of the loss sub-function Lpre corresponding to the first sample subset M Is less than the weight α of the loss sub-function Lreal corresponding to the at least one second sample subset QIn such a design, the detection effect of the image detection model for detecting images obtained through the acquisition method corresponding to the second sample subset can be improved.

[0117] In some other examples, the multiple sample subsets include multiple second sample subsets, where the scene to which the images in any one of the second sample subsets belong is different from the scene to which the images in the multiple second sample subsets other than the images in the any one of the second sample subsets belong. For example, the second sample subsets may include a second sample subset 1 and a second sample subset 2. The scene to which the images in the second sample subset 1 belong may be scene 1. The scene to which the images in the second sample subset 2 belong may be scene 2.

[0118] In the electronic device, a third loss function may be configured:

[0119]

[0120] wherein, the Lpre represents the loss sub-function corresponding to the first sample subset, and the α M represents the weight of the loss sub-function corresponding to the first sample subset; the NT represents the total number of the multiple second sample subsets, and the represents the loss sub-function corresponding to the Nth second sample subset among the multiple second sample subsets, and the α real_N represents the loss sub-function corresponding to the Nth second sample subset.

[0121] Optionally, the the wherein, the represents the image feature of the ath sample image in the first sample subset pre; the represents the image feature of the bth sample image in the Nth second sample subset;

[0122] the represents the target text feature corresponding to the image feature of the ath sample image in the first sample subset pre. Among the similarities between the image feature of the ath sample image in the first sample subset pre and the text features of the multiple texts, the similarity between the target text feature corresponding to the image feature of the ath sample image in the first sample subset pre and the image feature of the ath sample image in the first sample subset pre is the largest;

[0123] the Characterize the target text feature corresponding to the image feature of the b-th sample image in the N-th second sample subset. Among the similarities between the image feature of the b-th sample image in the N-th second sample subset and the text features of the multiple texts, the similarity between the target text feature corresponding to the b-th sample image in the N-th second sample subset and the image feature of the b-th sample image in the N-th second sample subset is the largest;

[0124] Characterize the Loss parameter of Characterize the Loss parameter of; where The F te_k Characterize the k-th text feature among the text features of the multiple texts; Characterize the And the Similarity of Characterize the And the Similarity of, where τ is a constant.

[0125] For each sample image in each sample subset, the electronic device can determine the text feature corresponding to the maximum value among the similarities between the image feature of the sample image and each text feature as the target text feature corresponding to the image feature of the sample image. And based on the similarity between the image feature of the sample image and each text feature, and the target text feature corresponding to the image feature of the sample image, determine the loss parameter of the sample image.

[0126] For example Loss parameter of For example Loss parameter

[0127] Based on the above third loss function and the loss parameter of each sample image in each sample subset, the electronic device can determine the loss value of this training, and based on this loss value, adjust the model and then perform the foregoing training process again. Optionally, the electronic device can end the training after determining that the loss value is less than the preset loss value threshold, and use the trained model as the trained image detection model.

[0128] In a possible design, in the above third loss function, the weight α of the loss sub-function Lpre corresponding to the first sample subset M Is less than the weight α of the loss sub-function corresponding to each second sample subset among the multiple second sample subsets Of real_NThe minimum value of (where N ranges from 1 to NT). For example, multiple second sample subsets may include second sample subset 1 and second sample subset 2. The loss sub-function corresponding to the second sample subset 1 has a weight denoted as α real_1 , and the loss sub-function corresponding to the second sample subset 2 has a weight denoted as α real_2 . The minimum value of α real_1 and α real_2 is greater than the weight α M of the loss sub-function Lpre corresponding to the first sample subset.

[0129] Optionally, in the above third loss function, the weights of the loss sub-functions corresponding to each second sample subset can be configured in combination with the actual application scenario. In some examples, the weights of the loss sub-functions corresponding to each second sample subset can be determined based on the number of samples in each second sample subset. For example, the weight of the loss sub-function corresponding to the second sample subset with a large number of samples can be less than the weight of the loss sub-function corresponding to the second sample subset with a small number of samples.

[0130] Figure 6 Exemplarily, a training process of an image detection model is shown. Among them, the electronic device can obtain a training sample set, where the training sample set may include a first sample subset and a second sample subset. Optionally, the first sample subset may be the pre-training sample set I in the related art pre , and the second sample subset may be the real training set I in the related art real .

[0131] The electronic device can use the Region Proposal Network (RPN) to find the target object. For example, in the aforementioned kitchen scenario, the target object may be a person. Among them, the region of interest found by the RPN contains the target object. The electronic device can use the image encoder to generate the image features of the regions of interest of each sample image, the image features of the sample image.

[0132] The electronic device can generate extended texts for each label based on the labels of the samples in the training sample set. Optionally, the electronic device can also generate extended texts for each label in combination with the acquisition method of the samples in the first sample subset, the acquisition method of the samples in the second sample subset, and the scenario to which the samples in the second sample subset belong. The electronic device can use the text encoder to generate the text features of each text.

[0133] For the image feature of each sample image, the electronic device can calculate the similarity between each image feature and each text feature respectively, and determine the loss value based on a preset loss function, so as to adjust the network parameters of the training model.

[0134] In some examples, the training sample set may include a first sample subset pre and a second sample subset real_1. The electronic device may be configured with a fourth loss function: L = 0.3 * Lpre + 0.7 * Lreal.

[0135] In the fourth loss function, the The Among them, the weight of the loss sub-function corresponding to the first sample subset is 0.3, and the weight of the loss sub-function corresponding to the second sample subset real_1 is 0.7.

[0136] Among them, the represents the image feature of the a-th sample image in the first sample subset pre, the real_N represents the N-th second sample subset among the multiple second sample subsets, and the represents the image feature of the b-th sample image in the second sample subset real_1; the represents the target text feature corresponding to the image feature of the a-th sample image in the first sample subset pre. Among the similarities between the image feature of the a-th sample image in the first sample subset pre and the text features of the multiple texts, the similarity between the target text feature corresponding to the image feature of the a-th sample image in the first sample subset pre and the image feature of the a-th sample image in the first sample subset pre is the largest;

[0137] The represents the target text feature corresponding to the image feature of the b-th sample image in the second sample subset real_1. Among the similarities between the image feature of the b-th sample image in the second sample subset real_1 and the text features of the multiple texts, the similarity between the target text feature corresponding to the b-th sample image in the second sample subset real_1 and the image feature of the b-th sample image in the second sample subset real_1 is the largest;

[0138] represents the loss parameter of represents the loss parameter of; among them, The F te_k represents the k-th text feature among the text features of the multiple texts; represents the similarity between and represents the similarity between and, where τ is a constant.

[0139] In some examples, the training sample set may include a first sample subset pre, a second sample subset real_1, and a second sample subset real_2. The electronic device may be configured with a fifth loss function:

[0140] L = 0.2 * Lpre + 0.4 * Lreal

[0141] In the fifth loss function, the The

[0142] NT = 2. Among them, the weight of the loss sub-function corresponding to the first sample subset is 0.2, and the weight of the loss sub-function corresponding to the second sample subset is 0.4.

[0143] Among them, the represents the image feature of the a-th sample image in the first sample subset pre, the real_N represents the N-th second sample subset among the multiple second sample subsets, and the represents the image feature of the b-th sample image in the second sample subset real_1; the represents the target text feature corresponding to the image feature of the a-th sample image in the first sample subset pre. Among the similarities between the image feature of the a-th sample image in the first sample subset pre and the text features of the multiple texts, the similarity between the target text feature corresponding to the image feature of the a-th sample image in the first sample subset pre and the image feature of the a-th sample image in the first sample subset pre is the largest;

[0144] The represents the target text feature corresponding to the image feature of the b-th sample image in the second sample subset real_1. Among the similarities between the image feature of the b-th sample image in the second sample subset real_1 and the text features of the multiple texts, the similarity between the target text feature corresponding to the b-th sample image in the second sample subset real_1 and the image feature of the b-th sample image in the second sample subset real_1 is the largest;

[0145] represents the loss parameter of, represents the loss parameter of; among them, The F te_k represents the k-th text feature among the text features of the multiple texts; represents the and the similarity of, represents the similarity with the , where τ is a constant.

[0146] The represents the image feature of the c-th sample image in the N-th second sample subset; the represents the target text feature corresponding to the image feature of the c-th sample image in the second sample subset real_2. Among the similarities between the image feature of the c-th sample image in the second sample subset real_2 and the text features of the multiple texts, the similarity between the target text feature corresponding to the c-th sample image in the second sample subset real_2 and the image feature of the c-th sample image in the second sample subset real_2 is the largest;

[0147] represents the loss parameter of the represents the similarity with the .

[0148] Figure 7 Exemplarily, an image detection method is shown, which can be executed by a processor or an electronic device. The method may include:

[0149] S701, obtain an image to be detected.

[0150] Taking the processor executing the image detection method provided by this application as an example for introduction. In some possible situations, the processor is configured with a pre-trained image detection model. Optionally, the training process of the image detection model can be performed on other electronic devices, and the processor stores the trained image detection model. Or the training process of the image detection model is performed in the processor. After the processor trains the image detection model, it can use the image detection model to perform the detection task.

[0151] S702, if the image to be detected includes a target object, extract the image features of the region including the target object.

[0152] Specifically, when implemented, the processor can use the RPN technology to obtain the region of interest, that is, the region including the target object. And use the features of this region as the image features of the image to be detected.

[0153] S703, determine the similarity between the image features and the text features of each text in multiple texts, where the multiple texts are pre-generated, and the multiple texts correspond to multiple object categories one by one, and each text includes the corresponding object category.

[0154] In specific implementation, the multiple texts pre-generated by the electronic device are respectively extended texts of multiple object categories. Each text may include the corresponding object category. The electronic device may extract the text features of each text.

[0155] S704, determine the object category included in the target text as the object category of the target object, where the similarity between the image feature and the text feature of the target text is the largest among the multiple similarities corresponding to the determined image feature.

[0156] In specific implementation, the electronic device may determine the text corresponding to the text feature with the largest similarity to the image feature of the image to be measured as the target text. The electronic device determines the object category included in the target text as the object category of the image to be measured, so as to implement the detection of the object category corresponding to the image to be measured.

[0157] In some examples, the number of characters in the text generated by the electronic device is greater than or equal to a preset character threshold. For example, the preset character threshold may be 16 characters. Optionally, the number of characters in the text generated by the electronic device may be greater than or equal to 128 characters.

[0158] In some examples, the image to be measured is collected by an image acquisition device in a kitchen scene; the target object may be a person. The multiple object categories include: wearing a chef's hat and a mask, wearing a chef's hat and not wearing a mask, not wearing a chef's hat and wearing a mask, not wearing a chef's hat and not wearing a mask.

[0159] Based on the same technical concept, the embodiment of the present application provides a training device, which can achieve the same technical effect as the foregoing training method, and will not be elaborated here. Refer to Figure 8 , the device includes a sample set acquisition module, a text feature generation module, and a model training module. Among them:

[0160] The sample set acquisition module is used to acquire a training sample set, the training sample set includes sample images of multiple categories, one sample image corresponds to one category, and the label of the sample image is the category corresponding to the sample image;

[0161] The text feature generation module is used to generate multiple texts according to the multiple categories, and extract the text features of each text, where the multiple texts correspond to the multiple categories one by one, and each text includes the category corresponding to each text;

[0162] A model training module is configured to perform the following operations on each image in the training sample set: determining a loss parameter for each image: inputting each image into a network model to be trained, and extracting first image features of a region containing a target object in each image; determining the similarity between the first image features and the text features of each text in the multiple texts; and determining the loss parameter for each image based on the multiple similarities corresponding to the determined first image features; and adjusting the parameters of the network model to be trained based on the loss parameters of all the images in the training sample set and a preset loss function.

[0163] In a possible implementation manner, the text feature generation module is specifically configured to generate each text according to at least one augmentation field and the category corresponding to each text, where the augmentation field includes one or more of the following:

[0164] The acquisition method of the images in the training sample set, and the scenes to which the images in the training sample set belong.

[0165] In a possible implementation manner, the training sample set includes: multiple sample subsets; the multiple sample subsets include a first sample subset and at least one second sample subset;

[0166] The acquisition method of the sample images in the first sample subset is different from the acquisition method of the sample images in the second sample subset;

[0167] Moreover, the total number of sample images in the first sample subset is greater than the total number of sample images in each of the second sample subsets in the at least one second sample subset.

[0168] In a possible implementation manner, in the training method of the image detection model provided by the embodiments of the present application, in the loss function, the weight of the loss sub-function corresponding to the first sample subset is less than the weight of the loss sub-function corresponding to the at least one second sample subset.

[0169] In a possible implementation manner, in the training method of the image detection model provided by the embodiments of the present application, the multiple sample subsets include multiple second sample subsets, where the scenes to which the images in any one of the second sample subsets belong are different from the scenes to which the images in the multiple second sample subsets other than the any one of the second sample subsets belong.

[0170] In a possible implementation manner, in the training method of the image detection model provided by the embodiments of the present application, the loss function includes the loss sub-function corresponding to the first sample subset and the loss sub-functions corresponding to each of the second sample subsets in the second sample subset;

[0171] Among them, the weight of the loss sub-function corresponding to the first sample subset is less than the minimum value of the weights of the loss sub-functions corresponding to the multiple second sample subsets.

[0172] Based on the same technical concept, an embodiment of the present application provides a training device, which can achieve the same technical effect as the foregoing detection method, and will not be elaborated here. Refer to Figure 9 , the device includes an image acquisition module and an image detection module. Among them:

[0173] The image acquisition module is used to acquire the image to be measured;

[0174] The image detection module is used to, if the image to be measured includes a target object, extract the image features of the region including the target object; determine the similarity between the image features and the text features of each text in multiple texts, where the multiple texts are pre-generated, and the multiple texts correspond to multiple object categories one by one, and each text includes the corresponding object category; determine the object category included in the target text as the object category of the target object, where the similarity between the image features and the text features of the target text is the largest among the multiple similarities corresponding to the determined image features.

[0175] In a possible implementation manner, the number of characters of each text is greater than a preset character number threshold.

[0176] In a possible implementation manner, the image to be measured is collected by an image acquisition device in a kitchen scene;

[0177] The multiple object categories include:

[0178] Wearing a chef's hat and a mask, wearing a chef's hat and not wearing a mask, not wearing a chef's hat and wearing a mask, not wearing a chef's hat and not wearing a mask.

[0179] Based on the same technical concept, an embodiment of the present application provides a first electronic device, which can execute the training method provided in the above embodiment and can achieve the same technical effect, and will not be elaborated here.

[0180] Refer to Figure 10 , the first electronic device includes a processor 1001, a memory 1002 and a communication interface 1003. The processor 1001, the memory 1002 and the communication interface are connected through a bus 1004. The communication interface 1003 is used to communicate with other electronic devices, including but not limited to interacting with sample images. The memory 1002 stores a computer program, and the processor 1001 executes the steps in the training method in the above embodiment according to the computer program.

[0181] Embodiments of the present application Figure 10The processor involved may be a central processing unit (CPU), a general-purpose processor, a graphics processing unit (GPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof.

[0182] Based on the same technical concept, an embodiment of this application provides a second electronic device. This second electronic device can execute the training method provided in the above embodiment and can achieve the same technical effects, which will not be elaborated here.

[0183] See Figure 11 , this second electronic device includes a processor 1101, a memory 1102, and a communication interface 1103. The processor 1101, the memory 1102, and the communication interface are connected through a bus 1104. The communication interface 1103 is used to communicate with other electronic devices, including but not limited to interacting with the image to be measured. The memory 1102 stores a computer program, and the processor 1101 executes the steps in the detection method in the above embodiment according to the computer program.

[0184] Optionally, the second electronic device is the same device as the foregoing first electronic device. Or, the second electronic device is not the same device as the first electronic device, and the second electronic device may store the image detection model trained by the first electronic device. The second electronic device can interact with the first electronic device for the relevant parameters of the image detection model.

[0185] Embodiments of this application Figure 11 The processor involved may be a central processing unit (CPU), a general-purpose processor, a graphics processing unit (GPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof.

[0186] In addition, the present application provides a computer-readable storage medium storing computer instructions, which, when run on a computer, cause the computer to execute the training method or the image detection method provided by any embodiment of the present application.

[0187] An embodiment of the present application provides a computer program product including a computer program, which, when executed by a computer, implements the steps of the training method or the image detection method provided by any of the above embodiments.

[0188] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0189] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the specified functions in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0190] These computer program instructions can also be stored in a computer-readable memory capable of guiding a computer or other programmable data processing devices to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the specified functions in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0191] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, such that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable devices provide means for implementing the specified functions in Figure 1 one or more flows and / or blocks Figure 1Steps of functions specified in one or more boxes.

[0192] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these changes and modifications.

Claims

1. A training method for an image detection model, characterized in that, the method includes: Obtain a training sample set, the training sample set includes sample images of multiple categories, one sample image corresponds to one category, and the label of the sample image is the category corresponding to the sample image; Generate multiple texts according to the multiple categories, and extract the text features of each text, wherein the multiple texts correspond to the multiple categories one by one, and each text includes the category corresponding to each text; For each image in the training sample set, perform the following operations to determine the loss parameter of each image: input each image into the network model to be trained, and extract the first image feature of the area containing the target object in each image; determine the similarity between the first image feature and the text features of each text in the multiple texts; and determine the loss parameter of each image based on the multiple similarities corresponding to the determined first image feature; Adjust the parameters of the network model to be trained based on the loss parameters of all images in the training sample set and a preset loss function; wherein, the training sample set includes: multiple sample subsets; the multiple sample subsets include a first sample subset and at least one second sample subset; The acquisition method of the sample images in the first sample subset is different from the acquisition method of the sample images in the second sample subset; And, the total number of sample images in the first sample subset is greater than the total number of sample images in each of the at least one second sample subset; wherein, in the loss function, the weight of the loss sub-function corresponding to the first sample subset is less than the weight of the loss sub-function corresponding to the at least one second sample subset.

2. The method according to claim 1, characterized in that, the generating multiple texts according to the multiple categories includes: Generating each text according to at least one expansion field and the category corresponding to each text, wherein the expansion field includes one or more of the following: The acquisition method of the images in the training sample set, the scene to which the images in the training sample set belong.

3. The method according to claim 1, characterized in that, the multiple sample subsets include multiple second sample subsets, wherein the scene to which the images in any one second sample subset belong is different from the scene to which the images in the multiple second sample subsets other than the any one second sample subset belong.

4. The method according to claim 3, characterized in that, the loss function includes the loss sub-function corresponding to the first sample subset and the loss sub-functions corresponding to each of the second sample subsets in the second sample subset; wherein, the weight of the loss sub-function corresponding to the first sample subset is less than the minimum value of the weights of the loss sub-functions corresponding to the multiple second sample subsets.

5. An image detection method, characterized in that, the method includes: Obtain the image to be detected; Use the pre-trained image detection model to perform the following operations, wherein the image detection model is trained by using the training method described in any one of claims 1-4: If the image to be measured includes a target object, extract the image features of the region including the target object; Determine the similarity between the image features and the text features of each text in multiple texts, where the multiple texts are pre-generated, and the multiple texts correspond one-to-one with multiple object categories, and each text includes the corresponding object category; Determine the object category included in the target text as the object category of the target object, where the similarity between the image features and the text features of the target text is the largest among the multiple similarities corresponding to the determined image features.

6. The method according to claim 5, wherein, The number of characters in each text is greater than a preset character number threshold.

7. The method according to claim 5 or 6, wherein, The image to be measured is collected by an image acquisition device in a kitchen scenario; The multiple object categories include: Wearing a chef's hat and a mask, wearing a chef's hat and not wearing a mask, not wearing a chef's hat and wearing a mask, not wearing a chef's hat and not wearing a mask.

8. An electronic device, wherein, Comprising a memory and a processor; The memory is used to store program instructions; The processor is configured to execute the program instructions to implement the method according to any one of claims 1-7.

9. A computer-readable storage medium, wherein, Comprising computer program instructions, when the computer program instructions are executed by a computer, the computer executes the method according to any one of claims 1-7.

10. A training device, wherein, Comprising: An acquisition sample set module, configured to acquire a training sample set, the training sample set includes sample images of multiple categories, one sample image corresponds to one category, and the label of the sample image is the category corresponding to the sample image; A text feature generation module, configured to generate multiple texts according to the multiple categories, and extract the text features of each text, where the multiple texts correspond one-to-one with the multiple categories, and each text includes the category corresponding to each text; A model training module, configured to perform the following operations on each image in the training sample set: determine the loss parameter of each image: input each image into a network model to be trained, and extract the first image features of the region including the target object in each image; determine the similarity between the first image features and the text features of each text in the multiple texts; and determine the loss parameter of each image based on the multiple similarities corresponding to the determined first image features; and adjust the parameters of the network model to be trained based on the loss parameters of all the images in the training sample set and a preset loss function; where the training sample set includes: multiple sample subsets; the multiple sample subsets include a first sample subset and at least one second sample subset; The acquisition method of the sample images in the first sample subset is different from the acquisition method of the sample images in the second sample subset; And, the total number of sample images in the first sample subset is greater than the total number of sample images in each of the at least one second sample subset; Among them, in the loss function, the weight of the loss sub-function corresponding to the first sample subset is less than the weight of the loss sub-function corresponding to the at least one second sample subset.

11. An image detection device, characterized in that it includes: an image acquisition module for acquiring an image to be detected; an image detection module for, if the image to be detected includes a target object, extracting the image features of the region including the target object; determining the similarity between the image features and the text features of each text in a plurality of texts, where the plurality of texts are pre-generated and the plurality of texts correspond one-to-one to a plurality of object categories, and each text includes the corresponding object category; determining the object category included in the target text as the object category of the target object, where the similarity between the image features and the text features of the target text is the largest among the plurality of similarities corresponding to the determined image features; wherein the training sample set includes: a plurality of sample subsets; the plurality of sample subsets include a first sample subset and at least one second sample subset; the acquisition method of the sample images in the first sample subset is different from the acquisition method of the sample images in the second sample subset; and, the total number of the sample images in the first sample subset is greater than the total number of the sample images in each of the at least one second sample subsets; wherein, in the loss function, the weight of the loss sub-function corresponding to the first sample subset is less than the weight of the loss sub-function corresponding to the at least one second sample subset.

Citation Information

Patent Citations

  • Detection model training method and device, detection model classification method and device and electronic equipment

    CN116129224A

  • Data processing method, image-text retrieval method, image classification method and related equipment

    CN116226688A