Model training method and device, equipment, storage medium and vehicle

By obtaining the most generalized category of the literary graph model and replacing the training text, the problem of high training cost and low generalization ability in the existing technology is solved, and the effect of improving the generalization ability of the literary graph model while reducing training cost is achieved.

CN120180112APending Publication Date: 2025-06-20BEIJING CO WHEELS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311743731.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-18
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

When the prior art improves the generalization ability of literary and graphic models, the training cost is high and the generalization ability is still low.

Method used

By obtaining the most generalized category of the target object, replacing the target category in the original description text, using the replaced description text and target image to train the initial literary graph model to generate an image corresponding to the target object.

Benefits of technology

While improving the generalization ability of the literary graph model, the training cost of the model is reduced, and the generalization ability of the most generalized category is inherited rather than adding training samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180112A_ABST
    Figure CN120180112A_ABST
Patent Text Reader

Abstract

The invention discloses a model training method and device, equipment, a storage medium and a vehicle. The method comprises the steps that a first image corresponding to a target object and a first description text corresponding to the first image are acquired, and the first description text comprises a target category corresponding to the target object; the most generalized category corresponding to the target object is obtained, the most generalized category is the category with the best image generation effect for the second description text in m categories corresponding to the target object, and the second description text is the description text irrelevant to existing scenes of the m categories; replacing the target category in the first description text with the most generalized category to obtain a third description text; and training the initial text graph model by using the first image and the third description text to obtain a first text graph model corresponding to the target object. According to the model training method provided by the embodiment of the invention, the generalization ability of the first text graph model corresponding to the target object can be improved, and meanwhile, the training cost of the model is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of computer vision technology, and particularly relates to a model training method, device, equipment, storage medium and vehicle. Background Art

[0002] With the development of Artificial Intelligence (AI), the application of AI drawing is becoming more and more widespread. The Diffusion Model is a pre-trained model for AI drawing. The Diffusion Model has achieved remarkable success in text-to-image generation and can generate high-quality images according to the description text. In order to enable the Diffusion Model to adapt to specific tasks or fields, it is often necessary to fine-tune the Diffusion Model to obtain a text-to-image model corresponding to the specific task or field. In addition, the generalization ability refers to the performance of a machine learning model on unseen new data. The better the generalization ability of the text-to-image model, the higher the quality of the generated images.

[0003] Currently, it is usually to train the text-to-image model with more training samples corresponding to the target object to improve the generalization ability of the text-to-image model corresponding to the target object. However, more training samples increase the training cost of the model, and the generalization ability of the text-to-image model trained by the above method is still low. Summary of the Invention

[0004] Embodiments of this application provide a model training method, device, equipment, storage medium and vehicle, which can improve the generalization ability of the first text-to-image model corresponding to the target object while reducing the training cost of the model.

[0005] In a first aspect, embodiments of this application provide a model training method, and the method includes:

[0006] Obtain a first image corresponding to the target object and a first description text corresponding to the first image, where the first description text includes a target category corresponding to the target object;

[0007] Obtain the most generalized category corresponding to the target object, where the most generalized category is the category corresponding to the target image among the m categories corresponding to the target object, the target image is an image generated based on a second description text and having a similarity greater than a preset threshold with the second description text, the second description text is a description text irrelevant to the existence scenarios of the m categories, and the m categories include the target category, and m is a positive integer;

[0008] Replace the target category in the first description text with the most generalized category to obtain a third description text;

[0009] Train the initial text-to-image model using the first image and the third description text to obtain a first text-to-image model corresponding to the target object, where the first text-to-image model is used to generate an image corresponding to the description text of the target object.

[0010] In one possible implementation, the obtaining the most generalized category corresponding to the target object includes:

[0011] Obtain a fourth description text corresponding to the target category, where the fourth description text is a description text not related to the existence scenario of the target category;

[0012] Successively replace the target category in the fourth description text with each of the m categories to obtain fifth description texts respectively corresponding to the m categories;

[0013] For each of the fifth description texts, perform text-to-image processing on the fifth description text using the initial text-to-image model to obtain a second image;

[0014] Determine the similarity between each of the fifth description texts and its corresponding second image to obtain m first similarities;

[0015] Among the m first similarities, determine the category corresponding to the first similarity with the largest value as the most generalized category.

[0016] In one possible implementation, there are n fourth description texts corresponding to the target category, and there are n×m first similarities corresponding to the n fourth description texts, where n is a positive integer greater than 1; after obtaining the n×m first similarities, the method further includes:

[0017] Among the n×m first similarities, determine n first similarities respectively corresponding to each of the m categories;

[0018] For each of the m categories, perform statistics on the n first similarities to obtain a second similarity;

[0019] The determining the category corresponding to the first similarity with the largest value among the m first similarities as the most generalized category includes:

[0020] Among the m second similarities, determine the category corresponding to the second similarity with the largest value as the most generalized category.

[0021] In one possible implementation, obtaining the first description text corresponding to the first image includes:

[0022] Use the image-to-text model to describe the content of the first image to obtain a content description text;

[0023] Use the image-text matching model to perform image-text matching on the first image and a preset style text;

[0024] Determine the preset style text that matches the first image as the style description text of the first image;

[0025] Concatenate the content description text and the style description text to obtain the first description text.

[0026] In a possible implementation manner, when there are multiple target objects, before replacing the target category in the first description text with the most generalized category, the method further includes:

[0027] Obtain an object identifier corresponding to each target object;

[0028] Concatenate the object identifier corresponding to each target object and the most generalized category to obtain the most generalized word corresponding to each target object;

[0029] The step of replacing the target category in the first description text with the most generalized category to obtain a third description text includes:

[0030] For each target object, replace the target category in the first description text with the most generalized word to obtain the third description text corresponding to each target object.

[0031] In a possible implementation manner, the step of obtaining the first image corresponding to the target object includes:

[0032] Obtain the first image corresponding to each target object;

[0033] The step of training the initial text-to-image model with the first image and the third description text to obtain a first text-to-image model corresponding to the target object includes:

[0034] Train the initial text-to-image model with the first image corresponding to each target object and the third description sample to obtain first text-to-image models corresponding to multiple target objects respectively;

[0035] Determine the multiple first text-to-image models as second text-to-image models corresponding to multiple target objects, where the second text-to-image models are used to generate images corresponding to the description texts of multiple target objects respectively.

[0036] In a second aspect, an embodiment of the present application provides a model training device, and the device includes:

[0037] A first acquisition module, configured to acquire a first image corresponding to a target object and a first description text corresponding to the first image, where the first description text includes a target category corresponding to the target object;

[0038] A second acquisition module, configured to acquire a most generalized category corresponding to the target object, where the most generalized category is a category corresponding to a target image among m categories corresponding to the target object, the target image is an image generated based on a second description text and having a similarity greater than a preset threshold with the second description text, the second description text is a description text not related to the existence scenarios of the m categories, the m categories include the target category, and m is a positive integer;

[0039] A replacement module, configured to replace the target category in the first description text with the most generalized category to obtain a third description text;

[0040] A training module, configured to train an initial text-to-image model by using the first image and the third description text to obtain a first text-to-image model corresponding to the target object, where the first text-to-image model is used to generate an image corresponding to the description text of the target object.

[0041] In a third aspect, an embodiment of the present application provides an electronic device, which includes: a processor and a memory storing computer program instructions;

[0042] When the processor executes the computer program instructions, the method in any possible implementation method in the first aspect above is implemented.

[0043] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method in any possible implementation method in the first aspect above is implemented.

[0044] In a fifth aspect, an embodiment of the present application provides a vehicle, which includes at least one of the following:

[0045] A model training device in any embodiment of the second aspect;

[0046] An electronic device in any embodiment of the third aspect;

[0047] A computer-readable storage medium in any embodiment of the fourth aspect.

[0048] In the model training method, apparatus, device, storage medium, and vehicle according to the embodiments of the present application, since the second description text is a description text that is not relevant to the existence scenarios of m categories, there may be no images corresponding to the second description text in real life. Based on this, since the most generalized category is the category corresponding to the target image, and the target image is an image generated based on the second description text and having a similarity greater than a preset threshold with the second description text, the most generalized category can be the category with the strongest generalization ability among the m categories. Therefore, by replacing the target category in the first description text with the most generalized category, a third description text is obtained, and then, based on the first image and the third description text, the initial text-to-image model is trained to obtain a first text-to-image model corresponding to the target object, enabling the first text-to-image model to inherit the generalization ability of the most generalized category to a certain extent, and thus improving the generalization ability of the first text-to-image model corresponding to the target object. Since the first text-to-image model improves its own generalization ability by inheriting the generalization ability of the most generalized category rather than by using more training samples, the training cost of the model can be reduced. In this way, through the embodiments of the present application, it is possible to improve the generalization ability of the first text-to-image model corresponding to the target object while reducing the training cost of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] To more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required to be used in the embodiments of the present application. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0050] Figure 1 is a flowchart of a model training method provided by an embodiment of the present application;

[0051] Figure 2 is a flowchart of another model training method provided by an embodiment of the present application;

[0052] Figure 3 is a structural diagram of a model training apparatus provided by an embodiment of the present application;

[0053] Figure 4 is a structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0054] In order to better understand the above objects, features, and advantages of the present application, the following further describes the solutions of the present application. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other.

[0055] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present application. However, the present application may be practiced in other ways than those described herein. Obviously, the embodiments in the specification are only a part of the embodiments of the present application, rather than all of the embodiments.

[0056] It should be noted that, in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements that are inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.

[0057] As described in the background art section, in order to solve the problems of the prior art, embodiments of the present application provide a model training method, apparatus, device, storage medium, and vehicle.

[0058] First, the model training method provided by the embodiments of the present application will be introduced below.

[0059] Figure 1 The flowchart of a model training method provided by an embodiment of the present application is shown. This method can be executed by any processor or server including a model training module, which is not limited herein. As Figure 1 shown, the model training method provided by the embodiments of the present application includes the following steps:

[0060] S110. Obtain a first image corresponding to a target object and a first description text corresponding to the first image, where the first description text includes a target category corresponding to the target object;

[0061] S120. Obtain the most generalized category corresponding to the target object, where the most generalized category is the category corresponding to the target object among m categories corresponding to the target object and corresponding to a target image, the target image is an image generated based on a second description text and having a similarity greater than a preset threshold with the second description text, the second description text is a description text not related to the existence scenarios of the m categories, the m categories include the target category, and m is a positive integer;

[0062] S130. Replace the target category in the first description text with the most generalized category to obtain a third description text;

[0063] S140. Use the first image and the third description text to train the initial text-to-image model to obtain a first text-to-image model corresponding to the target object. The first text-to-image model is used to generate an image corresponding to the description text of the target object.

[0064] In the model training method of the embodiments of the present application, since the second description text is a description text not related to the existence scenarios of m categories, there may be no image corresponding to the second description text in real life. Based on this, since the most generalized category is the category corresponding to the target image, and the target image is generated based on the second description text and has a similarity greater than the preset threshold with the second description text, the most generalized category can be the category with the strongest generalization ability among the m categories. Therefore, by replacing the target category in the first description text with the most generalized category, the third description text is obtained, and then the initial text-to-image model is trained based on the first image and the third description text to obtain a first text-to-image model corresponding to the target object, enabling the first text-to-image model to inherit the generalization ability of the most generalized category to a certain extent, and thus improving the generalization ability of the first text-to-image model corresponding to the target object. Since the first text-to-image model improves its own generalization ability by inheriting the generalization ability of the most generalized category rather than by more training samples, the training cost of the model can be reduced. In this way, through the embodiments of the present application, while improving the generalization ability of the first text-to-image model corresponding to the target object, the training cost of the model can be reduced.

[0065] The specific implementation manners of the above steps are introduced below.

[0066] In some embodiments, in S110, the target object can be any one of a person, an animal, a plant, and an object. When the target object is an object, the target object can be specific to the manufacturer and model of the object. When the target object is an animal or a plant, the target object can be specific to the variety of the animal or plant. In addition, the first image corresponding to the target object can be an image with the target object as the main body. For example, if image A includes object A and object B, and the space occupied by object A in image A is larger and the space occupied by object B is smaller, then image A has object A as the main body. Therefore, when the target object is object A, the first image corresponding to the target object can include image A.

[0067] As an example, obtaining the first image corresponding to the target object may include: first obtaining multiple third images including the target object, and then screening the multiple third images to obtain the first image with the target object as the main body. Among them, the method of image screening can be manual screening or screening through a model. The steps of screening through a model may, for example, include: first performing object detection on the third image to determine the detection box corresponding to each object in the third image; if there is one object in the third image, then this object may be the target object, and this third image may be the first image corresponding to the target object; if there are multiple objects in the third image, then the perimeters of the detection boxes corresponding to the multiple objects can be determined, and the object corresponding to the detection box with the largest perimeter can be determined as the main body of the third image. If the object corresponding to the detection box with the largest perimeter is the target object, then this third image may be the first image corresponding to the target object.

[0068] In addition, the first image may be an image with a short side size not less than 768px and a clear main body. Multiple first images may be obtained, and each first image may correspond to a first description text. Specifically, the number of first images may, for example, be not less than 20.

[0069] In addition, the first description text may be a description of the first image. The subject in the first description text may correspond to the main body in the first image. Among them, the subject in the first description text may, for example, be the target category corresponding to the target object. If the target object is vehicle A, then the target category may, for example, be any one of car, electric vehicle, motor vehicle, off-road vehicle, commercial vehicle, and sports utility vehicle.

[0070] In addition, the first description text may only include the content description text, or may also include the content description text and the style description text at the same time. Among them, the content description text may be a text obtained by describing the content of the first image. The style description text may be a text obtained by describing the style of the first image. For example, if the first image is an image taken of vehicle A on the road, then the content description text of the first image may, for example, be "A car is running on the road", and the style description text of the first image may, for example, be "realistic style". Based on this, the first description text may be "A car is running on the road", or may also be "A car is running on the road, realistic style".

[0071] As an example, after obtaining the first image, the first description text (manually tagging the first image) can be obtained by manually performing content description and style description on the first image. Manual tagging can ensure the accuracy of the first description text, but the efficiency is low.

[0072] Based on this, in order to improve the efficiency of determining the first description text, in some embodiments, the above-mentioned obtaining the first description text corresponding to the first image may specifically include:

[0073] Use an image-to-text model to describe the content of the first image to obtain a content description text;

[0074] Determine the content description text as the first description text.

[0075] Here, the image-to-text model can be a neural network model that can describe the content of an image and generate a content description text. The image-to-text model can be, for example, the Bootstrapping Language-Image Pre-training (BLIP) model.

[0076] As an example, after obtaining the first image, the first image can be input into the BLIP model. The BLIP model can describe the content of the first image and obtain and output a content description text. The content description text can be output in the form of a string. If the content description text is denoted as string A, taking the example that the first image is an image taken of vehicle A on the road, it can be A = "A car is running on the road". In this example, if the first description text is denoted as string C, it can also be C = "A car is running on the road".

[0077] As an example, after obtaining multiple first images, the multiple first images can be input into the BLIP model. The BLIP model can batch process the multiple first images to obtain the content description texts corresponding to each of the multiple first images.

[0078] In this way, by using the image-to-text model to automatically tag the first image, compared with manual tagging, the efficiency of determining the first description text can be improved.

[0079] Based on this, in order to enrich the first description text, in some embodiments, the above-mentioned obtaining the first description text corresponding to the first image may specifically further include:

[0080] Use an image-to-text model to describe the content of the first image to obtain a content description text;

[0081] Use an image-text matching model to perform image-text matching on the first image and a preset style text;

[0082] Determine the preset style text that matches the first image as the style description text of the first image;

[0083] Concatenate the content description text and the style description text to obtain the first description text.

[0084] Here, the image-text matching model can be a neural network model capable of matching images and texts and determining the similarity between the two. The image-text matching model can be, for example, a Contrastive Language–Image Pre-training (CLIP) model. Additionally, the preset style texts can include, for example, "realistic style", "virtual animation style", "oil painting style", "sketch style", etc.

[0085] As an example, after obtaining the first image, the first image and multiple preset style texts can be jointly input into the CLIP model. The CLIP model can calculate the similarity between the first image and each preset style text to obtain multiple similarities. The preset style text corresponding to the maximum similarity value among the multiple similarities is determined as the style description text of the first image. Among them, the manifestation form of the style description text can be a string. If the style description text is denoted as string B, taking the first image as an image obtained by taking a picture of vehicle A on the road as an example, it can be B = "realistic style".

[0086] As an example, after obtaining string A and string B, for example, python can be used to splice string A and string B to obtain string C. After obtaining string C, the content in string C can also be manually corrected to obtain the final first description text.

[0087] In this way, by adding the style description text to the first description text, the first description text can be enriched, and further, by performing model training based on the first description text subsequently, the quality of the images generated by the model can be improved.

[0088] In some embodiments, in S120, the second description text can be a description text unrelated to the existence scenarios of the m categories. For example, all m categories belong to vehicles, and the existence scenario of the vehicles can be land. Since vehicles generally do not swim in water, the second description text can be, for example, "A car is speeding at the bottom of the water, surrounded by a large number of fish". In other words, there may be no image corresponding to the second description text in real life. That is to say, the second description text can be a description text constructed artificially, rather than a description text obtained by labeling existing images.

[0089] Additionally, the corresponding relationship between the target object and the m categories can be set in advance. If the target object is vehicle A, the m categories can include, for example, cars, electric vehicles, motor vehicles, off-road vehicles, commercial vehicles, sport utility vehicles, etc.

[0090] As an example, the second description text may include categories. If the m categories include cars and off-road vehicles, the second description text may, for example, include "A car is speeding underwater, surrounded by a large number of fish" and "An off-road vehicle is speeding underwater, surrounded by a large number of fish". If text-to-image processing is performed on the two second description texts respectively, and the image generated by the former is better than the image generated by the latter, the most generalized category corresponding to vehicle A can be "car" rather than "off-road vehicle". Among them, a better image effect can mean a higher reduction of the image to the description text. Whether the image effect is good or bad can be determined by the user or by the model, which is not limited here. If the image effect is determined by the model, the target image with the best image effect can be an image generated based on the second description text and having a similarity greater than a preset threshold to the second description text. Among them, the preset threshold can be a similarity threshold preset for determining the target image. That is, if the similarity between the image and the second description text is greater than the preset threshold, the image can be the target image.

[0091] Based on this, in some embodiments, the above S120 may specifically include:

[0092] Obtain a fourth description text corresponding to the target category, where the fourth description text is a description text not related to the existence scenario of the target category;

[0093] Successively replace the target category in the fourth description text with each of the m categories to obtain fifth description texts corresponding to the m categories respectively;

[0094] For each fifth description text, use the initial text-to-image model to perform text-to-image processing on the fifth description text respectively to obtain a second image;

[0095] Determine the similarity between each fifth description text and its corresponding second image to obtain m first similarities;

[0096] Among the m first similarities, determine the category corresponding to the first similarity with the largest value as the most generalized category.

[0097] Here, the fourth description text can be description text that is not relevant to the existence scenario of the target category. The fourth description text can be constructed manually. For the introduction of the fourth description text, reference can be made to the relevant introduction of the second description text above, and details will not be elaborated here. For example, if m = 3, the target category is cars, and the fourth description text is "A car is speeding underwater, surrounded by a large number of fish", and the 3 categories include cars, electric cars, and off-road cars, then after replacing "car" in the fourth description text with "car", the fifth description text "A car is speeding underwater, surrounded by a large number of fish" can be obtained. After replacing "car" in the fourth description text with "electric car", the fifth description text "An electric car is speeding underwater, surrounded by a large number of fish" can be obtained. After replacing "car" in the fourth description text with "off-road car", the fifth description text "An off-road car is speeding underwater, surrounded by a large number of fish" can be obtained. That is, the 3 fifth description texts can include "A car is speeding underwater, surrounded by a large number of fish", "An electric car is speeding underwater, surrounded by a large number of fish", and "An off-road car is speeding underwater, surrounded by a large number of fish".

[0098] As an example, the initial text-to-image model can be a pre-trained model for text-to-image. The initial text-to-image model can, for example, include a diffusion model (Diffussion Model) and a stable diffusion model (Stable Diffusion Model). After inputting the fifth description text into the initial text-to-image model, the initial text-to-image model can generate a second image corresponding to the fifth description text. After obtaining the second image, the similarity between the fifth description text and its corresponding second image can be determined manually, or the similarity between the fifth description text and its corresponding second image can be determined through a text-image matching model, and the first similarity corresponding to each of the m fifth description texts can be obtained. Among them, the larger the value of the first similarity, the better the reduction of the second image to the fifth description text, and the stronger the generalization ability of the category corresponding to the fifth description text. Furthermore, the category corresponding to the fifth description text can be determined as the most generalized category.

[0099] As an example, among the m first similarities, determining the category corresponding to the first similarity with the largest value as the most generalized category can specifically include:

[0100] Sort the m similarities in descending order to obtain a first similarity sequence;

[0101] Determine the category corresponding to the first similarity in the first similarity sequence as the most generalized category.

[0102] As another example, among the m first similarities, determining the category corresponding to the first similarity with the largest value as the most generalized category can specifically include:

[0103] Sort the m similarities in ascending order to obtain a second similarity sequence;

[0104] Determine the category corresponding to the last similarity in the second similarity sequence as the most generalized category.

[0105] In addition, the fourth description text corresponding to the target category can be one or more. Among them, the more the number of the fourth description texts, the more the number of the fifth description texts, and the more the number of the first similarities. Based on this, by determining the most generalized category according to more first similarities, the reliability of the most generalized category can be improved.

[0106] As an example, if the fourth description text corresponding to the target category is n (n≥2), the number of the first similarities corresponding to the n fourth description texts can be n×m. For example, assume n = 2, m = 3, the target category is a car, and the 2 fourth description texts include "a car swimming in the water" and "a car flying in the sky", and the 3 categories include cars, electric cars, and off-road cars. Then for the first fourth description text, by sequentially replacing the car in the fourth description text with a car, an electric car, and an off-road car, 3 fifth description texts "a car swimming in the water", "an electric car swimming in the water", and "an off-road car swimming in the water" can be obtained. For the second fourth description text, by sequentially replacing the car in the fourth description text with a car, an electric car, and an off-road car, 3 fifth description texts "a car flying in the sky", "an electric car flying in the sky", and "an off-road car flying in the sky" can be obtained. In this way, a total of 6 (2×3) fifth description texts can be obtained. For each fifth description text, generate a second image and determine the first similarity between the fifth description text and the second image, and 6 (2×3) first similarities can be obtained.

[0107] Based on this, in order to improve the reliability of the most generalized category, in some embodiments, after obtaining n×m first similarities, it may further include:

[0108] Among the n×m first similarities, determine the n first similarities corresponding to each of the m categories respectively;

[0109] For each of the m categories, count the n first similarities to obtain a second similarity.

[0110] Based on this, the above-mentioned determining the category corresponding to the first similarity with the largest value among the m first similarities as the most generalized category may specifically include:

[0111] Among the m second similarities, determine the category corresponding to the second similarity with the largest value as the most generalized category.

[0112] Here, each category may correspond to n first similarities. By summing the n first similarities or calculating the average value of the n first similarities, a second similarity can be obtained. Each category may correspond to one second similarity.

[0113] In this way, by determining the category corresponding to the largest second similarity value among the m second similarities as the most generalized category, the reliability of the most generalized category can be improved.

[0114] In some embodiments, in S130, after obtaining the most generalized category, the target category in the first description text can be directly replaced with the most generalized category. Here, the target category and the most generalized category may be the same or different. When the target category is the same as the most generalized category, the third description text may be the same as the first description text. Additionally, after obtaining the most generalized category, the target category and the most generalized category can be compared first. When the target category is different from the most generalized category, the target category in the first description text can be replaced with the most generalized category. When the target category is the same as the most generalized category, the first description text can be determined as the third description text.

[0115] As an example, after determining the most generalized category, the find() method in python can be used to check whether the first description text contains a category description corresponding to a non - most - generalized category. The non - most - generalized category can be any one of the m categories other than the most generalized category. For example, if the m categories include cars and off - road vehicles, and the off - road vehicle is the most generalized category, then the car can be determined as the non - most - generalized category. When the first description text includes the non - most - generalized category, the replace() method in python can be used to replace the non - most - generalized category with the most generalized category. For example, if the first description text is "A car is speeding on the road, realistic style", then the third description text can be "An off - road vehicle is speeding on the road, realistic style".

[0116] After obtaining the third description text, the third description text can be determined as the image label of the first image, and the multiple first images and their respective corresponding third description samples can be determined as the training samples for subsequent model training.

[0117] In some embodiments, in S140, the first text - to - image model can perform text - to - image processing on the description text corresponding to the target object to generate a third image with the target object as the main body. By training the initial text - to - image model using the first image and the third description text, the first text - to - image model can improve both the generalization ability of generating the target object and the restoration ability of the target object.

[0118] In the embodiments of the present application, the initial text-to-image model can be, for example, the StableDiffusion model. Based on this, model training can be performed based on the Low-Rank Adaptation of Large Language Models (LoRA) fine-tuning algorithm. The principle of the LoRA fine-tuning algorithm can be: by modifying the cross-attention layer in the UNet network architecture of the StableDiffusion model to reduce the model size, it is possible to ensure both the training performance of the model and adjust the image styles generated by the diffusion model. Specifically, modifying the cross-attention layer in the UNet network architecture of the StableDiffusion model can include: training the model parameters (low-rank matrices B and A) of the LoRA model, and adjusting the initial parameters W0 of the StableDiffusion model through the formula W = W0 + BA to obtain the final parameters W of the StableDiffusion model. Among them, if W0 is a p×q matrix, the size of matrix A is p×k, and the size of matrix B is k×q, where k is much smaller than p. During the process of model training based on the LoRA fine-tuning algorithm, W0 can be frozen and not updated with gradients, but the trainable parameters included in A and B are adjusted.

[0119] As an example, the specific process of training the StableDiffusion model using a training sample set (including multiple first images and their respective corresponding third description texts) can include:

[0120] Input the training samples into the StableDiffusion model, and use the StableDiffusion model to perform text-to-image processing on each third description text in the training samples to obtain predicted images;

[0121] In the case where the similarity between the predicted image and the first image is less than a preset threshold, adjust the model parameters in the manner of W = W0 + AB, and repeat the above steps until the similarity between the predicted image and the first image is greater than the preset threshold, to obtain the first text-to-image model corresponding to the target object. The first text-to-image model can not only generate images corresponding to the target object, but also has high generalization ability.

[0122] Based on the above embodiments, a first text-to-image model corresponding to a single target object can be generated. If there are multiple target objects, a second text-to-image model corresponding to multiple target objects can be generated.

[0123] Based on this, in order to improve the generalization ability of the second text-to-image model for multiple target objects, in some embodiments, before the above S130, it may further include:

[0124] Obtain an object identifier corresponding to each target object;

[0125] Concatenate the object identifier and the most generalized category corresponding to each target object to obtain the most generalized word corresponding to each target object.

[0126] Based on this, the above S130 may specifically include:

[0127] For each target object, replace the target category in the first description text with the most generalized word to obtain the third description text corresponding to each target object.

[0128] Here, in the case where there are multiple target objects, in order to distinguish different target objects, a unique object identifier can be set for each target object. Additionally, the m categories corresponding to multiple target objects can be the same, and thus the most generalized categories corresponding to multiple target objects can be the same. Therefore, in order to improve the generalization ability of the second text-to-image model corresponding to multiple target objects while distinguishing different target objects, for each target object, concatenate the object identifier and the most generalized category to obtain the most generalized word.

[0129] In this way, by replacing the target category in the first description text with the most generalized word to obtain the third description text, and then using the first image and the third description text corresponding to each target object for model training subsequently, the generalization ability of the second text-to-image model for multiple target objects can be improved.

[0130] Based on this, in the case where there are multiple target objects, in order to make the second text-to-image model have better reducibility and generalization ability for multiple target objects, in some embodiments, the above S110 may specifically include:

[0131] Obtain the first image corresponding to each target object;

[0132] Based on this, the above S140 may specifically include:

[0133] Use the first image and the third description sample corresponding to each target object to train the initial text-to-image model to obtain the first text-to-image models corresponding to multiple target objects respectively;

[0134] Determine the multiple first text-to-image models as the second text-to-image model corresponding to multiple target objects, and the second text-to-image model is used to generate images corresponding to the description texts of multiple target objects respectively.

[0135] Here, each target object may correspond to a training sample set. The training sample set may include multiple training samples. Each training sample may include a first image and its corresponding third description text. The number of samples in the training samples corresponding to different target objects may be the same or different, and this is not limited herein.

[0136] In this way, by determining multiple first text-to-image generation models as the second text-to-image generation models corresponding to multiple target objects, it can be ensured that the second text-to-image generation models can generate images corresponding to the description texts of multiple target objects respectively, so that the second text-to-image generation models have better reduction and generalization capabilities for multiple target objects.

[0137] To better describe the entire solution, based on the above embodiments, some specific examples are given.

[0138] For example, as Figure 2 shown, if there is one target object, to improve the generalization ability of the first text-to-image generation model, the training samples can be prepared through the following steps:

[0139] S21. Obtain the first image corresponding to the target object;

[0140] S22. Use the BLIP model to determine the content description text corresponding to the first image, and use the CLIP model to determine the style description text corresponding to the first image;

[0141] S23. Concatenate the content description text and the style description text to obtain the first description text;

[0142] S24. Obtain the most generalized category corresponding to the target object;

[0143] S25. Replace the target category corresponding to the target object in the first description text with the most generalized category to obtain the third description text;

[0144] S26. Determine the first image and its corresponding third description text as the training samples.

[0145] After obtaining multiple training samples in the above manner, the multiple training samples can be determined as a training sample set, and the initial text-to-image generation model can be trained using the training sample set to obtain the first text-to-image generation model.

[0146] Thus, by replacing the category in the description text, it is possible to perform a conceptual substitution of the target object, enabling the model to inherit the generalization ability of the most generalized category to a certain extent, making the target object generated by the first text-to-image generation model easier to combine with other concepts and improving the generalization ability of the first text-to-image generation model.

[0147] Based on the model training method provided in the above embodiments, correspondingly, the present application also provides a specific implementation manner of the model training device. Please refer to the following embodiments.

[0148] As Figure 3 shown, the model training device 300 provided in the embodiments of the present application includes the following modules:

[0149] The first acquisition module 310 is configured to acquire a first image corresponding to a target object and a first description text corresponding to the first image, where the first description text includes a target category corresponding to the target object;

[0150] The second acquisition module 320 is configured to acquire the most generalized category corresponding to the target object, where the most generalized category is the category corresponding to the target object among the m categories corresponding to the target object and corresponding to a target image, the target image is an image generated based on a second description text and having a similarity greater than a preset threshold with the second description text, the second description text is a description text not related to the existence scenarios of the m categories, the m categories include the target category, and m is a positive integer;

[0151] The replacement module 330 is configured to replace the target category in the first description text with the most generalized category to obtain a third description text;

[0152] The training module 340 is configured to train an initial text-to-image model using the first image and the third description text to obtain a first text-to-image model corresponding to the target object, and the first text-to-image model is used to generate an image corresponding to the description text of the target object.

[0153] The above model training apparatus 300 will be described in detail below, as follows:

[0154] In some embodiments, the second acquisition module 320 may specifically include:

[0155] The first acquisition sub-module is configured to acquire a fourth description text corresponding to the target category, where the fourth description text is a description text not related to the existence scenario of the target category;

[0156] The first replacement sub-module is configured to sequentially replace the target category in the fourth description text with each of the m categories to obtain fifth description texts corresponding to the m categories respectively;

[0157] The processing sub-module is configured to, for each fifth description text, perform text-to-image processing on the fifth description text using the initial text-to-image model to obtain a second image;

[0158] The first determination sub-module is configured to determine the similarity between each fifth description text and its corresponding second image to obtain m first similarities;

[0159] The second determination sub-module is configured to, among the m first similarities, determine the category corresponding to the first similarity with the largest value as the most generalized category.

[0160] In some embodiments, there are n fourth description texts corresponding to the target category, and the first similarities corresponding to the n fourth description texts are n×m. n is a positive integer greater than 1. Based on this, the second acquisition module 320 may specifically further include:

[0161] A third determination sub-module, configured to, after obtaining n×m first similarities, determine, from the n×m first similarities, n first similarities respectively corresponding to each of the m categories.

[0162] A first statistics sub-module, configured to, for each of the m categories, perform statistics on the n first similarities to obtain a second similarity.

[0163] Based on this, the second determination sub-module may specifically include:

[0164] A determination unit, configured to, among the m second similarities, determine the category corresponding to the second similarity with the largest value as the most generalized category.

[0165] In some embodiments, the first acquisition module 310 may specifically include:

[0166] A descriptor sub-module, configured to use a graph-to-text model to perform content description on the first image to obtain a content description text;

[0167] A matching sub-module, configured to use an image-text matching model to perform image-text matching on the first image and a preset style text;

[0168] A fourth determination sub-module, configured to determine the preset style text that matches the first image as the style description text of the first image;

[0169] A splicing sub-module, configured to splice the content description text and the style description text to obtain a first description text.

[0170] In some embodiments, there are multiple target objects. Based on this, the model training device 300 may further include:

[0171] A third acquisition module, configured to, before replacing the target category in the first description text with the most generalized category, acquire an object identifier corresponding to each target object;

[0172] A splicing module, configured to splice the object identifier corresponding to each target object and the most generalized category to obtain a most generalized word corresponding to each target object.

[0173] Based on this, the replacement module 330 may specifically include:

[0174] A second replacement sub-module, configured to, for each target object, replace the target category in the first description text with the most generalized word to obtain a third description text corresponding to each target object.

[0175] In some embodiments, the first acquisition module 310 may specifically include:

[0176] The second acquisition sub-module is configured to acquire the first image corresponding to each target object.

[0177] Based on this, the training module 340 may specifically include:

[0178] The training sub-module is configured to train the initial text-to-image model by using the first image corresponding to each target object and the third description sample, so as to obtain the first text-to-image models corresponding to multiple target objects respectively;

[0179] The fifth determination sub-module is configured to determine the multiple first text-to-image models as the second text-to-image models corresponding to multiple target objects, and the second text-to-image models are used to generate the images corresponding to the description texts of multiple target objects respectively.

[0180] In the model training device according to the embodiment of the present application, since the second description text is a description text that is not related to the existence scenarios of m categories, there may be no images corresponding to the second description text in real life. Based on this, since the most generalized category is the category corresponding to the target image, and the target image is an image generated based on the second description text and has a similarity greater than a preset threshold with the second description text, the most generalized category may be the category with the strongest generalization ability among the m categories. Therefore, by replacing the target category in the first description text with the most generalized category, the third description text is obtained, and then the initial text-to-image model is trained based on the first image and the third description text, so as to obtain the first text-to-image model corresponding to the target object, enabling the first text-to-image model to inherit the generalization ability of the most generalized category to a certain extent, and further improving the generalization ability of the first text-to-image model corresponding to the target object. Since the first text-to-image model improves its own generalization ability by inheriting the generalization ability of the most generalized category, rather than by using more training samples, the training cost of the model can be reduced. In this way, through the embodiment of the present application, the generalization ability of the first text-to-image model corresponding to the target object can be improved while reducing the training cost of the model.

[0181] Based on the model training method provided in the above embodiment, the embodiment of the present application also provides a specific implementation manner of an electronic device. Figure 4 The schematic diagram of the electronic device 400 provided by the embodiment of the present application is shown.

[0182] The electronic device 400 may include a processor 410 and a memory 420 storing computer program instructions.

[0183] Specifically, the above-mentioned processor 410 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured as one or more integrated circuits for implementing the embodiments of the present application.

[0184] The memory 420 may include a mass storage for data or instructions. By way of example and not limitation, the memory 420 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disc, a magneto-optical disc, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. In a suitable case, the memory 420 may include removable or non-removable (or fixed) media. In a suitable case, the memory 420 may be internal or external to the electronic device 400. In a specific embodiment, the memory 420 is a non-volatile solid-state memory.

[0185] The memory may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk storage media device, an optical storage media device, a flash memory device, an electrical, optical, or other physical / tangible memory storage device. Thus, generally, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to the first aspect of the present application.

[0186] The processor 410 reads and executes the computer program instructions stored in the memory 420 to implement any one of the model training methods in the above embodiments.

[0187] In one example, the electronic device 400 may further include a communication interface 430 and a bus 440. As shown, Figure 4 the processor 410, the memory 420, and the communication interface 430 are connected via the bus 440 and complete communication with each other.

[0188] The communication interface 430 is mainly used to implement communication between the various modules, devices, units, and / or devices in the embodiments of the present application.

[0189] The bus 440 includes hardware, software, or both, and couples components of the electronic device together. By way of example and not limitation, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or a combination of two or more of these. Where appropriate, the bus 440 may include one or more buses. Although embodiments of the present application describe and illustrate specific buses, the present application contemplates any suitable bus or interconnect.

[0190] Exemplarily, the electronic device 400 may be a mobile phone, a tablet computer, a laptop computer, a handheld computer, an in-vehicle electronic device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc.

[0191] The electronic device may execute the model training method in the embodiments of the present application, so as to implement the combination of Figures 1 to 3 the described model training method and apparatus.

[0192] In addition, in combination with the model training method in the above embodiments, embodiments of the present application may provide a computer-readable storage medium to implement. Computer program instructions are stored on the computer-readable storage medium; when the computer program instructions are executed by a processor, any one of the model training methods in the above embodiments is implemented.

[0193] In addition, embodiments of the present application further provide a vehicle, which may include at least one of the following:

[0194] A model training device as in any one of the embodiments of the second aspect;

[0195] An electronic device as in any one of the embodiments of the third aspect;

[0196] A computer-readable storage medium as in any one of the embodiments of the fourth aspect. Details will not be described herein again.

[0197] It should be clear that the present application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present application.

[0198] The functional blocks shown in the above-described structural block diagrams can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, and so on. When implemented in software, the elements of the present application are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted via a data signal carried in a carrier wave over a transmission medium or a communication link. A "machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical discs, hard disks, fiber optic media, radio frequency (RF) links, and so on. The code segment can be downloaded via a computer network such as the Internet, an intranet, and so on.

[0199] It should also be noted that the exemplary embodiments mentioned in the present application describe some methods or systems based on a series of steps or devices. However, the present application is not limited to the order of the above steps, that is, the steps can be executed in the order mentioned in the embodiments, or different from the order in the embodiments, or several steps can be executed simultaneously.

[0200] The various aspects of the present application have been described above with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each block in the flowchart and / or block diagram, and the combination of blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing devices to produce a machine such that the instructions executed by the processor of the computer or other programmable data processing devices enable the implementation of the functions / actions specified in one or more blocks of the flowchart and / or block diagram. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can also be implemented by dedicated hardware that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0201] As described above, this is only the specific implementation manner of the present application. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, modules, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein. It should be understood that the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present application.

Claims

1. A model training method, characterized in that, Including: Obtain a first image corresponding to a target object and a first description text corresponding to the first image, where the first description text includes a target category corresponding to the target object; Obtain the most generalized category corresponding to the target object, where the most generalized category is the category corresponding to the target object in m categories corresponding to the target object, the target image is an image generated based on a second description text and having a similarity greater than a preset threshold with the second description text, the second description text is a description text not related to the existence scenarios of the m categories, the m categories include the target category, and m is a positive integer; Replace the target category in the first description text with the most generalized category to obtain a third description text; Use the first image and the third description text to train an initial text-to-image model to obtain a first text-to-image model corresponding to the target object, where the first text-to-image model is used to generate an image corresponding to the description text of the target object.

2. The method according to claim 1, characterized in that, The obtaining the most generalized category corresponding to the target object includes: Obtain a fourth description text corresponding to the target category, where the fourth description text is a description text not related to the existence scenario of the target category; Successively replace the target category in the fourth description text with each of the m categories to obtain fifth description texts corresponding to the m categories respectively; For each of the fifth description texts, perform text-to-image processing on the fifth description text respectively using the initial text-to-image model to obtain a second image; Determine the similarity between each of the fifth description texts and its corresponding second image to obtain m first similarities; Among the m first similarities, determine the category corresponding to the first similarity with the largest value as the most generalized category.

3. The method according to claim 2, characterized in that, There are n fourth description texts corresponding to the target category, and there are n×m first similarities corresponding to the n fourth description texts, and n is a positive integer greater than 1; After obtaining the n×m first similarities, the method further includes: Among the n×m first similarities, determine n first similarities corresponding to each of the m categories respectively; For each of the m categories, perform statistics on the n first similarities to obtain a second similarity; The determining the category corresponding to the first similarity with the largest value among the m first similarities as the most generalized category includes: Among the m second similarities, determine the category corresponding to the second similarity with the largest value as the most generalized category.

4. The method according to claim 1, characterized in that, Obtaining the first description text corresponding to the first image includes: Use an image-to-text model to describe the content of the first image to obtain a content description text; Use an image-text matching model to perform image-text matching on the first image and a preset style text; Determine the preset style text matching the first image as the style description text of the first image; Concatenate the content description text and the style description text to obtain the first description text.

5. The method according to claim 1, characterized in that, The target objects are multiple. Before replacing the target category in the first description text with the most generalized category, the method further includes: Obtaining an object identifier corresponding to each of the target objects; Concatenating the object identifier corresponding to each of the target objects and the most generalized category to obtain a most generalized word corresponding to each of the target objects; The replacing the target category in the first description text with the most generalized category to obtain a third description text includes: For each of the target objects, replacing the target category in the first description text with the most generalized word to obtain the third description text corresponding to each of the target objects.

6. The method according to claim 5, characterized in that, The obtaining the first image corresponding to the target object includes: Obtaining the first image corresponding to each of the target objects; The training the initial text-to-image model with the first image and the third description text to obtain a first text-to-image model corresponding to the target object includes: Training the initial text-to-image model with the first image corresponding to each of the target objects and the third description sample to obtain first text-to-image models corresponding to the multiple target objects respectively; Determining the multiple first text-to-image models as second text-to-image models corresponding to the multiple target objects, where the second text-to-image models are used to generate images corresponding to the description texts of the multiple target objects respectively.

7. A model training device, characterized in that, The apparatus includes: A first obtaining module, configured to obtain a first image corresponding to a target object and a first description text corresponding to the first image, where the first description text includes a target category corresponding to the target object; A second obtaining module, configured to obtain a most generalized category corresponding to the target object, where the most generalized category is the category corresponding to the target object among the m categories corresponding to the target object and corresponding to a target image, the target image is an image generated based on a second description text and having a similarity greater than a preset threshold with the second description text, the second description text is a description text not related to the existence scenarios of the m categories, the m categories include the target category, and m is a positive integer; A replacing module, configured to replace the target category in the first description text with the most generalized category to obtain a third description text; A training module, configured to train an initial text-to-image model with the first image and the third description text to obtain a first text-to-image model corresponding to the target object, where the first text-to-image model is used to generate an image corresponding to the description text of the target object.

8. An electronic device, characterized in that, The electronic device includes: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, the model training method according to any one of claims 1-6 is implemented.

9. A computer-readable storage medium, characterized in that, Computer program instructions are stored on a computer-readable storage medium, and when the computer program instructions are executed by a processor, the model training method according to any one of claims 1-6 is implemented.

10. A vehicle, characterized in that, Including at least one of the following: The model training apparatus according to claim 7; The electronic device according to claim 8; The computer-readable storage medium according to claim 9.