Model training method and device, picture generation method and device, equipment and storage medium

By generating multiple pictures, determining sample pictures with high object similarity, and constructing training samples to train the LoRA model, the problem of inconsistent objects in multiple pictures generated in the existing technology is solved, and the consistency of objects in the generated pictures is achieved.

CN120672881APending Publication Date: 2025-09-19BEIJING QIYI CENTURY SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510619969.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

It is difficult with existing technologies to ensure the consistency of objects contained in multiple images generated based on the same text prompt word.

Method used

By generating multiple pictures, determining sample pictures whose object similarity is higher than a preset threshold, obtaining sample labels of the sample pictures, and constructing training samples based on these sample pictures and labels to train the LoRA model and obtain the target LoRA model.

Benefits of technology

Improves the consistency of objects contained in multiple images generated by the same text prompt word, ensuring that the similarity of objects in the generated images is higher than a preset threshold.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120672881A_ABST
    Figure CN120672881A_ABST
Patent Text Reader

Abstract

The invention relates to a model training method, a picture generation method and device, equipment and a storage medium, multiple pictures (including objects described by text cue words) are generated through the same text cue word, multiple sample pictures are determined from the multiple pictures, the similarity of the objects contained in the multiple sample pictures is made to be higher than a preset threshold value, and the similarity of the objects contained in the multiple sample pictures is determined to be higher than the preset threshold value. According to the method, a plurality of sample pictures are obtained, sample labels (the sample labels are used for describing states of objects contained in the sample pictures) of the sample pictures are obtained, so that training samples are constructed based on the sample pictures and the sample labels corresponding to the sample pictures, the LoRA model is trained, and the target LoRA model obtained through training can help the text generation graph model to improve the text generation efficiency based on the same text cue word. The same object can be understood that the similarity of the objects is higher than the preset threshold value, and therefore the consistency of the objects contained in the multiple pictures generated by the same text prompt word is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to a model training method, an image generation method, an apparatus, a device, and a storage medium. Background Art

[0002] With the development of artificial intelligence generated content (AIGC), AI models (such as StableDiffusion and Dalle3) can generate images based on text prompts. For example, after a prompt describing an object's characteristics is input into the AI ​​model, the AI ​​model can generate multiple images containing the object based on the prompt. However, the similarity of the objects contained in these images cannot be guaranteed. That is, the objects in multiple images generated based on the same text prompt may be inconsistent. Therefore, how to ensure the consistency of objects in multiple images generated based on the same text prompt is a technical problem that needs to be solved. Summary of the Invention

[0003] In order to solve the above technical problems or at least partially solve the above technical problems, the embodiments of the present disclosure provide a model training method, an image generation method, an apparatus, a device and a storage medium.

[0004] In a first aspect, an embodiment of the present disclosure provides a model training method, the method comprising:

[0005] generating a plurality of images based on the same text prompt word, wherein the plurality of images include the object described by the text prompt word;

[0006] Determining a plurality of sample images from the plurality of images, wherein the similarity of the objects in the plurality of sample images is higher than a preset threshold;

[0007] Obtaining a sample label of the sample image, where the sample label is used to describe the state of the object;

[0008] A training sample is constructed based on the multiple sample images and the sample label corresponding to each sample image to train the LoRA model to obtain a target LoRA model.

[0009] In some implementations, determining a plurality of sample images from the plurality of images includes:

[0010] extracting features of the object from the plurality of images;

[0011] Determining similarities of the objects in the plurality of images based on features of the objects;

[0012] A sample picture is determined from the pictures whose similarity is higher than a preset threshold.

[0013] In some embodiments, determining a sample image from the images having a similarity greater than a preset threshold includes:

[0014] Grouping the pictures whose similarity is higher than a preset threshold into one group to obtain at least one picture group;

[0015] determining a group of pictures from the at least one group of pictures as a target group of pictures;

[0016] The pictures in the target picture group are used as sample pictures.

[0017] In some embodiments, determining a group of pictures from the at least one group of pictures as a target group of pictures includes:

[0018] Deleting a picture group in which the number of pictures is smaller than a preset number from the at least one picture group;

[0019] A target group of pictures is determined from the remaining groups of pictures.

[0020] In some embodiments, determining a group of pictures from the at least one group of pictures as a target group of pictures includes:

[0021] For each picture group, calculating the intra-class distance of the picture group;

[0022] The image group with the smallest intra-class distance is determined as the target image group.

[0023] In a second aspect, an embodiment of the present disclosure provides a method for generating an image, the method comprising:

[0024] Acquire text prompt words, where the text prompt words are used to describe features of the object;

[0025] The text prompt word is input into the text-generated graph model, and the target LoRA model is called by the text-generated graph model to generate multiple pictures containing the same object, and the target LoRA model is trained using any method as in the first aspect.

[0026] In a third aspect, an embodiment of the present disclosure provides a model training device, the device comprising:

[0027] A generating module, configured to generate a plurality of images based on the same text prompt word, wherein the plurality of images include the object described by the text prompt word;

[0028] a determination module, configured to determine a plurality of sample images from the plurality of images, wherein the similarity of the objects in the plurality of sample images is higher than a preset threshold;

[0029] A first acquisition module is used to acquire a sample label of the sample image, where the sample label is used to describe the state of the object;

[0030] The training module is used to construct a training sample based on the multiple sample images and the sample label corresponding to each sample image to train the LoRA model to obtain a target LoRA model.

[0031] In some embodiments, the determining module is configured to:

[0032] extracting features of the object from the plurality of images;

[0033] Determining similarities of the objects in the plurality of images based on features of the objects;

[0034] A sample picture is determined from the pictures whose similarity is higher than a preset threshold.

[0035] In some embodiments, the determining module is configured to:

[0036] Grouping the pictures whose similarity is higher than a preset threshold into one group to obtain at least one picture group;

[0037] determining a group of pictures from the at least one group of pictures as a target group of pictures;

[0038] The pictures in the target picture group are used as sample pictures.

[0039] In some embodiments, the determining module is configured to:

[0040] Deleting a picture group in which the number of pictures is smaller than a preset number from the at least one picture group;

[0041] A target group of pictures is determined from the remaining groups of pictures.

[0042] In some embodiments, the determining module is configured to:

[0043] For each picture group, calculating the intra-class distance of the picture group;

[0044] The image group with the smallest intra-class distance is determined as the target image group.

[0045] In a fourth aspect, an embodiment of the present disclosure provides an image generation device, including:

[0046] A second acquisition module is used to acquire text prompt words, where the text prompt words are used to describe the characteristics of the object;

[0047] The image generation module is used to input the text prompt word into the text graph model, and call the target LoRA model through the text graph model to generate multiple pictures containing the same object. The target LoRA model is trained using any method as in the first aspect.

[0048] In a fifth aspect, an embodiment of the present disclosure provides a computer device, the computer device comprising:

[0049] Memory;

[0050] processor; and

[0051] computer programs;

[0052] The computer program is stored in the memory and is configured to be executed by the processor to implement the method as described in the first aspect or the second aspect.

[0053] In a sixth aspect, an embodiment of the present disclosure provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the method described in the first aspect or the second aspect.

[0054] The model training method, image generation method, device, equipment and storage medium provided by the embodiments of the present disclosure generate multiple images (the multiple images include the objects described by the text prompt words) through the same text prompt word, determine multiple sample images from the multiple images, so that the similarity of the objects contained in the multiple sample images is higher than a preset threshold, and obtain sample labels of the sample images (the sample labels are used to describe the state of the objects contained in the sample images), thereby constructing training samples based on the multiple sample images and the sample labels corresponding to each sample image to train the LoRA model. The target LoRA model obtained by training can control the images generated by the text image model, so that the images generated based on the same text prompt word contain the same objects (the same objects can be understood as the similarity of the objects being higher than the preset threshold), thereby improving the consistency of the objects contained in the multiple images generated by the same text prompt word. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0056] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0057] Figure 1 is a flowchart of a model training method provided by an embodiment of the present disclosure;

[0058] Figure 2 is a flowchart of a method for determining a sample image provided by an embodiment of the present disclosure;

[0059] Figure 3 is a flowchart of a method for generating an image provided by an embodiment of the present disclosure;

[0060] Figure 4 is a structural diagram of a model training device provided by an embodiment of the present disclosure;

[0061] Figure 5 is a structural diagram of a picture generating device provided by an embodiment of the present disclosure;

[0062] Figure 6 A schematic diagram of the structure of a computer device embodiment provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0063] In order to more clearly understand the above-mentioned objectives, features and advantages of the present disclosure, the scheme of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features therein can be combined with each other in the absence of conflict.

[0064] In the following description, many specific details are set forth to facilitate a full understanding of the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; it is obvious that the embodiments in the specification are only part of the embodiments of the present disclosure, rather than all of the embodiments.

[0065] For example, Figure 1 This is a flow chart of a model training method provided by an embodiment of the present disclosure. The method can be executed by a computer device. The computer device can be exemplarily any device with computing and processing capabilities, such as a mobile phone, a computer, a distributed computing node, and a server for providing resources such as videos and pictures to users, but is not limited to the devices listed here. Figure 1 As shown, in some implementations, the model training method provided by the embodiments of the present disclosure may include steps 101 to 104.

[0066] Step 101: Generate multiple pictures based on the same text prompt word, where the multiple pictures include the object described by the text prompt word.

[0067] The text prompt words referred to in the embodiments of the present disclosure refer to prompt words in text form used to describe the characteristics of an object. The prompt words may include one or more phrases and / or one or more sentences.

[0068] The object referred to in the embodiments of the present disclosure can be exemplarily understood as any objective existence (such as a person, animal, plant, or object, etc.) or virtual existence (such as a cartoon image, but not limited to a cartoon image).

[0069] In some embodiments, a text prompt word can be input into a preset text graph model so that the text graph model generates multiple pictures based on the text prompt word. The text graph model referred to in the embodiments of the present disclosure can be understood as any text graph model in the relevant technology, such as the SDXL model or the Flux model, etc., but is not limited to the SDXL model or the Flux model. For example, in some examples, a text prompt word can be used to describe the appearance characteristics of an object. In this case, if the text prompt word is input into the preset text graph model, the output of the text graph model can be multiple pictures containing a solid color background of the object generated based on the appearance characteristics of the object.

[0070] It should be noted that the multiple pictures generated by the text graph model based on the same text prompt word all contain the object described by the text prompt word, but the objects contained in these pictures (for example, the appearance or appearance of the object) may not be consistent, that is, the objects contained in some pictures may have low similarity with the objects in other pictures.

[0071] Step 102: Determine a plurality of sample images from the plurality of images, wherein the similarity of objects generated by the text prompt words in the plurality of sample images is higher than a preset threshold.

[0072] In some embodiments, the multiple images generated in step 101 can be input into a pre-trained recognition model, and the recognition model is trained to identify whether the multiple images contain objects whose similarity is higher than a preset threshold, so that the recognition model can identify images whose object similarity is higher than the preset threshold, and use images whose object similarity is higher than the preset threshold as sample images. The training method of the recognition model can be obtained by training using the model training method provided by the prior art. The training samples of the recognition model can include multiple images and sample labels for each image. The multiple images in the training samples can be generated based on the same text prompt word. In the training samples, images whose object similarity is higher than the preset threshold are marked with the same sample label, and images whose object similarity is lower than or equal to the preset threshold are marked with different sample labels.

[0073] In other embodiments, features of objects contained in an image may be extracted, and then based on the features of the objects, the similarity of the objects contained in the image may be determined, and images whose object similarity exceeds a preset threshold may be determined as sample images. For example, in one example, features of objects contained in an image may be extracted based on a pre-trained feature extractor, such as Facebook or dinov2-base. Feature similarity of the objects may then be calculated based on a similarity calculation method such as cosine similarity, thereby using the feature similarity of the objects as the object similarity. Images whose object similarity exceeds a preset threshold may be determined as sample images.

[0074] In the embodiment of the present disclosure, multiple pictures are generated based on the same text prompt word, the similarity of the objects is determined based on the features of the objects contained in the pictures, and at least some of the pictures whose object similarity exceeds a preset threshold are determined as sample pictures, thereby ensuring the consistency of the objects in the sample pictures.

[0075] Step 103: Obtain a sample label of the sample image. The sample label is used to describe the state of the object.

[0076] In the embodiments of the present disclosure, the sample labels of the sample images are used to describe the state of the object. The state of the object may include one or more of, but not limited to, the following: hairstyle, clothing, expression, posture, color, height, proportion, weight, and shape. The sample labels of different sample images can be the same or different.

[0077] In some embodiments, after obtaining the sample image, the sample image can be displayed to the user through a preset human-computer interface. The user configures a sample label for the sample image through the human-computer interface, thereby obtaining the sample label of the sample image through the human-computer interface.

[0078] Step 104: construct a training sample based on multiple sample images and a sample label corresponding to each sample image to train the LoRA model to obtain a target LoRA model.

[0079] The LoRA (Low-Rank Adaptation) model is a lightweight plug-in used to fine-tune large language models (such as StableDiffusion). It can achieve specific control over the generated results through training with a small amount of data. In the disclosed embodiment, the images generated by the text graph model are controlled by the trained target LoRA model. The similarity of the objects contained in multiple images generated by the text graph model based on the same text prompt word is higher than a preset threshold, but the state of the objects can be different. For example, the objects may have the same appearance but different hairstyles, expressions, clothing, postures, etc.

[0080] Among them, a training sample is constructed based on multiple sample images and the sample label corresponding to each sample image. The method of training the LoRA model based on the training sample can refer to the supervised model training method, which will not be repeated in the embodiment of this disclosure.

[0081] In an embodiment of the present disclosure, multiple pictures are generated through the same text prompt word (the multiple pictures include objects described by the text prompt word), multiple sample pictures are determined from the multiple pictures, so that the similarity of the objects included in the multiple sample pictures is higher than a preset threshold, and sample labels of the sample pictures are obtained (the sample labels are used to describe the status of the objects included in the sample pictures), so as to construct training samples based on the multiple sample pictures and the sample labels corresponding to each sample picture to train the LoRA model. The target LoRA model obtained by training can help the text graph model generate pictures containing the same object (the same object can be understood as the object similarity is higher than a preset threshold) based on the same text prompt word, thereby improving the consistency of the objects included in the multiple pictures generated by the same text prompt word.

[0082] Figure 2 is a flow chart of a method for determining a sample image provided by an embodiment of the present disclosure, such as Figure 2 As shown, in some implementations, the sample image can be determined through the method of steps 201 to 203.

[0083] Step 201: Group together multiple pictures whose object similarity is higher than a preset threshold into one group to obtain at least one picture group.

[0084] In some implementations, a pre-trained feature extractor can be used to extract features of objects contained in images. A similarity calculation method, such as cosine similarity, can then be used to calculate feature similarity of the objects. This similarity can then be used as the object similarity. Images with object similarity exceeding a preset threshold are then grouped together to produce at least one image group.

[0085] In other embodiments, multiple pictures can also be input into a pre-trained grouping model, and the grouping model is trained to group pictures, dividing pictures containing similar or identical objects into a group, so that multiple pictures can be grouped by the grouping model to obtain at least one picture group. Among them, the grouping model can be trained using the model training method provided by the relevant technology. The training samples of the grouping model may include multiple picture groups, each picture group includes at least two pictures, pictures in the same group contain objects that are the same or have a similarity higher than a preset threshold, and pictures in different groups contain objects with a similarity lower than or equal to a preset threshold. The sample labels of pictures in the same group are the same, and the sample labels of pictures in different groups are different. The sample labels of pictures are used to indicate the picture group to which the pictures belong.

[0086] In some implementations, to ensure sufficient sample quantity, after obtaining the at least one picture group, picture groups containing less than a preset number of pictures may be deleted based on the number of pictures contained in the picture groups, thereby determining a target picture group from the remaining picture groups.

[0087] Step 202: Determine a picture group from the at least one picture group as a target picture group.

[0088] In some implementations, a picture group may be randomly determined from the at least one picture group obtained in step 201 as the target picture group.

[0089] In other implementations, the picture group containing the largest number of pictures among the at least one picture group obtained in step 201 may be determined as the target picture group.

[0090] In some further embodiments, the intra-class distance of sample images in each image group may be calculated, and the image group with the smallest intra-class distance may be determined as the target image group. The method for calculating the intra-class distance may be, for example, Euclidean distance, but is not limited to Euclidean distance. For example, in other embodiments, the intra-class distance may be calculated using cosine similarity. Methods for calculating the intra-class distance using Euclidean distance or cosine similarity can be found in the prior art and will not be further described here.

[0091] Step 203: Use the pictures in the target picture group as sample pictures.

[0092] In the embodiment of the present disclosure, at least one picture group is obtained by clustering pictures whose object similarity in multiple pictures is higher than a preset threshold, and then a picture group is determined from the at least one picture group obtained by clustering as a target picture group, and the pictures in the target picture group are used as sample pictures, so that the consistency of the objects contained in the sample pictures can be ensured.

[0093] Figure 3 is a flow chart of a method for generating an image provided by an embodiment of the present disclosure, such as Figure 3 As shown, in some implementations, a picture can be generated through the method of steps 301 and 302.

[0094] Step 301: Acquire text prompt words, which are used to describe the features of the object.

[0095] Step 302: Input the text prompt word into the text-generated graph model, and use the text-generated graph model to call the target LoRA model to generate multiple pictures containing the same object.

[0096] The target LoRA model can be trained using the model training method described in any of the above embodiments. The target LoRA model can control the images generated by the text graph model, ensuring that the similarity of objects in multiple images generated by the text graph model based on the same text prompt word exceeds a preset threshold. However, the objects may differ in state. For example, the objects may have the same appearance but different hairstyles, expressions, clothing, and postures. The specific training methods and beneficial effects of the target LoRA model can be found in the above embodiments related to model training and will not be further elaborated here.

[0097] Figure 4 is a structural diagram of a model training device provided by an embodiment of the present disclosure, such as Figure 4 As shown, the model training device 40 includes:

[0098] A generating module 41 is configured to generate multiple images based on the same text prompt word, wherein the multiple images include the object described by the text prompt word;

[0099] a determination module 42, configured to determine a plurality of sample images from the plurality of images, wherein the similarity of the objects in the plurality of sample images is higher than a preset threshold;

[0100] A first acquisition module 43 is configured to acquire a sample label of the sample image, where the sample label is used to describe the state of the object;

[0101] The training module 44 is used to construct a training sample based on the multiple sample images and the sample label corresponding to each sample image to train the LoRA model to obtain a target LoRA model.

[0102] In some embodiments, the determination module 42 is configured to:

[0103] extracting features of the object from the plurality of images;

[0104] Determining similarities of the objects in the plurality of images based on features of the objects;

[0105] A sample picture is determined from the pictures whose similarity is higher than a preset threshold.

[0106] In some embodiments, the determination module 42 is configured to:

[0107] Grouping the pictures whose similarity is higher than a preset threshold into one group to obtain at least one picture group;

[0108] determining a group of pictures from the at least one group of pictures as a target group of pictures;

[0109] The pictures in the target picture group are used as sample pictures.

[0110] In some embodiments, the determination module 42 is configured to:

[0111] Deleting a picture group in which the number of pictures is smaller than a preset number from the at least one picture group;

[0112] A target group of pictures is determined from the remaining groups of pictures.

[0113] In some embodiments, the determination module 42 is configured to:

[0114] For each picture group, calculating the intra-class distance of the picture group;

[0115] The image group with the smallest intra-class distance is determined as the target image group.

[0116] The model training device provided in the embodiment of the present disclosure can execute the method of the above-mentioned model training embodiment. Its execution method and beneficial effects are similar and will not be repeated here.

[0117] Figure 5 is a structural diagram of a picture generating device provided by an embodiment of the present disclosure, such as Figure 5 As shown, the image generating device 50 includes:

[0118] The second acquisition module 51 is used to acquire text prompt words, where the text prompt words are used to describe the characteristics of the object;

[0119] The image generation module 52 is used to input the text prompt word into the text graph model, and call the target LoRA model through the text graph model to generate multiple images containing the same object. The target LoRA model is trained using any method in the first aspect.

[0120] The image generation device provided by the embodiment of the present disclosure can perform the above Figure 3 The method of the embodiment has similar execution mode and beneficial effects, which will not be described in detail here.

[0121] It should also be noted that the division of modules in the above-mentioned model training device and image generation device in the embodiment of the present disclosure is schematic and is only a logical functional division. There may be other division methods in actual implementation. In addition, the functional modules in the various embodiments of the present application can be integrated into a processing module, or each module can exist physically alone, or two or more modules can be integrated into one module. The above-mentioned integrated modules can be implemented in the form of hardware or in the form of software functional modules.

[0122] If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a processor-readable storage medium. Based on this understanding, the technical solution of the present disclosure, or the part that contributes to the existing technology, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present disclosure.

[0123] Figure 6 This is a schematic diagram of the structure of a computer device embodiment provided by the present disclosure. Figure 6 As shown, the computer device includes a memory 121 and a processor 122 .

[0124] Memory 121 is used to store programs. In addition to the aforementioned programs, memory 121 may also be configured to store various other data to support operations on the computer device. Examples of such data include instructions for any application or method operating on the computer device, contact member data, phonebook member data, messages, images, videos, etc.

[0125] The memory 121 can be implemented by any type of volatile or non-volatile memory device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0126] The processor 122 is coupled to the memory 121 and executes the program stored in the memory 121 to implement the method of any of the above method embodiments.

[0127] Further, if Figure 6 As shown, the computer device may further include: a communication component 123, a power component 124, an audio component 125, a display 126 and other components. Figure 6 Only some components are shown schematically, which does not mean that the computer equipment only includes Figure 6 Components shown.

[0128] The communication component 123 is configured to facilitate wired or wireless communication between the computer device and other devices. The computer device can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, or a combination thereof. In an exemplary embodiment, the communication component 123 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 123 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared member data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0129] The power supply component 124 provides power to various components of the computer device. The power supply component 124 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the computer device.

[0130] The audio component 125 is configured to output and / or input audio signals. For example, the audio component 125 includes a microphone (MIC), which is configured to receive external audio signals when the computer device is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 121 or transmitted via the communication component 123. In some embodiments, the audio component 125 also includes a speaker for outputting audio signals.

[0131] The display 126 includes a screen, which may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, slides, and gestures on the touch panel. The touch sensor can not only sense the boundaries of a touch or slide action, but also detect the duration and pressure associated with the touch or slide operation.

[0132] In addition, an embodiment of the present disclosure further provides a computer-readable storage medium on which a computer program is stored. The computer program is executed by a processor to implement the method described in any of the above method embodiments.

[0133] In the embodiments of the present disclosure, the above-mentioned computer-readable storage medium can be any available medium or data storage device that can be accessed by the processor, including but not limited to magnetic storage (such as floppy disks, hard disks, magnetic tapes, magneto-optical disks (MO)), optical storage (such as CDs, DVDs, BDs, HVDs, etc.), and semiconductor storage (such as ROM, EPROM, EEPROM, non-volatile memory (NAND FLASH), solid-state drives (SSDs)), etc.

[0134] Those skilled in the art will appreciate that the embodiments of the present disclosure may be provided as methods, systems, or computer program products. Therefore, the present disclosure may take the form of a fully hardware embodiment, a fully software embodiment, or an embodiment combining software and hardware. Furthermore, the embodiments of the present disclosure may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage and optical storage) containing computer-usable program code.

[0135] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.

[0136] The foregoing description is intended only to provide specific embodiments of the present disclosure, intended to enable those skilled in the art to understand and implement the present disclosure. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the embodiments described herein, but rather to be construed in the broadest manner consistent with the principles and novel features disclosed herein.

Claims

1. A model training method, characterized in that: The method comprises: generating a plurality of images based on the same text prompt word, wherein the plurality of images include the object described by the text prompt word; Determining a plurality of sample images from the plurality of images, wherein the similarity of the objects in the plurality of sample images is higher than a preset threshold; Obtaining a sample label of the sample image, where the sample label is used to describe the state of the object; A training sample is constructed based on the multiple sample images and the sample label corresponding to each sample image to train the LoRA model to obtain a target LoRA model.

2. The method according to claim 1, characterized in that The determining a plurality of sample pictures from the plurality of pictures includes: extracting features of the object from the plurality of images; Determining similarities of the objects in the plurality of images based on features of the objects; A sample picture is determined from the pictures whose similarity is higher than a preset threshold.

3. The method according to claim 2, characterized in that The determining of the sample pictures from the pictures whose similarity is higher than a preset threshold comprises: Grouping the pictures whose similarity is higher than a preset threshold into one group to obtain at least one picture group; determining a group of pictures from the at least one group of pictures as a target group of pictures; The pictures in the target picture group are used as sample pictures.

4. The method according to claim 3, characterized in that The determining a group of pictures from the at least one group of pictures as a target group of pictures includes: Deleting a picture group in which the number of pictures is smaller than a preset number from the at least one picture group; A target group of pictures is determined from the remaining groups of pictures.

5. The method according to claim 3, characterized in that The determining a group of pictures from the at least one group of pictures as a target group of pictures includes: For each picture group, calculating the intra-class distance of the picture group; The image group with the smallest intra-class distance is determined as the target image group.

6. A method for generating an image, characterized in that: include: Acquire text prompt words, where the text prompt words are used to describe features of the object; The text prompt word is input into a text graph model, and a target LoRA model is called by the text graph model to generate multiple pictures containing the same object, wherein the target LoRA model is trained using the method according to any one of claims 1 to 5.

7. A model training device, characterized in that: include: A generating module, configured to generate a plurality of images based on the same text prompt word, wherein the plurality of images include the object described by the text prompt word; a determination module, configured to determine a plurality of sample images from the plurality of images, wherein the similarity of the objects in the plurality of sample images is higher than a preset threshold; A first acquisition module is used to acquire a sample label of the sample image, where the sample label is used to describe the state of the object; The training module is used to construct a training sample based on the multiple sample images and the sample label corresponding to each sample image to train the LoRA model to obtain a target LoRA model.

8. A picture generating device, characterized in that: include: A second acquisition module is used to acquire text prompt words, where the text prompt words are used to describe the characteristics of the object; The image generation module is used to input the text prompt word into the text graph model, and call the target LoRA model through the text graph model to generate multiple pictures containing the same object, and the target LoRA model is trained using the method described in any one of claims 1-5.

9. A computer device, characterized in that: The computer device comprises: Memory; processor; and computer programs; The computer program is stored in the memory and configured to be executed by the processor to implement the method according to any one of claims 1 to 6.

10. A computer-readable storage medium, characterized in that The storage medium stores a computer program, wherein the computer program implements the method according to any one of claims 1 to 6 when executed by a processor.