Method and device for determining graph generation style of text graph model, medium and equipment

By acquiring matching text with multiple predicted styles and sub-styles generated by the text-to-image model, and combining it with the image-text matching model, the style of the image generated by the text-to-image model is automatically evaluated. This solves the problem of difficulty in quickly and accurately determining the style of images in existing technologies, and achieves efficient and accurate style evaluation.

CN122020189APending Publication Date: 2026-05-12BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2024-11-06
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

The style of images generated by existing textual image models is difficult to determine quickly and accurately, resulting in a waste of time and manpower.

Method used

By acquiring the original text, a pre-set large model is used to generate matching text with various prediction styles and sub-styles. Combined with the image-text matching model, the image style generated by the text-to-image model is automatically evaluated to determine the true style.

Benefits of technology

It enables the rapid and automated determination of the style of images generated by the text-based image model, reducing time and labor costs, and improving the accuracy and scalability of the evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122020189A_ABST
    Figure CN122020189A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a method and device for determining a graphics generation style of a graphics generation model, a medium and equipment. The method comprises the following steps: acquiring an original text indicating to generate an image, and based on a preset large model, obtaining one or more prediction styles of the image generated according to the original text and N matching texts corresponding to N subdivision styles of each prediction style; inputting the original text into the target text generation graph model to obtain a target image; and for each prediction style, combining the target image with each matching text and the original text to obtain N + 1 image text pairs, inputting each image text pair into a preset image-text matching model to obtain an image-text matching score corresponding to each image text pair, and according to the image-text matching score, obtaining the image-text matching score corresponding to the image-text pair. And determining whether the predicted style is a real style of the text graph model generated image. Through the method, the style of the text graph model generated image can be efficiently and automatically determined, and the consumed time and labor cost are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to a method, apparatus, medium, and device for determining the style of a text-to-image model. Background Technology

[0002] Text-to-image (TTO) generation models are artificial intelligence models capable of generating corresponding image content based on text descriptions. They have wide applications in fields such as painting, architectural design, and educational assistance. However, due to differences in training data sources and model configuration parameters, current TTO models often inevitably generate images with specific stylistic characteristics. For example, some TTO models generate images with specific painting styles, such as realistic, cartoonish, or abstract. Others typically generate images with specific compositional styles, such as symmetrical or rule-of-thirds compositions. Users often need to spend considerable time understanding the style of the generated images when using TTO models, making it difficult to efficiently and accurately determine their style. Summary of the Invention

[0003] This disclosure describes a method, apparatus, medium, and device for determining the style of a text-to-image model.

[0004] According to the first aspect, a method for determining the style of a text-to-image model is provided, including:

[0005] Obtain the original text that indicates the generation of the image. Based on a preset large model, obtain one or more prediction styles of the image generated from the original text, and N matching texts corresponding to N sub-styles of each prediction style, where N is a natural number greater than or equal to 1.

[0006] The original text is input into the target text-to-image model to obtain the target image. For each prediction style, the target image is combined with each matching text corresponding to the prediction style and the original text to obtain N+1 image-text pairs. Each image-text pair is input into a preset image-text matching model to obtain the image-text matching score corresponding to each image-text pair. Based on the image-text matching score, it is determined whether the prediction style is the true style of the image generated by the text-to-image model.

[0007] According to the second aspect, an apparatus for determining the style of a text-to-image model is provided, comprising:

[0008] The acquisition unit is configured to acquire the original text indicating the generation of the image, and based on a preset large model, obtain one or more predicted styles of the image generated from the original text, and N matching texts corresponding to N sub-styles of each predicted style, where N is a natural number greater than or equal to 1.

[0009] The judgment unit is configured to input the original text into the target text-to-image model to obtain the target image; for each prediction style, combine the target image with each matching text corresponding to the prediction style and the original text to obtain N+1 image-text pairs; input each image-text pair into a preset image-text matching model to obtain the image-text matching score corresponding to each image-text pair; and determine whether the prediction style is the true style of the image generated by the text-to-image model based on the image-text matching score.

[0010] According to a third aspect, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed in a computer, causes the computer to perform the method of the first aspect.

[0011] According to a fourth aspect, an electronic device is provided, including a memory and a processor, wherein the memory stores executable code, and the processor executes the executable code to implement the method of the first aspect.

[0012] The present disclosure provides an apparatus, device, and medium. First, original text instructing the generation of an image is acquired. Based on a preset large model, one or more predicted styles of the image generated from the original text are obtained, along with multiple matching texts corresponding to multiple sub-styles of each predicted style. Then, the original text is input into a target text-to-image model to obtain a target image. For each predicted style, the target image is combined with each matching text corresponding to that predicted style and the original text to obtain multiple image-text pairs. Each image-text pair is input into a preset image-text matching model to obtain an image-text matching score for each pair. Based on the image-text matching score, it is determined whether the predicted style is the true style of the image generated by the text-to-image model. This method can efficiently and automatically determine the style of the image generated by the text-to-image model, improving the speed of style determination and significantly reducing the time and labor costs. Attached Figure Description

[0013] Figure 1 This diagram illustrates the manual verification of the raw image style of the raw image model.

[0014] Figure 2 A schematic diagram of a method for determining the style of a text-to-image model according to an embodiment of the present disclosure is shown;

[0015] Figure 3 A flowchart illustrating a method for determining the style of a text-to-image model according to an embodiment of the present disclosure is shown.

[0016] Figure 4A schematic diagram illustrating the saving of predicted styles to a preset database according to an embodiment of the present disclosure is shown;

[0017] Figure 5 A schematic diagram illustrating the determination of true style based on image-text matching score according to an embodiment of the present disclosure is shown;

[0018] Figure 6 A schematic block diagram of an apparatus for determining the style of a text-to-image model according to an embodiment of the present disclosure is shown;

[0019] Figure 7 A schematic diagram of the structure of an electronic device suitable for implementing embodiments of the present disclosure is shown;

[0020] Figure 8 A schematic diagram of the structure of a storage medium suitable for implementing embodiments of the present disclosure is shown. Detailed Implementation

[0021] The technical solutions provided in this specification will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the relevant invention and not intended to limit the invention. Furthermore, it should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.

[0022] In the description of the implementations disclosed herein, the term "comprising" and similar terms should be understood as open inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one / an implementation" or "the implementation" should be understood as "at least one / an implementation". The term "some implementations" should be understood as "at least some implementations". Other explicit and implicit definitions may also be included below.

[0023] As mentioned earlier, text-to-image (TTO) generation models refer to artificial intelligence models that can generate corresponding image content based on text descriptions, and they have wide applications in many fields. However, due to differences in training data sources or model configuration parameters, current TTO models inevitably generate images with specific stylistic characteristics. For example, some TTO models generate images with specific painting styles, such as realistic, cartoonish, or abstract. Others typically generate images with specific compositional styles, such as symmetrical or rule-of-thirds compositions. Still others generate images with specific color styles, such as vibrant, bright, or soft. However, users often need to use TTO models for an extended period and identify a large number of images generated by them to gradually confirm the model's inherent style. Figure 1 As shown. The problems with this approach are threefold. First, it consumes a significant amount of the user's time to determine the style of the images generated by the model, which is neither quick nor efficient and incurs substantial time costs. Second, manually determining the style of images generated by the text-based image model is limited by the knowledge and subjective judgment of human users, making it difficult to obtain accurate and quantitative evaluation results regarding the style of the model-generated images. Third, this approach cannot automatically identify the style of images generated by the text-based image model, requiring substantial human resources.

[0024] To address the aforementioned technical problems, this disclosure provides a method for determining the style of a text-to-image model. Figure 2 A schematic diagram is shown of a method for determining the image style of a text-to-image model according to an embodiment of the present disclosure. For example... Figure 2 As shown, in some embodiments, the original text used to generate the image can be input into the text-to-image model under test to obtain, for example, the target image. The original text can be input into a pre-defined large model to obtain one or more predicted styles (or simply predicted styles, e.g., ...) that the image generated from the original text might possess. Figure 2The model generates images for each of the following: Predicted style 1, predicted style 2, etc., and multiple sub-styles for each predicted style (e.g., multiple sub-styles of predicted style 1, sub-style D1, sub-style D2, sub-style D3). The matching texts for each sub-style are then used to generate images corresponding to its respective sub-style. For each predicted style, the matching texts for each sub-style, along with the original text, are combined with the target image to form image-text pairs. These image-text pairs are then input into the image-text matching model to obtain an image-text matching score. Based on the image-text matching scores of each pair, it can be determined whether the predicted style represents the true style of the image generated by the tested text-to-image model. For example, for predicted style 1, multiple image-text pairs are obtained: {target image, matching text T1}, {target image, matching text T2}, {target image, matching text T3}, and {target image, original text}. These image-text pairs can be input into the image-text matching model to obtain multiple image-text matching scores for each pair, such as image-text matching score S1, image-text matching score S2, image-text matching score S3, and image-text matching score S4. Based on these scores, it can be determined whether predicted style 1 represents the true style of the image generated by the text-to-image model. Similarly, it can be determined whether predicted style 2 represents the true style of the image generated by the text-to-image model.

[0025] The advantages of this method are as follows: First, compared to manually determining the style of images generated by the text-to-image model, this method can efficiently and automatically determine the style of images generated by the text-to-image model, improving the speed of style determination and significantly reducing the time and manpower costs. Second, it is not limited by the knowledge reserves and subjective judgment of human users, and can obtain more accurate and quantitative determination results for the style of images generated by the text-to-image model through image-text matching models. Third, by automatically expanding the predicted style of the generated images through a large model, and determining the true style of the images generated by the text-to-image model based on the predicted style and image-text matching model, this method has good scalability and the potential to automatically identify styles and their sub-styles that are currently unknown or have not yet appeared, or that users are not yet aware of or have difficulty being aware of, thus continuously improving the quality of the determination results over a long operating cycle.

[0026] The following describes the detailed process of this method.

[0027] Figure 3 A flowchart illustrating a method for determining the image style of a text-to-image model according to an embodiment of the present disclosure is shown. Figure 3As shown, the method includes at least the following steps:

[0028] Step S301: Obtain the original text indicating the generation of the image; based on the preset large model, obtain one or more prediction styles of the image generated from the original text, and N matching texts corresponding to N sub-styles of each prediction style, where N is a natural number greater than or equal to 1.

[0029] Step S303: Input the original text into the target text-to-image model to obtain the target image; for each prediction style, combine the target image with each matching text corresponding to the prediction style and the original text to obtain N+1 image-text pairs; input each image-text pair into a preset image-text matching model to obtain the image-text matching score corresponding to each image-text pair; and determine whether the prediction style is the true style of the image generated by the text-to-image model based on the image-text matching score.

[0030] First, in step S301, the original text indicating the generation of the image is obtained. Based on a preset large model, one or more prediction styles of the image generated from the original text are obtained, and N matching texts corresponding to N sub-styles of each prediction style are obtained. N can be a natural number greater than or equal to 1.

[0031] In different embodiments, the original text can be used to generate images with different specific content, and this specification does not limit this. Predicted style refers to the style that the image generated by the text-to-image model (TPA) after inputting the original text into the tested TPA model (e.g., the target TPA model), as inferred by a large model, may possess. In different embodiments, the predicted style obtained from the original statement can be one or more. Sub-style refers to further refined style categories included within a specific type of predicted style.

[0032] In different embodiments, the specific methods for obtaining various predicted styles of the image generated from the original text, various sub-styles of each predicted style, and various matching texts corresponding to each sub-style can differ based on a preset large model. In one example, a first prompt word can be constructed, comprising the original text and a first prompt portion. The first prompt portion indicates one or more predicted styles of the generated image inferred from the original text, and N matching texts corresponding to N sub-styles of each predicted style. The first prompt word is input into the preset large model to obtain the one or more predicted styles and the N matching texts corresponding to the N sub-styles of each predicted style. In this way, various predicted styles of the image generated from the original text, various sub-styles of each predicted style, and various matching texts corresponding to each sub-style can be efficiently obtained using a preset large model without consuming significant time and manpower costs.

[0033] A large model typically refers to an artificial intelligence model with hundreds of millions or more parameters, pre-trained on massive datasets. In different embodiments, the pre-set large model can be of different specific types or have different neural network structures. A prompt is the text information input to the large model, intended to guide it in generating corresponding content based on the prompt. In different embodiments, the first prompt can be a prompt word guiding the large model to generate different specific content; this specification does not limit this.

[0034] In another embodiment, the first prompt word further includes a second prompt portion, which indicates example text for generating the image, style examples of the example image generated based on the example text, and examples of matching text corresponding to multiple sub-styles of the style examples. In this way, by providing the large model with example text, as well as the predicted style of the image generated based on the example text, the sub-styles of the predicted style, and examples of matching text corresponding to the sub-styles, the large model's understanding of the content to be generated can be improved, further enhancing the accuracy of the large model's output.

[0035] Specifically, in different embodiments, the predicted style obtained from the original statement can be different, and different predicted styles can have different sub-styles. For example, in one example, the predicted style for the original statement is a painting style, which includes several further refined style categories, such as realistic, cartoon, and abstract. In another example, for example, one predicted style for the original statement is a color style, which also includes several further refined style categories, such as vibrant, bright, and soft. Each matching text corresponding to each sub-style refers to the matching text of the original text that has that sub-style. In different embodiments, the specific representation of the matching text of each sub-style can be different. In one example, for example, the original text is "a little boy wearing Hanfu", and the predicted style for the original text is a painting style, the sub-categories of which include, for example, realistic, cartoon, and abstract. Among them, the matching text corresponding to the realistic sub-category (or simply realistic style) can be, for example, "a little boy wearing Hanfu in a realistic style", the matching text corresponding to the cartoon sub-category (or simply cartoon style) can be, for example, "a little boy wearing Hanfu in a cartoon style", and the matching text corresponding to the abstract sub-category (or simply abstract style) can be, for example, "a little boy wearing Hanfu in an abstract style".

[0036] Furthermore, in one embodiment, the predicted style and multiple matching texts corresponding to multiple sub-styles of the predicted style can be retrieved from a preset knowledge base based on the original text. The knowledge base pre-associates and stores the original text and multiple matching texts corresponding to one or more predicted styles and multiple sub-styles of each predicted style obtained based on a preset large model. In this way, the predicted style of the original text and the matching texts corresponding to the sub-styles of the predicted style can be obtained in advance through a large model and associated with the original text and stored in the knowledge base, such as... Figure 4 As shown. Therefore, when it is desired to determine the raw graph style of the text graph model under test, the predicted style of the original text and the matching text corresponding to the sub-styles of the predicted style can be directly retrieved from the knowledge base based on the original text, further accelerating the determination of the raw graph style of the text graph model.

[0037] Then, in step S303, the original text can be input into the target text-to-image model to obtain the target image. The text-to-image model is a neural network model that can generate a matching image based on the input text description. The target text-to-image model is the text-to-image model being tested. In different embodiments, the target text-to-image model can be of different specific types or have different neural network structures.

[0038] After acquiring the target image, for each predicted style obtained in step S301, the target image can be combined with each matching text corresponding to that predicted style and the original text to obtain N+1 image-text pairs. Each image-text pair is then input into a preset image-text matching model to obtain an image-text matching score for each pair. Based on the image-text matching score, it is determined whether the predicted style is the true style of the image generated by the text-to-image model. The image-text matching model is a multimodal neural network model designed to understand the semantic relationship between image and text data of different modalities and to determine whether they match. In different embodiments, the preset image-text matching model can be different specific models.

[0039] In different embodiments, the specific method for determining whether the predicted style is the true style of the image generated by the target text-to-image model based on the image-text matching score can vary. In one embodiment, for each predicted style, the first image-text pair with the highest image-text matching score can be determined from multiple image-text pairs based on the image-text matching scores corresponding to each image-text pair of that predicted style. If the first image-text pair is obtained by combining any one of the N matching texts of that predicted style with the target image, then that predicted style can be determined as the true style of the image generated by the text-to-image model. Here, the true style, or habitual style, can be the style that the image generated by the text-to-image model has a high probability of occurring. For example... Figure 5 As shown, for predicted style 1 (specifically, painting style), the image-text pair consisting of the target image generated by the target text-to-image model and the matching text "a realistic-style little boy wearing Hanfu" achieves the highest matching score of 0.94 among all image-text pairs for the predicted styles. Therefore, it can be confirmed that the image generated by the target text-to-image model has a habitual painting style, or in other words, a specific painting style preference. In this way, the true style of the image generated by the target text-to-image model can be accurately determined from various predicted styles.

[0040] In another embodiment, if the first image-text pair is obtained by combining the original text with the target image, then it can be determined that the predicted style is not the true style of the image generated by the text-to-image model. For example Figure 5 As shown, for prediction style 2 (specifically, color style), the image-text pair consisting of the target image generated by the target text-to-image model and the original text "a little boy wearing Hanfu" achieves the highest matching score of 0.85 among all image-text pairs in the prediction styles. Therefore, it can be confirmed that the image generated by the target text-to-image model does not have a habitual color style, that is, it does not have a specific color preference. In this way, styles that the target text-to-image model does not favor can be accurately excluded from various prediction styles.

[0041] In some scenarios, further determining the sub-styles of the true style of the image generated by the target text-to-image model can, for example, help the user determine whether to use the target text-to-image model to generate an image with a specific sub-style. Therefore, in one embodiment, for various predicted styles, if it is determined that the predicted style is the true style of the image generated by the text-to-image model, then based on the sub-styles corresponding to the matching text in the first image-text pair, the sub-styles in the true style of the image generated by the text-to-image model can also be determined. For example... Figure 5 As shown, for prediction style 1 (specifically, painting style), the image-text pair consisting of the target image and the matching text in the realistic style has the highest matching score. Therefore, it can be determined that among various specific painting styles, the target text-to-image model prefers to generate images in the realistic style.

[0042] Figure 6 A schematic block diagram of an apparatus for determining the style of a text-to-image model according to an embodiment of the present disclosure is provided. The apparatus is used to perform, for example... Figure 3 The method shown. (As shown) Figure 6 As shown, the device 600 includes:

[0043] The acquisition unit 601 is configured to acquire the original text indicating the generation of the image, and based on a preset large model, obtain one or more prediction styles of the image generated from the original text, and N matching texts corresponding to N sub-styles of each prediction style, where N is a natural number greater than or equal to 1.

[0044] The judgment unit 602 is configured to input the original text into the target text-to-image model to obtain the target image; for each prediction style, combine the target image with each matching text corresponding to the prediction style and the original text to obtain N+1 image-text pairs; input each image-text pair into a preset image-text matching model to obtain the image-text matching score corresponding to each image-text pair; and determine whether the prediction style is the true style of the image generated by the text-to-image model based on the image-text matching score.

[0045] This disclosure also provides an electronic device, including a memory and a processor. The memory stores executable code, and when the processor executes the executable code, it implements, for example... Figure 3 The method shown.

[0046] The following can also be referenced Figure 7 It shows a schematic diagram of the structure of an electronic device 700 suitable for implementing embodiments of the present disclosure. Figure 7 The illustrated electronic device 700 is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0047] like Figure 7 As shown, the electronic device 700 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 701. The aforementioned processing device 701 may be a general-purpose processor, a digital signal processor (DSP), a microprocessor, or a microcontroller, and may further include an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can perform various appropriate actions and processes based on a program stored in read-only memory (ROM) 702 or a program loaded from storage device 708 into random access memory (RAM) 703. RAM 703 also stores various programs and data required for the operation of the electronic device 700. The processing device 701, ROM 702, and RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0048] Typically, the following devices can be connected to I / O interface 705: input devices 706 including, for example, touchscreens, touchpads, keyboards, mice, etc.; output devices 707 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 708 including, for example, magnetic tapes, hard disks, etc.; and communication devices 709. Communication device 709 allows electronic device 700 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 7 An electronic device 700 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively. Figure 7 Each box shown can represent a device or multiple devices as needed.

[0049] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 709, or installed from storage device 708, or installed from ROM 702. When the computer program is executed by processing device 701, it performs the functions defined in the method for determining the style of a text-based image model provided by embodiments of this disclosure.

[0050] This disclosure also provides a computer-readable storage medium storing a computer program thereon, which, when executed in a computer, causes the computer to perform the functions provided in this disclosure. Figure 3 The method shown is for determining the style of a text-to-image model. Figure 8 A schematic diagram illustrating a storage medium for implementing an embodiment of this disclosure. For example, such as... Figure 8As shown, the storage medium 800 can be a non-transitory computer-readable storage medium used to store non-transitory computer-executable instructions 801. When the non-transitory computer-executable instructions 801 are executed by a processor, a method for determining the image style of a text-to-image model provided in this disclosure embodiment can be implemented. For example, when the non-transitory computer-executable instructions 801 are executed by a processor, one or more steps in the method for determining the image style of a text-to-image model provided in this disclosure embodiment can be performed. For example, the storage medium 800 can be applied in the above-mentioned electronic device. For example, the storage medium 800 can include the memory in the electronic device. The description of the storage medium 800 can be found in the description of the memory in the embodiments of the electronic device, and will not be repeated here. The specific functions and technical effects of the storage medium 800 can be found in the description of the method for determining the image style of a text-to-image model provided in this disclosure embodiment, and will not be repeated here.

[0051] It should be noted that the computer-readable medium in the embodiments of this disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a memory card of a smartphone, a storage component of a tablet computer, a portable computer disk, a hard disk of a personal computer, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In the embodiments of this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the embodiments of this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, wherein computer-readable program code is carried. The transmitted data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (Radio Frequency), etc., or any suitable combination thereof.

[0052] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device. The aforementioned computer-readable medium carries one or more programs, which, when executed by the server, cause the electronic device to implement the method for determining the style of a text-to-image model provided in the embodiments of this disclosure.

[0053] Computer program code for performing the operations of embodiments of this disclosure can be written in one or more programming languages ​​or a combination thereof. Programming languages ​​include object-oriented programming languages—such as Java, Smalltalk, and C++—and conventional procedural programming languages—such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0054] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions. The units described in the embodiments of the present disclosure may be implemented in software or hardware. The names of the units do not necessarily constitute a limitation on the unit itself. The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), Systems-on-Chip (SOCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0055] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments for storage media and computing devices are basically similar to the method embodiments, so they are described more simply; relevant parts can be referred to the descriptions of the method embodiments.

[0056] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in this disclosure. Furthermore, although the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of a single embodiment may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0057] The above detailed embodiments further illustrate the purpose, technical solution, and beneficial effects of the embodiments of the present invention. Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative forms of implementing the claims. It should be understood that the above are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made on the basis of the technical solutions of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for determining the style of a text-to-image model, comprising: Obtain the original text that indicates the generation of the image. Based on a preset large model, obtain one or more prediction styles of the image generated from the original text, and N matching texts corresponding to N sub-styles of each prediction style, where N is a natural number greater than or equal to 1. The original text is input into the target text-to-image model to obtain the target image. For each prediction style, the target image is combined with each matching text corresponding to the prediction style and the original text to obtain N+1 image-text pairs. Each image-text pair is input into a preset image-text matching model to obtain the image-text matching score corresponding to each image-text pair. Based on the image-text matching score, it is determined whether the prediction style is the true style of the image generated by the text-to-image model.

2. The method according to claim 1, wherein, Based on a pre-defined large model, one or more predicted styles of the image generated from the original text are obtained, along with N matching texts corresponding to N sub-styles of each predicted style, including: Construct a first prompt word, which includes original text and a first prompt portion. The first prompt portion indicates one or more predicted styles of the generated image inferred from the original text, and N matching texts corresponding to N sub-styles of each predicted style. Input the first prompt word into a preset large model to obtain the one or more predicted styles and N matching texts corresponding to N sub-styles of each predicted style.

3. The method according to claim 1, wherein, Based on the image-text matching score, the target style of the image generated by the text-to-image model is determined, including: Based on the image-text matching scores corresponding to each image-text pair, the first image-text pair with the highest image-text matching score is determined from multiple image-text pairs. If the first image-text pair is obtained by combining any one of the N matching texts with the target image, then the predicted style is determined to be the true style of the image generated by the text-to-image model.

4. The method according to claim 3, further comprising: If the first image-text pair is obtained by combining the original text with the target image, then the predicted style is determined to be the true style of the image generated by the text-to-image model.

5. The method according to claim 2, further comprising: If the predicted style is determined to be the true style of the image generated by the text-to-image model, then the sub-style in the true style of the image generated by the text-to-image model is determined according to the sub-style corresponding to the matching text in the first image-text pair.

6. The method according to claim 1, wherein, The first prompt word also includes a second prompt portion, which indicates example text for generating the image, style examples of the example image generated based on the example text, and examples of multiple matching texts corresponding to multiple sub-styles of the style examples.

7. The method according to claim 1, wherein, Based on a pre-defined large model, one or more predicted styles of the image generated from the original text are obtained, along with N matching texts corresponding to N sub-styles of each predicted style, including: Based on the original text, one or more prediction styles and N matching texts corresponding to N sub-styles of each prediction style are retrieved from a preset knowledge base. The knowledge base contains the original text, the one or more prediction styles obtained based on a preset large model, and multiple matching texts corresponding to multiple sub-styles of each prediction style.

8. An apparatus for determining the style of a text-to-image model, comprising: The acquisition unit is configured to acquire the original text indicating the generation of the image, and based on a preset large model, obtain one or more predicted styles of the image generated from the original text, and N matching texts corresponding to N sub-styles of each predicted style, where N is a natural number greater than or equal to 1. The judgment unit is configured to input the original text into the target text-to-image model to obtain the target image; for each prediction style, combine the target image with each matching text corresponding to the prediction style and the original text to obtain N+1 image-text pairs; input each image-text pair into a preset image-text matching model to obtain the image-text matching score corresponding to each image-text pair; and determine whether the prediction style is the true style of the image generated by the text-to-image model based on the image-text matching score.

9. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to perform the method of any one of claims 1-7.

10. An electronic device comprising a memory and a processor, wherein the memory stores executable code, and the processor, when executing the executable code, implements the method of any one of claims 1-7.