Image-text question generation method and electronic equipment

By identifying text elements and determining image positions during the image-text question generation process, and using an image generation model to output and fill the target image, the problem of unstable image elements in AI model-generated image-text questions is solved, achieving more efficient and stable image-text question generation.

CN121147352APending Publication Date: 2025-12-16WANGYIYOUDAO INFORMATION TECH BEIJING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511013309.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-22
Publication Date
2025-12-16

AI Technical Summary

Technical Problem

In existing technologies, the image elements in text and image questions generated by AI models are unstable, making them difficult to apply in practice, and the training and adjustment costs of the models are high.

Method used

By identifying target text elements from the text title, determining the image position in the title display template, inputting the target text elements into a preset image generation model, outputting the target image, filling the specified position, and generating a text and image title.

Benefits of technology

It reduces the model's understanding cost and task complexity, improves the accuracy of image content and the uniformity of layout, and makes the image and text titles have higher practical application value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121147352A_ABST
    Figure CN121147352A_ABST
Patent Text Reader

Abstract

The invention provides an image-text question generation method and electronic equipment. The method comprises the following steps: acquiring a text topic; wherein the text question is in a text format; identifying a target text element from the text question, and determining a question display template corresponding to the text question based on the target text element; wherein the question display template comprises at least one picture position; the picture position is located at a specified position in the question display template, and the picture position is used for filling a picture; determining a corresponding relationship between the picture position and the target text element; inputting the target text element into a preset picture generation model, and outputting a target picture corresponding to the target text element; and filling a picture position corresponding to the target text element with the target picture to obtain an image-text topic corresponding to the text topic. According to the mode, the picture layout is more uniform, the effect stability of the picture elements is improved, and the image-text questions have higher practical application value.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of content generation, and in particular to a method for generating a picture-text question and an electronic device. BACKGROUND

[0002] In a picture-text question, picture elements can appear in the stem and the options. In related technologies, a text prompt word is input into an AI model, and the AI model can automatically generate a picture-text question according to the text prompt word. However, the effect of the picture elements in the picture-text question generated by the AI model is unstable, which makes it difficult for the picture-text question to be practically applied. SUMMARY

[0003] Therefore, the present application aims to provide a method for generating a picture-text question and an electronic device to improve the stability of the effect of picture elements and make the picture-text question have higher practical application value.

[0004] In a first aspect, an embodiment of the present application provides a method for generating a picture-text question. The method comprises: obtaining a text question; wherein the text question is in a text format; identifying a target text element from the text question, and determining a question display template corresponding to the text question based on the target text element; wherein the question display template comprises at least one picture position; the picture position is located at a specified position in the question display template, and the picture position is used to fill a picture; determining a corresponding relationship between the picture position and the target text element; inputting the target text element into a preset picture generation model to output a target picture corresponding to the target text element; and filling the target picture into the picture position corresponding to the target text element to obtain a picture-text question corresponding to the text question.

[0005] In a second aspect, an embodiment of the present application provides a device for generating a picture-text question. The device comprises: a text question obtaining module configured to obtain a text question; wherein the text question is in a text format; a template determining module configured to identify a target text element from the text question, and determine a question display template corresponding to the text question based on the target text element; wherein the question display template comprises at least one picture position; the picture position is located at a specified position in the question display template, and the picture position is used to fill a picture; a corresponding relationship determining module configured to determine a corresponding relationship between the picture position and the target text element; a picture output module configured to input the target text element into a preset picture generation model to output a target picture corresponding to the target text element; and a picture-text question generating module configured to fill the target picture into the picture position corresponding to the target text element to obtain a picture-text question corresponding to the text question.

[0006] In a third aspect, an electronic device is provided, which includes a processor and a memory. The memory stores computer executable instructions capable of being executed by the processor. The processor executes the computer executable instructions to implement the method for generating a graphic-text question.

[0007] In a fourth aspect, a computer readable storage medium is provided, which stores computer executable instructions. When the computer executable instructions are invoked and executed by a processor, the computer executable instructions cause the processor to implement the method for generating a graphic-text question.

[0008] The embodiments of the present application have the following beneficial effects:

[0009] The method for generating a graphic-text question and the electronic device, obtain a text question; wherein the text question is in a text format; identify a target text element from the text question, determine a question display template corresponding to the text question based on the target text element; wherein the question display template includes at least one picture position; the picture position is located at a specified position in the question display template, and the picture position is used to fill a picture; determine a correspondence between the picture position and the target text element; input the target text element into a preset picture generation model, output a target picture corresponding to the target text element; fill the target picture into the picture position corresponding to the target text element, and obtain a graphic-text question corresponding to the text question.

[0010] In the above manner, after the target text element is identified from the text question, the target picture corresponding to the target text element is generated through the picture generation model, and then the target picture is filled into the question display template to obtain the graphic-text question. The picture generation model only needs to output the target picture corresponding to the target text element, which reduces the understanding cost and task complexity of the model, makes the content of the output picture more accurate, and at the same time, the target picture is filled into the specified position in the question display template, which makes the picture layout more uniform, improves the effect stability of the picture element, and the graphic-text question has higher practical application value.

[0011] Other features and advantages of the present application will be further described in the following specification, and in part will become apparent to those skilled in the art from the following specification, or will be learned by practice of the present application. The objects and other advantages of the present application will be realized and achieved by the structure particularly pointed out in the specification, claims and drawings.

[0012] In order to make the above objectives, characteristics and advantages of the present application more apparent and easy to understand, the following preferred embodiments are specifically described below, and the accompanying drawings are referred to. BRIEF DESCRIPTION OF DRAWINGS

[0013] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0014] Figure 1 This is a diagram illustrating text and image questions generated by AI models in related technologies.

[0015] Figure 2 A flowchart illustrating a method for generating graphic titles according to an embodiment of the present invention;

[0016] Figure 3 A schematic diagram of a question display template provided in an embodiment of the present invention;

[0017] Figure 4 A schematic diagram of another question display template provided in an embodiment of the present invention;

[0018] Figure 5 A schematic diagram of the original target image and the processed target image provided for embodiments of the present invention;

[0019] Figure 6 A schematic diagram of the title provided for an embodiment of the present invention;

[0020] Figure 7 A schematic diagram of a graphic title generation device provided in an embodiment of the present invention;

[0021] Figure 8 This is a schematic diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0023] In related technologies, when generating text-based questions using AI models, the content, layout, and style of the image elements in the questions may be unstable. As an example, the text prompts include: "Generate questions, converting text to images: [Question Stem] Petals -> Flower: Keyboard -> ? [Options] A. Computer B. Typing C. Screen D. Wire E. Program".

[0024] After inputting the above text prompts into the AI ​​model, the AI ​​model outputs... Figure 1 The three types of image-text questions shown in (a), (b), and (c) differ in content, layout, and style; the effects of the image elements are unstable, making the image-text questions difficult to apply in practice.

[0025] To obtain image-text questions that can be practically applied, it is usually necessary to train the model with a large number of image-text questions and adjust the model parameters of the AI ​​model. Alternatively, the AI ​​model can be called through an API (Application Programming Interface) to continuously adjust the text prompts until the AI ​​model generates image-text questions that meet the application requirements. However, this method has high labor and time costs.

[0026] Based on this, the method and electronic device for generating graphic and textual titles provided in this embodiment of the invention can be applied to the generation of graphic and textual titles in various scenarios.

[0027] To facilitate understanding of this embodiment, a method for generating graphic titles disclosed in this invention will first be described in detail, such as... Figure 2 As shown, the method for generating this graphic title includes the following steps:

[0028] Step S202: Obtain the text title; wherein the text title is in text format;

[0029] This text question is in text format and does not contain images. In one example, the text question includes:

Question Stem

Options

[0030] The text title can include one or more, and when there are multiple text titles, they can be arranged in a list. The text titles can be pre-saved or generated by a pre-trained title generation model.

[0031] Step S204: Identify target text elements from the text question, and determine the question display template corresponding to the text question based on the target text elements; wherein, the question display template includes at least one image position; the image position is located at a specified position in the question display template, and the image position is used to fill the image;

[0032] The target text element can be understood as a text element that needs to be converted into an image element. For example, 'branch', 'tree', 'hand', 'two hands', 'arm', 'fingers', and 'fingernails' in the example above can all be used as target text elements.

[0033] In this embodiment, the target text element is used to determine the question display template corresponding to the text question. Specifically, the question display template corresponding to the text question can be determined according to the number and type of the target text element, the number and type of the target text element in the stem, the number and type of the target text element in the options, and the like.

[0034] The number and position of the picture position in the question display template match the number and position of the target text element in the text question. In the question display template, the picture position is located at a specified position, which can be determined by coordinates. The picture position is used to fill the picture corresponding to the target text element.

[0035] The purpose of this embodiment is to convert the target text element in the text question into a picture and fill the picture into the question display template corresponding to the text question to obtain a graphic text question matching the text question. Since the picture position in the question display template is relatively fixed, the picture layout in the generated graphic text question is relatively fixed. When the text question includes multiple text questions, the picture layout in the graphic text questions of the multiple text questions is relatively uniform.

[0036] Step S206: Determine the correspondence between the picture position and the target text element.

[0037] When the target text element includes one and the picture position includes one, the target text element corresponds to the picture position. When the target text element includes multiple and the picture position includes multiple, the correspondence between the picture position and the target text element needs to be determined in advance. Specifically, the correspondence between the picture position and the target text element can be determined according to the arrangement order of the target text element in the text question and the arrangement order of the picture position in the question display template.

[0038] For example, the multiple target text elements are encoded according to the arrangement order of the target text element in the text question, and the multiple picture positions are encoded according to the arrangement order of the picture position in the question display template. The target text element and the picture position with the same code have a corresponding relationship.

[0039] Step S208: Input the target text element into a preset picture generation model to output a target picture corresponding to the target text element.

[0040] The picture generation model can be a pre-trained AI model, such as a diffusion model, a generative adversarial network model, a cross-modal model, and the like. In this embodiment, each target text element is input into the picture generation model one by one to output a target picture corresponding to the target text element. For example, the target text element is 'flower', and the target picture includes a flower pattern. For another example, the target text element is 'keyboard', and the target picture includes a keyboard pattern.

[0041] In the related art, the AI model directly outputs the graphic text title, while in the present embodiment, the picture generation model only needs to output the target picture corresponding to the target text element, reducing the understanding cost of the model, reducing the complexity of the model output task, and making the output picture content more accurate.

[0042] In step S210, the target picture is filled into the picture position corresponding to the target text element to obtain the graphic text title corresponding to the text title.

[0043] The aforementioned correspondence between the picture position and the target text element is determined, and after the picture generation model outputs the target picture corresponding to each target text element, the target picture is filled into the corresponding picture position according to the correspondence to obtain the graphic text title.

[0044] It should be noted that the question display template can also include part of the text, for example, option sub-identifiers such as 'A', 'B', 'C', etc.; The question display template can also include a text position for filling in text content other than the target text element in the text title.

[0045] The above method for generating a graphic text title obtains a text title; wherein the text title is in a text format; identifies a target text element from the text title, and determines a question display template corresponding to the text title based on the target text element; wherein the question display template includes at least one picture position; the picture position is located at a specified position in the question display template, and the picture position is used to fill in a picture; determines the correspondence between the picture position and the target text element; inputs the target text element into a preset picture generation model to output a target picture corresponding to the target text element; and fills the target picture into the picture position corresponding to the target text element to obtain a graphic text title corresponding to the text title.

[0046] In the above manner, after identifying the target text element from the text title, the target picture corresponding to the target text element is generated by the picture generation model, and then the target picture is filled into the question display template to obtain the graphic text title; the picture generation model only needs to output the target picture corresponding to the target text element, reducing the understanding cost and task complexity of the model, making the output picture content more accurate; at the same time, the question display template fills the target picture into the specified position, making the picture layout more uniform, improving the effect stability of the picture element, and the graphic text title has higher practical application value.

[0047] In one implementation, a prompt word for generating a text title is obtained; wherein the prompt word includes at least one of a knowledge point examined by the question, a logical requirement of the question, a reference question, and an output format; and the prompt word is input into a preset reasoning model to output the text title.

[0048] The above topic examines the knowledge point used to indicate the knowledge point examined by the generated text topic, for example, the analogy topic examines the synonym relationship, the inclusion relationship, the antonym relationship, etc. In a specific example, the prompt words corresponding to the topic knowledge point include: "In the language analogy question, see that there is only a'synonym' relationship between the first two words; select the words with the same relationship from the answer options."

[0049] The above topic logic requirement is used to indicate the number of correct answers in the options, the meanings of correct answers and incorrect answers need to be mutually exclusive, the difficulty of the options, etc. In an example, the prompt words corresponding to the topic logic requirement include: "Ensure that the answer is unique; the options other than the correct answer cannot have a synonym relationship with the correct answer; the other options other than the correct answer cannot have another logical relationship with the words in the stem so that this option can also be an answer, and the answer must be unique; difficulty classification" Among them, the difficulty classification can be specific to grade, age, etc.

[0050] The above reference question can be a specific text question example, including stem, options, answers, analysis, difficulty, etc. For example:

Stem

Options

Answer

Analysis

Difficulty label

[0051] The above output format is used to indicate the format of the text question to be output, which can include identification, symbols, etc. For example:

Stem

Options

Answer

Analysis

Difficulty label

[0052] The above reasoning model needs to be pre-trained, which can be a deepseek model, a chatGPT model, etc.

[0053] The above text question includes multiple; the text question includes a stem part and an option part; among them, the stem part is associated with a stem identifier, and the stem identifier is used to indicate the position of the stem part in the text question; the option part is associated with an option identifier, and the option identifier is used to indicate the position of the option part in the text question.

[0054] The stem identifier can be a specific symbol, word, etc. For example, the stem identifier is '

Stem

Options

[0055] In an implementation manner, the multiple text questions are arranged continuously, a text part between the stem identifier and the next option identifier of the stem identifier is a stem part corresponding to the stem identifier, and a text part between the option identifier and the next stem identifier of the option identifier is an option part corresponding to the option identifier.

[0056] If the text question is output by the reasoning model, the reasoning model can be controlled to output the corresponding stem identifier and option identifier while outputting the text question. If the text question is saved, the stem identifier and the option identifier can be set manually or automatically recognized by the model, and then set.

[0057] According to the arrangement order of the stem identifier and the option identifier in the multiple text questions, the stem part and the option part belonging to the same text question are recognized.

[0058] When the multiple text questions are arranged in a list or other forms, the stem identifier and the option identifier in the multiple text questions can be recognized. For a certain stem identifier, the adjacent next identifier of the stem identifier is the option identifier, and the stem part corresponding to the stem identifier and the option part corresponding to the option identifier constitute a text question. Through the above manner, the stem part and the option part belonging to the same text question can be recognized. When generating a target picture subsequently, the stem part and the option part belonging to the same text question can assist the picture generation model to understand the text scene environment where the target picture is located, thereby improving the accuracy of the picture content.

[0059] In addition, by recognizing the stem part and the option part belonging to the same text question, the number of questions of the multiple text questions can be counted, and whether the number of questions meets the requirement can be determined. If not, the text question needs to be generated continuously by the reasoning model.

[0060] Further, the stem part is further provided with a stem format indication identifier. The stem format indication identifier is used to indicate the format of the stem part. The stem indication identifier can be a word, a symbol, etc. For example, in “

stem

[0061] For each text question, the stem text element is recognized from the stem part based on the format indication identifier of the stem part. For example, the text at both ends of the format indication identifier is the target text element, and for another example, the text after the format indication identifier is the target text element. A regular expression can be set according to the format indication identifier, and the stem text element is extracted from the stem part by the regular expression.

[0062] The option text element is identified from the option part based on a structured format corresponding to the option part. The structured format of the option part includes: an option sub-identifier and an option content corresponding to the option sub-identifier. The option sub-identifier can be "A", "B", "C", etc., or "1", "2", "3", etc. Each option sub-identifier corresponds to an option content. For example, the option content corresponding to the option sub-identifier A is 'flower', and the option content corresponding to the option sub-identifier B is 'potted plant', etc. After the option sub-identifier is identified, the text that is continuous after the option sub-identifier is determined as the option text element. The aforementioned target text element includes the option text element and the stem text element.

[0063] In an implementation manner, a question display template corresponding to the text question is determined based on the element quantity and / or element distribution position of the target text element. The number of positions of the picture positions in the question display template matches the element quantity of the target text element.

[0064] In determining the question display template, the element quantity of the target text element can be used alone, or the element quantity and the element distribution position of the target text element can be used together, or the element quantity and the element distribution position of the target text element are used simultaneously.

[0065] For example, the number of positions of the picture positions in the question display template matches the element quantity of the target text element. For another example, the number of positions of the picture positions in the stem part of the question display template matches the element quantity of the stem text element in the target text element, and the number of positions of the picture positions in the option part of the question display template matches the element quantity of the option text element in the target text element.

[0066] For another example, the arrangement manner between the target text elements in the text question is identified, so that the arrangement manner of the picture positions in the determined question display template is the same as the arrangement manner between the target text elements. In an example, the stem text element in the stem part of the text question is two per row and two rows in total, and then the picture positions in the stem part of the determined question display template are also two per row and two rows in total.

[0067] Figure 3 As an example, there are six picture positions corresponding to six target text elements, in which three picture positions are located in the stem part, and the other three picture positions are located in the option part.

[0068] Figure 4 As another example, there are eight picture positions corresponding to eight target text elements, in which three picture positions are located in the stem part, and the other five picture positions are located in the option part.

[0069] The multiple text questions correspond to the same or similar question display templates, which can be the layout and style of the final output of the graphic text questions, so that the graphic text questions have higher practical application value.

[0070] Further, the text question and the target text element are filled into the preset prompt word template to generate a picture prompt word; and the picture prompt word is input into a preset picture generation model to determine a picture belonging scene of the picture based on the text question by using the picture generation model, and output a target picture corresponding to the target text element based on the picture belonging scene.

[0071] The prompt word template provides a filling position of the text question and a filling position of the target text element, wherein the filling position of the text question is associated with a text question identifier to assist the picture generation model in identifying which prompt words are the text question; and the filling position of the target text element is associated with a text element identifier to assist the picture generation model in identifying which prompt words are the target text element.

[0072] The picture generation model can run on a terminal device or a server. In actual implementation, the picture generation model can be implemented by RPA (Robotic Process Automation), and the picture prompt word is input into a webpage to obtain the target picture.

[0073] In the process of generating the target picture, the user performs operations by simulating the above-mentioned RPA mode, which reduces the question generation cost and has high economic efficiency compared with the mode of continuously adjusting the prompt words to generate graphic text questions by API calling an AI model in the related art.

[0074] In actual implementation, one text question corresponds to a group of picture prompt words, and one text question can include one or more target text elements. After the picture prompt words are input into the picture generation model, the picture belonging scene is first determined based on the text question, for example, a human body scene, a fruit scene, a home scene, a landscape scene, etc.; and then a target picture matching the picture belonging scene is output, and the picture content of the target picture matches the target text element.

[0075] In another mode, after the picture prompt words are input into the picture generation model, the picture generation model first outputs a target picture with picture content matching the target text element based on the target text element, and then determines the picture belonging scene based on the text question to determine whether the target picture matches the picture belonging scene. If not, the target picture corresponding to the target text element is regenerated. Multiple target pictures can also be generated for one target text element, and the target picture matching the picture belonging scene is selected from the multiple target pictures, and the target picture not matching the picture belonging scene is deleted.

[0076] In one example, the target text element is 'apple', if it is known from the text title that the picture of the target text element belongs to a fruit scene, then the target picture is a fruit apple; if it is known from the text title that the picture of the target text element belongs to an electronic device scene, then the target picture is an apple phone.

[0077] In the above manner, the text title and the target text element are jointly input as prompt words into the picture generation model, so that the picture content of the output target picture matches the target text element, and the scene to which the target picture belongs matches the text title, thereby improving the accuracy of the target picture.

[0078] In one specific implementation, the target picture is evaluated based on a preset evaluation standard to obtain an evaluation parameter; if the evaluation parameter does not reach a preset parameter range, the target picture is deleted, and the step of inputting the target text element into the preset picture generation model to output the target picture corresponding to the target text element is continued until the evaluation parameter reaches the preset parameter range.

[0079] The evaluation standard can be preset, for example, picture clarity, picture background, picture style, picture content, etc.; an algorithm corresponding to the evaluation standard can be set, the target picture is input into the algorithm to output the evaluation parameter; or the target picture and reference information are input into the algorithm to output the evaluation parameter; wherein the reference information can include reference pictures, reference texts, etc.

[0080] The above preset parameter range can be set according to the evaluation standard and the evaluation parameter, for example, the higher the evaluation parameter of the picture clarity, the higher the clarity of the target picture, and the preset parameter range can be greater than 50, greater than 100, etc. For another example, the closer the picture background to the reference picture, the higher the evaluation parameter of the picture background, and the preset parameter range can be greater than 0.5, greater than 0.8, etc., and the closer the evaluation parameter to 1, the closer the picture background to the reference picture.

[0081] For the target picture whose evaluation parameter does not reach the preset parameter range, it needs to be deleted and regenerated through the foregoing picture generation model.

[0082] Further, the step of evaluating the target picture based on a preset evaluation standard to obtain an evaluation parameter includes at least one of the following:

[0083] (1) identifying picture background content of the target picture, determining an evaluation parameter of the picture background content based on the picture background content and a preset reference background content; the closer the picture background content and the reference background content, the higher the evaluation parameter of the picture background content; the picture background content can be a picture background color, a picture content color, etc.; taking the picture background color as an example, the reference background content is a reference background color, for example, white; after identifying the picture background color of the target picture, the closer the picture background color of the target picture and the reference background color, the higher the evaluation parameter of the picture background content.

[0084] In one example, the reference background color is white, if the picture background color of the target picture is white, the evaluation parameter of the picture background content is the highest, if the picture background color of the target picture is gray, the evaluation parameter of the picture background content is in the middle, and if the picture background color of the target picture is black, the evaluation parameter of the picture background content is the lowest.

[0085] In actual implementation, the edge line in the target picture can be identified first, the picture background area and the picture content area are determined based on the edge line, the picture background content such as the background color, the background texture, etc. is identified in the picture background area.

[0086] (2) obtaining the picture definition of the target picture, determining an evaluation parameter of the picture definition; the higher the picture definition, the higher the evaluation parameter of the picture definition;

[0087] The picture definition can be calculated by a preset definition algorithm, for example, the number of pixel points and the number of noise points in the target picture are obtained, the difference between the number of pixel points and the number of noise points is calculated, and then the ratio of the difference to the number of pixel points is calculated to obtain the picture definition. The picture definition can be directly used as the evaluation parameter of the picture definition, or the picture definition can be mapped according to a preset mapping relationship to obtain the evaluation parameter of the picture definition.

[0088] In one example, the picture definition of the target picture A is 0.9, and the picture definition of the target picture B is 0.8, then the evaluation parameter of the picture definition of the target picture A is greater than the evaluation parameter of the picture definition of the target picture B.

[0089] (3) identifying the picture style of the target picture, determining an evaluation parameter of the picture style based on the picture style and a preset reference style; the closer the picture style and the reference style, the higher the evaluation parameter of the picture style;

[0090] The target picture can be input into a pre-trained style recognition model to recognize the picture style of the target picture; the picture style can include ink painting style, retro style, simple sketch style, cartoon style, etc.; if the picture style is the same as the preset reference style, the evaluation parameter of the picture style is higher, and if the picture style is different from the preset reference style, the evaluation parameter of the picture style is lower.

[0091] For example, the reference style is a simple sketch style, if the picture style is also a simple sketch style, the evaluation parameter of the picture style is 1, and if the picture style is an ink painting style, not a simple sketch style, the evaluation parameter of the picture style is 0.

[0092] (4) The picture content of the target picture is recognized, and based on the picture content and the target text element, the evaluation parameter of the picture content is determined; the closer the picture content and the target text element, the higher the evaluation parameter of the picture content.

[0093] The target picture can be input into a pre-trained semantic recognition model to recognize the picture content of the target picture; for example, the target picture includes a flower pattern, and the recognized picture content is a text format 'flower'; the picture content is compared with the target text element, if the picture content and the target text element are the same, the evaluation parameter of the picture content is the highest; if the picture content and the target text element are partially the same, the evaluation parameter of the picture content is in the middle; if the picture content and the target text element are completely different, the evaluation parameter of the picture content is the lowest.

[0094] In actual implementation, when the evaluation criteria include multiple, each evaluation criterion is pre-set with a corresponding preset parameter range; for a target picture, each evaluation criterion corresponds to an evaluation parameter, when the evaluation parameter corresponding to each evaluation criterion meets the preset parameter range, it is determined that the evaluation parameter of the target picture meets the preset parameter range; if the evaluation parameter corresponding to any evaluation criterion does not meet the preset parameter range, it is determined that the evaluation parameter of the target picture does not meet the preset parameter range.

[0095] In another way, when the evaluation criteria include multiple, for a target picture, each evaluation criterion corresponds to an evaluation parameter, and the sum of the evaluation parameters corresponding to the multiple evaluation criteria is calculated, when the sum meets the preset parameter range, it is determined that the evaluation parameter of the target picture meets the preset parameter range; if the sum does not meet the preset parameter range, it is determined that the evaluation parameter of the target picture does not meet the preset parameter range.

[0096] In this way, the target picture can be evaluated by the evaluation criteria, and the target picture whose evaluation parameter does not meet the preset parameter range can be regenerated by the picture generation model; the picture quality of the target picture can be ensured to meet the requirements by the evaluation criteria, and the picture generation quality is further improved.

[0097] In one implementation, after outputting the target image corresponding to the target text element, it also includes at least one of the following:

[0098] (1) Remove the background content of the target image and update the preset baseline background content to the background content of the target image;

[0099] In practical implementation, edge detection algorithms can be used to identify edge lines in the target image, determine the background area and content area of ​​the image based on the edge lines, remove the background content in the background area, and then fill the background area with the target image, for example, with a white image.

[0100] (2) Identify the shadowed areas of the target image and remove the shadowed areas;

[0101] In actual implementation, the edge lines in the target image can be identified first. Then, areas with brightness below a preset brightness threshold can be identified around the edge lines. These areas are shadow areas. The content within the shadow area is removed, and the referenced area is filled with the background area outside the shadow area.

[0102] (3) Identify edge noise in the target image and remove edge noise;

[0103] In actual implementation, we can first identify the edge lines in the target image, identify noise points around the edge lines, delete the noise points, and fill the noise points with pixels from the background area.

[0104] (4) Adjust the thickness of the edge lines of the target image and smooth the edge lines.

[0105] The thickness of edge lines can be adjusted using morphological algorithms such as erosion and dilation. If the edge lines contain jagged edges, they can be smoothed using smoothing and filtering algorithms.

[0106] exist Figure 5 In the example, the background of the original target image contained color, and the background of the processed target image was replaced with a white background; the shadows and edge noise of the edges were removed, and the color and line thickness of the image content were adjusted.

[0107] By using the above methods to post-process the target image, the image quality and style of the target image are made consistent, thereby improving the stability and usability of the image quality.

[0108] In actual implementation, the above image position has pre-set position coordinates; the image position corresponding to the target text element is determined; the target image is placed on the position coordinates corresponding to the image position to obtain the image and text title corresponding to the text title.

[0109] The position coordinates can be coordinates of the position center of the picture position, or coordinates of a specified position such as the upper left corner or the lower right corner of the picture position. After the corresponding relationship between the picture position and the target text element is queried based on the foregoing corresponding relationship, the target picture corresponding to the target text element is placed at the position coordinates corresponding to the picture position.

[0110] Figure 6 As an example of a picture-text question, the question stem part includes three target pictures, and the option part includes five target pictures. Taking the target picture corresponding to option A as an example, the target text element corresponding to the target picture is a single palm, and the picture position corresponding to the target text element is picture position 4, that is, the picture position corresponding to option A. In the target picture corresponding to the target text element, a pattern including a single palm is generated, and the target picture is placed at picture position 4.

[0111] By the foregoing manner of this embodiment, a picture-text question with stable quality can be obtained at low cost, the generation cost of the picture-text question is reduced, and the picture-text question has wide applicability.

[0112] Referring to Figure 7 A schematic diagram of a picture-text question generation device is shown in the figure. The device includes:

[0113] A text question acquisition module 70 is configured to acquire a text question. The text question is in a text format.

[0114] A template determination module 72 is configured to identify a target text element from the text question, and determine a question display template corresponding to the text question based on the target text element. The question display template includes at least one picture position. The picture position is located at a specified position in the question display template, and the picture position is used to fill in a picture.

[0115] A corresponding relationship determination module 74 is configured to determine a corresponding relationship between the picture position and the target text element.

[0116] A picture output module 76 is configured to input the target text element into a preset picture generation model, and output a target picture corresponding to the target text element.

[0117] A picture-text question generation module 78 is configured to fill the target picture into the picture position corresponding to the target text element, and obtain a picture-text question corresponding to the text question.

[0118] The above-mentioned image-text question generation device obtains a text question; wherein the text question is in a text format; identifies a target text element from the text question, determines a question display template corresponding to the text question based on the target text element; wherein the question display template includes at least one picture position; the picture position is located at a specified position in the question display template, and the picture position is used to fill a picture; determines a corresponding relationship between the picture position and the target text element; inputs the target text element into a preset picture generation model to output a target picture corresponding to the target text element; and fills the target picture into the picture position corresponding to the target text element to obtain an image-text question corresponding to the text question.

[0119] In the above manner, after the target text element is identified from the text question, the target picture corresponding to the target text element is generated through the picture generation model, and then the target picture is filled into the question display template to obtain the image-text question; the picture generation model only needs to output the target picture corresponding to the target text element, thereby reducing the understanding cost and task complexity of the model and making the output picture content more accurate; at the same time, the target picture is filled into the specified position in the question display template, thereby making the picture layout more uniform, improving the effect stability of the picture element, and making the image-text question have higher practical application value.

[0120] The above-mentioned text question obtaining module is configured to: obtain a prompt word used to generate a text question; wherein the prompt word includes at least one of a knowledge point examined by the question, a logic requirement of the question, a reference question, and an output format; and input the prompt word into a preset reasoning model to output the text question.

[0121] The above-mentioned text question includes a plurality of text questions; each text question includes a stem part and an option part; wherein the stem part is associated with a stem identifier, and the stem identifier is used to indicate the position of the stem part in the text question; the option part is associated with an option identifier, and the option identifier is used to indicate the position of the option part in the text question; and the above-mentioned device further includes a question identification module configured to: identify the stem part and the option part belonging to the same text question according to the arrangement order of the stem identifier and the option identifier in the plurality of text questions.

[0122] The above-mentioned stem part is further provided with a stem format indication identifier; the stem format indication identifier is used to indicate the format of the stem part; and the above-mentioned template determination module is configured to: for each text question, identify a stem text element from the stem part based on the format indication identifier of the stem part; identify an option text element from the option part based on a structured format corresponding to the option part; and wherein the option text element and the stem text element constitute a target text element.

[0123] The template determination module is configured to determine a question display template corresponding to the text question based on the number of elements and / or the element distribution position of the target text element, wherein the number of positions of the picture positions in the question display template matches the number of elements of the target text element.

[0124] The picture output module is configured to fill the text question and the target text element into the preset prompt word template to generate a picture prompt word, and input the picture prompt word into the preset picture generation model to determine a picture belonging scene based on the text question by using the picture generation model, and output the target picture corresponding to the target text element based on the picture belonging scene.

[0125] The device further comprises an evaluation module configured to evaluate the target picture based on a preset evaluation standard to obtain an evaluation parameter, and if the evaluation parameter does not reach a preset parameter range, delete the target picture, and continue to execute the step of inputting the target text element into the preset picture generation model to output the target picture corresponding to the target text element until the evaluation parameter reaches the preset parameter range.

[0126] The evaluation module is configured to identify picture background content of the target picture, determine an evaluation parameter of the picture background content based on the picture background content and a preset reference background content, wherein the closer the picture background content and the reference background content, the higher the evaluation parameter of the picture background content; obtain picture definition of the target picture, and determine an evaluation parameter of the picture definition, wherein the higher the picture definition, the higher the evaluation parameter of the picture definition; identify picture style of the target picture, determine an evaluation parameter of the picture style based on the picture style and a preset reference style, wherein the closer the picture style and the reference style, the higher the evaluation parameter of the picture style; identify picture content of the target picture, and determine an evaluation parameter of the picture content based on the picture content and the target text element, wherein the closer the picture content and the target text element, the higher the evaluation parameter of the picture content.

[0127] The device further comprises a picture processing module configured to remove background content of the target picture, update the preset reference background content to the background content of the target picture, identify a shadow part of the target picture and delete the shadow part, identify edge noise of the target picture and delete the edge noise, and adjust thickness of an edge line of the target picture and perform smoothing processing on the edge line.

[0128] The picture position is provided with a position coordinate in advance, and the image-text question generation module is configured to determine a picture position corresponding to the target text element, and place the target picture on the position coordinate corresponding to the picture position to obtain an image-text question corresponding to the text question.

[0129] The embodiment also provides an electronic device, comprising a processor and a memory, the memory storing computer executable instructions capable of being executed by the processor, and the processor executes the computer executable instructions to implement the above-mentioned generation method of the title of the image-text.

[0130] Referring to Figure 8 The electronic device shown in the figure comprises a processor 100 and a memory 101, the memory 101 storing computer executable instructions capable of being executed by the processor 100, and the processor 100 executes the computer executable instructions to implement the above-mentioned generation method of the title of the image-text.

[0131] Further, Figure 8 The electronic device shown in the figure further comprises a bus 102 and a communication interface 103, and the processor 100, the communication interface 103 and the memory 101 are connected through the bus 102.

[0132] The memory 101 can contain a high-speed random access memory (RAM, Random Access Memory) and can also include a non-volatile memory, for example, at least one disk memory. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 103 (which can be wired or wireless), and the Internet, a wide area network, a local area network, a metropolitan area network, etc. can be used. The bus 102 can be an ISA bus, a PCI bus or an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 8 Only one bidirectional arrow is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0133] The processor 100 can be an integrated circuit chip with signal processing capability. In implementation, the steps of the above method can be completed by integrated logic circuits or instructions in the form of software in the processor 100. The processor 100 described above can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component. The disclosed methods, steps and logic block diagrams in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as a hardware coding processor for execution, or a combination of hardware and software modules in the coding processor for execution. The software module can be located in the random access memory, the flash memory, the read-only memory, the programmable read-only memory or the electrically erasable programmable memory, the register, or other mature storage mediums in the art. The storage medium is located in the storage 101, and the processor 100 reads the information in the storage 101, and combines the hardware to complete the steps of the method of the above embodiments.

[0134] The processor in the above electronic device can implement the following operations in the above generation method of the image-text question by executing computer executable instructions:

[0135] A generation method of an image-text question, obtaining a text question; wherein the text question is in a text format; identifying a target text element from the text question, determining a question display template corresponding to the text question based on the target text element; wherein the question display template includes at least one picture position; the picture position is located at a specified position in the question display template, and the picture position is used to fill in a picture; determining a corresponding relationship between the picture position and the target text element; inputting the target text element into a preset picture generation model to output a target picture corresponding to the target text element; filling the target picture into the picture position corresponding to the target text element to obtain an image-text question corresponding to the text question.

[0136] Obtaining a prompt word for generating a text question; wherein the prompt word includes at least one of a knowledge point examined by the question, a logic requirement of the question, a reference question and an output format; inputting the prompt word into a preset reasoning model to output the text question.

[0137] The text question includes multiple text questions; the text question includes a question stem part and an option part; the question stem part is associated with a question stem identifier, and the question stem identifier is used to indicate the position of the question stem part in the text question; the option part is associated with an option identifier, and the option identifier is used to indicate the position of the option part in the text question; according to the arrangement order of the question stem identifier and the option identifier in the multiple text questions, the question stem part and the option part belonging to the same text question are identified.

[0138] The question stem part is also provided with a question stem format indication identifier; the question stem format indication identifier is used to indicate the format of the question stem part; for each text question, the question stem text element is identified from the question stem part based on the question stem format indication identifier; the option text element is identified from the option part based on the structured format corresponding to the option part; wherein the option text element and the question stem text element constitute the target text element.

[0139] Based on the number and / or distribution position of the elements of the target text element, a question display template corresponding to the text question is determined; wherein the number of picture positions in the question display template matches the number of elements of the target text element.

[0140] The text question and the target text element are filled into a preset prompt word template to generate a picture prompt word; the picture prompt word is input into a preset picture generation model to determine the scene to which the picture belongs based on the text question by the picture generation model, and to output the target picture corresponding to the target text element based on the scene to which the picture belongs.

[0141] Based on the preset evaluation standard, the target picture is evaluated to obtain an evaluation parameter; if the evaluation parameter does not reach the preset parameter range, the target picture is deleted, and the step of inputting the target text element into the preset picture generation model to output the target picture corresponding to the target text element is continued until the evaluation parameter reaches the preset parameter range.

[0142] The picture background content of the target picture is identified, and based on the picture background content and the preset reference background content, the evaluation parameter of the picture background content is determined; wherein the closer the picture background content and the reference background content, the higher the evaluation parameter of the picture background content; the picture definition of the target picture is obtained, and the evaluation parameter of the picture definition is determined; wherein the higher the picture definition, the higher the evaluation parameter of the picture definition; the picture style of the target picture is identified, and based on the picture style and the preset reference style, the evaluation parameter of the picture style is determined; wherein the closer the picture style and the reference style, the higher the evaluation parameter of the picture style; the picture content of the target picture is identified, and based on the picture content and the target text element, the evaluation parameter of the picture content is determined; wherein the closer the picture content and the target text element, the higher the evaluation parameter of the picture content.

[0143] The background content of the target picture is removed, a preset reference background content is updated as the background content of the target picture, a shadow part of the target picture is identified and deleted, edge noise of the target picture is identified and deleted, and the thickness of the edge line of the target picture is adjusted and the edge line is smoothed.

[0144] The picture position is pre-set with a position coordinate, the picture position corresponding to the target text element is determined, the target picture is placed on the position coordinate corresponding to the picture position, and a graphic text question corresponding to the text question is obtained.

[0145] In the above manner, after the target text element is identified from the text question, the target picture corresponding to the target text element is generated by the picture generation model, and the target picture is filled into the question display template to obtain the graphic text question. The picture generation model only needs to output the target picture corresponding to the target text element, which reduces the understanding cost and task complexity of the model, makes the output picture content more accurate, and at the same time, the question display template fills the target picture into the specified position, which makes the picture layout more uniform, improves the effect stability of the picture element, and the graphic text question has higher practical application value.

[0146] The embodiment also provides a computer readable storage medium, which stores computer executable instructions. When the computer executable instructions are called and executed by a processor, the computer executable instructions cause the processor to implement the above-mentioned graphic text question generation method.

[0147] The computer executable instructions stored in the above-mentioned computer readable storage medium can implement the following operations in the above-mentioned graphic text question generation method by executing the computer executable instructions.

[0148] A graphic text question generation method, a text question is obtained; wherein the text question is in a text format; a target text element is identified from the text question, and a question display template corresponding to the text question is determined based on the target text element; wherein the question display template includes at least one picture position; the picture position is located at a specified position in the question display template, and the picture position is used to fill a picture; a corresponding relationship between the picture position and the target text element is determined; the target text element is input into a preset picture generation model, and a target picture corresponding to the target text element is output; the target picture is filled into the picture position corresponding to the target text element, and a graphic text question corresponding to the text question is obtained.

[0149] A prompt word used for generating a text question is obtained; wherein the prompt word includes at least one of a knowledge point examined by the question, a logic requirement of the question, a reference question, and an output format; the prompt word is input into a preset reasoning model, and a text question is output.

[0150] The text question includes multiple text questions; the text question includes a question stem part and an option part; the question stem part is associated with a question stem identifier, and the question stem identifier is used to indicate the position of the question stem part in the text question; the option part is associated with an option identifier, and the option identifier is used to indicate the position of the option part in the text question; according to the arrangement order of the question stem identifier and the option identifier in the multiple text questions, the question stem part and the option part belonging to the same text question are identified.

[0151] The question stem part is also provided with a question stem format indication identifier; the question stem format indication identifier is used to indicate the format of the question stem part; for each text question, the question stem text element is identified from the question stem part based on the question stem format indication identifier; the option text element is identified from the option part based on the structured format corresponding to the option part; wherein the option text element and the question stem text element constitute the target text element.

[0152] Based on the number and / or distribution position of the elements of the target text element, a question display template corresponding to the text question is determined; wherein the number of picture positions in the question display template matches the number of elements of the target text element.

[0153] The text question and the target text element are filled into a preset prompt word template to generate a picture prompt word; the picture prompt word is input into a preset picture generation model to determine the scene to which the picture belongs based on the text question by the picture generation model, and to output the target picture corresponding to the target text element based on the scene to which the picture belongs.

[0154] Based on the preset evaluation standard, the target picture is evaluated to obtain an evaluation parameter; if the evaluation parameter does not reach the preset parameter range, the target picture is deleted, and the step of inputting the target text element into the preset picture generation model to output the target picture corresponding to the target text element is continued until the evaluation parameter reaches the preset parameter range.

[0155] The picture background content of the target picture is identified, and based on the picture background content and the preset reference background content, the evaluation parameter of the picture background content is determined; wherein the closer the picture background content and the reference background content, the higher the evaluation parameter of the picture background content; the picture definition of the target picture is obtained, and the evaluation parameter of the picture definition is determined; wherein the higher the picture definition, the higher the evaluation parameter of the picture definition; the picture style of the target picture is identified, and based on the picture style and the preset reference style, the evaluation parameter of the picture style is determined; wherein the closer the picture style and the reference style, the higher the evaluation parameter of the picture style; the picture content of the target picture is identified, and based on the picture content and the target text element, the evaluation parameter of the picture content is determined; wherein the closer the picture content and the target text element, the higher the evaluation parameter of the picture content.

[0156] The background content of the target picture is removed, the preset reference background content is updated to the background content of the target picture, the shadow part of the target picture is identified and deleted, the edge noise of the target picture is identified and deleted, and the thickness of the edge line of the target picture is adjusted and the edge line is smoothed.

[0157] The picture position is provided with a position coordinate in advance, the picture position corresponding to the target text element is determined, and the target picture is placed on the position coordinate corresponding to the picture position to obtain a graphic text question corresponding to the text question.

[0158] In the above manner, after the target text element is identified from the text question, the target picture corresponding to the target text element is generated through the picture generation model, and then the target picture is filled into the question display template to obtain the graphic text question; the picture generation model only needs to output the target picture corresponding to the target text element, thereby reducing the understanding cost and task complexity of the model, making the output picture content more accurate; meanwhile, the target picture is filled into the specified position by the question display template, so that the picture layout is more uniform, the effect stability of the picture element is improved, and the graphic text question has higher practical application value.

[0159] The computer program product of the graphic text question generation method and the electronic equipment provided by the embodiment of the application includes a computer readable storage medium storing program codes, the program codes include instructions for executing the method described in the foregoing method embodiment, and specific implementation can be referred to the method embodiment, and details are not described herein.

[0160] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the system and the device described above can refer to the corresponding process in the foregoing method embodiment, and details are not described herein.

[0161] In addition, in the description of the embodiment of the application, unless otherwise explicitly specified and limited, the terms "mounting", "connection" and "connection" should be understood in a broad sense, for example, can be fixed connection, can also be detachable connection, or integral connection; can be mechanical connection, can also be electrical connection; can be directly connected, can also be indirectly connected through an intermediate medium, can be the communication between two elements. For those skilled in the art, the specific meaning of the above terms in the application can be understood according to the specific circumstances.

[0162] If the functions are realized in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application or the parts of the present application that essentially contribute to the prior art or the parts of the technical solutions can be embodied in the form of software products. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0163] In the description of the present application, it should be noted that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer" and the like indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the purpose of facilitating the description of the present application and simplifying the description, and do not indicate or imply that the devices or elements referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as limiting the present application. In addition, the terms "first", "second", "third" are only for the purpose of description, and cannot be understood as indicating or implying relative importance.

[0164] Finally, it should be noted that: the above embodiments are only specific embodiments of the present application, which are used to illustrate the technical solutions of the present application, and are not limited thereto, the protection scope of the present application is not limited thereto, although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art within the technical range disclosed by the present application can modify or easily think of changes to the technical solutions recorded in the foregoing embodiments, or make equivalent replacement to part of the technical features; and these modifications, changes or replacements do not make the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for generating graphic and textual titles, characterized in that, The method includes: Obtain the text title; wherein the text title is in text format; Identify target text elements from the text question, and determine a question display template corresponding to the text question based on the target text elements; wherein, the question display template includes at least one image position; the image position is located at a specified position in the question display template, and the image position is used to fill an image; Determine the correspondence between the image location and the target text element; The target text element is input into a preset image generation model, and the target image corresponding to the target text element is output. The target image is filled into the image position corresponding to the target text element to obtain the image and text title corresponding to the text title.

2. The method according to claim 1, characterized in that, The text questions include multiple parts; each text question includes a stem and an option part; wherein, the stem part is associated with a stem identifier, which indicates the position of the stem part in the text question; the option part is associated with an option identifier, which indicates the position of the option part in the text question; Before the step of identifying target text elements from the text questions, the method further includes: identifying the stem and option parts belonging to the same text question according to the arrangement order of the stem identifiers and option identifiers in the multiple text questions.

3. The method according to claim 2, characterized in that, The question stem section also includes a question stem format indicator; the question stem format indicator is used to indicate the format of the question stem section. The step of identifying target text elements from the text title includes: For each text question, the question stem text elements are identified from the question stem based on the format indicator of the question stem portion; Based on the structured format corresponding to the option section, option text elements are identified from the option section; wherein, the option text elements and the question stem text elements constitute the target text element.

4. The method according to claim 1, characterized in that, The steps for determining the question display template corresponding to the text question based on the target text element include: Based on the number of elements and / or the distribution position of the elements of the target text element, a question display template corresponding to the text question is determined; wherein, the number of positions of the image in the question display template matches the number of elements of the target text element.

5. The method according to claim 1, characterized in that, The steps of inputting the target text element into a preset image generation model and outputting the target image corresponding to the target text element include: The text title and the target text element are filled into a preset prompt template to generate an image prompt; The image prompt is input into a preset image generation model, which determines the scene to which the image belongs based on the text title and outputs the target image corresponding to the target text element based on the scene to which the image belongs.

6. The method according to claim 1, characterized in that, After the steps of inputting the target text element into a preset image generation model and outputting the target image corresponding to the target text element, the method further includes: The target image is evaluated based on preset evaluation criteria to obtain evaluation parameters; If the evaluation parameters do not reach the preset parameter range, delete the target image and continue to execute the steps of inputting the target text element into the preset image generation model and outputting the target image corresponding to the target text element, until the evaluation parameters reach the preset parameter range.

7. The method according to claim 6, characterized in that, The step of evaluating the target image based on preset evaluation criteria to obtain evaluation parameters includes at least one of the following: Identify the background content of the target image, and determine the evaluation parameters of the background content based on the background content and a preset benchmark background content; wherein, the closer the background content of the image is to the benchmark background content, the higher the evaluation parameters of the background content of the image. The image clarity of the target image is obtained, and the evaluation parameter of the image clarity is determined; wherein, the higher the image clarity, the higher the evaluation parameter of the image clarity; Identify the image style of the target image, and determine the evaluation parameters of the image style based on the image style and a preset benchmark style; wherein, the closer the image style is to the benchmark style, the higher the evaluation parameters of the image style. Identify the image content of the target image, and determine the evaluation parameters of the image content based on the image content and the target text element; wherein, the closer the image content and the target text element are, the higher the evaluation parameters of the image content.

8. The method according to claim 1, characterized in that, After the steps of inputting the target text element into a preset image generation model and outputting the target image corresponding to the target text element, the method further includes at least one of the following: Remove the background content of the target image and update the preset baseline background content to the background content of the target image; Identify the shadowed areas of the target image and delete the shadowed areas; Identify edge noise in the target image and remove the edge noise; Adjust the thickness of the edge lines of the target image and smooth the edge lines.

9. An electronic device, characterized in that, The method includes a processor and a memory, the memory storing computer-executable instructions that can be executed by the processor, the processor executing the computer-executable instructions to implement the method for generating graphic titles according to any one of claims 1-8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when invoked and executed by a processor, cause the processor to implement the method for generating graphic titles as described in any one of claims 1-8.