Image generation method, apparatus and system, and device, medium and vehicle
By identifying the user's image generation instructions, determining the content category and text expansion, the problem of low image quality in the vehicle is solved, and more accurately matching user needs and improving image quality is achieved.
Patent Information
- Application Number
- PCT/CN2024/140731
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-19
- Filing Date
- 2024-12-19
- Publication Date
- 2025-06-26
AI Technical Summary
In the prior art, the image generated in the vehicle is of low quality and cannot accurately match the content requirements in the user's image generation instructions.
By identifying the user's image generation instructions, the content category of the target image is determined, and the content rich information corresponding to the content category is used to expand the prompt words text, and finally convert the expanded text into the target image using the image generation model.
Through text augmentation, the image generation model understands the needs more accurately, thereby generating high-quality images that are more in line with user expectations.
Smart Images

Figure CN2024140731_26062025_PF_FP_ABST
Abstract
Description
Image generation method, device, system, equipment, medium and vehicle
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to the Chinese patent application filed with the China Patent Office on December 19, 2023, with application number 202311756165.7 and application name “Image Generation Method, Device, System, Equipment, Medium and Vehicle”, the entire contents of which are incorporated by reference into this application.
[0003] This application claims priority to the Chinese patent application filed with the China Patent Office on December 19, 2023, with application number 202311756157.2 and application name “Method, device, medium and equipment for optimizing prompt words for generating text images”, the entire contents of which are incorporated by reference into this application.
[0004] This application claims priority to the Chinese patent application filed with the China Patent Office on December 19, 2023, with application number 202311757386.6 and application name “Image Generation Method, Device, Equipment and Storage Medium”, the entire contents of which are incorporated by reference into this application.
[0005] This application claims priority to the Chinese patent application filed with the China Patent Office on December 19, 2023, with application number 202311757393.6 and application name “Image Optimization Method, Device, Equipment, Storage Medium and Vehicle”, the entire contents of which are incorporated by reference into this application. Technical Field
[0006] The present application belongs to the field of artificial intelligence technology, and in particular relates to an image generation method, apparatus, system, equipment, medium and vehicle. Background Art
[0007] Text-based image generation is a computer-generated task that aims to convert image generation instructions or natural language text into corresponding images. In this task, the image generation model in the computer needs to understand the user's image generation instructions and generate images that match the image generation instructions.
[0008] In the related art, algorithms related to the image generation model can be transplanted into the vehicle's controller. Then, the user can instruct the controller to generate a target image through simple image generation instructions in the vehicle's cabin, and display the target image on the vehicle's display screen or application. However, in the related art, the prompt words converted from the user's image generation instructions can be input into a pre-set image generation model, and the image generation model is used to convert the prompt words into an image that matches the prompt words. Since the image generation instructions proposed by the user are often relatively simple, they cannot clearly express the user's complete requirements for the target image, which can easily lead to the content of the generated target image not matching the content requirements indicated by the user in the image generation instructions. Therefore, the content of the target image generated in the related art often does not meet the user's requirements, resulting in low quality of the generated image. Summary of the Invention
[0009] Embodiments of the present application provide an image generation method, apparatus, system, device, medium, and vehicle, which can solve the problem of low quality of images generated in a vehicle.
[0010] In a first aspect, an embodiment of the present application provides an image generation method, the method comprising:
[0011] Obtaining an image generation instruction from a user, where the image generation instruction is used to instruct generation of a target image;
[0012] Identifying the image generation instruction, determining the content category of the target image based on the identification result, and obtaining content enrichment information corresponding to the content category;
[0013] Using the content enrichment information, the prompt word corresponding to the image generation instruction is expanded to obtain a final text;
[0014] The final text is converted into a target image using an image generation model corresponding to the content category.
[0015] In a second aspect, an embodiment of the present application provides an image generating device, comprising:
[0016] An acquisition part configured to acquire an image generation instruction from a user, wherein the image generation instruction is used to instruct generation of a target image;
[0017] an identification part configured to identify the image generation instruction, determine the content category of the target image based on the identification result, and obtain content enrichment information corresponding to the content category;
[0018] an expansion part configured to perform text expansion on the prompt word corresponding to the image generation instruction using the content enrichment information to obtain a final text;
[0019] The conversion part is configured to convert the final text into a target image by using an image generation model corresponding to the content category.
[0020] In a third aspect, an embodiment of the present application provides an image generation system, the system comprising:
[0021] The server is configured to obtain an image generation instruction from a user, where the image generation instruction is used to instruct generation of a target image;
[0022] The application side is configured to identify the image generation instruction, determine the content category of the target image based on the identification result, and obtain content enrichment information corresponding to the content category;
[0023] The application end is further configured to use the content enrichment information to perform text expansion on the prompt word corresponding to the image generation instruction to obtain a final text;
[0024] The application end is further configured to convert the final text into a target image using an image generation model corresponding to the content category.
[0025] In a fourth aspect, an embodiment of the present application provides an image generating device, the device comprising: a processor and a memory storing computer program instructions;
[0026] When the processor executes the computer program instructions, the above image generation method is implemented.
[0027] In a fifth aspect, an embodiment of the present application provides a computer storage medium having computer program instructions stored thereon, which implement the above-mentioned image generation method when the computer program instructions are executed by a processor.
[0028] In a sixth aspect, an embodiment of the present application provides a vehicle, which includes the above-mentioned image generating apparatus, image generating system, image generating device and storage medium.
[0029] In this application, the user's image generation instruction is identified, and then the content category of the target image generated by the image generation instruction is determined based on the identification result. The prompt word corresponding to the image generation instruction is expanded using the content enrichment information corresponding to the content category to obtain the final text. The final text is then converted into the target image using the image generation model corresponding to the content category. In this way, the expanded final text can more clearly express the user's complete requirements for image content, and the image generation model corresponding to the content category can more accurately understand the user's complete requirements for image content, thereby obtaining a target image that better meets the user's needs and improves the quality of the target image generated in the vehicle. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments of the present application. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0031] FIG1 is a flow chart of an image generation method according to an embodiment of the present application;
[0032] FIG2 is a schematic diagram of an exemplary display image art style provided by an embodiment of the present application;
[0033] FIG3 is a flowchart of an exemplary method for optimizing prompt words for generating images from text provided in an embodiment of the present application;
[0034] FIG4 is a schematic diagram of an exemplary text-to-image prompt word optimization process provided by an embodiment of the present application;
[0035] FIG5 is a flow chart of an exemplary image optimization method provided in an embodiment of the present application;
[0036] FIG6 is a flowchart diagram 1 of an exemplary image generation method provided by an embodiment of the present application;
[0037] FIG7 is a flowchart diagram 1 of an exemplary image generation method provided by an embodiment of the present application;
[0038] FIG8 is a second flow chart of an exemplary image generation method provided in an embodiment of the present application;
[0039] FIG9 is a second flowchart of an exemplary image generation method provided by an embodiment of the present application;
[0040] FIG10 is a third flowchart of an exemplary image generation method provided by an embodiment of the present application;
[0041] FIG11 is a schematic structural diagram of an image generating device provided in one embodiment of the present application;
[0042] FIG12 is a schematic diagram of the hardware structure of an image generating device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0043] The features and exemplary embodiments of various aspects of the present application will be described in detail below. In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than to limit the present application. For those skilled in the art, the present application can be implemented without the need for some of these specific details. The following description of the embodiments is merely to provide a better understanding of the present application by illustrating the examples of the present application.
[0044] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0045] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The embodiments will be described in detail below with reference to the accompanying drawings.
[0046] Specifically, in order to solve the problems of the prior art, the embodiments of the present application provide an image generation method, apparatus, system, device, medium and vehicle. The image generation method provided by the embodiments of the present application is first introduced below.
[0047] Figure 1 shows a flow chart of an image generation method provided by one embodiment of the present application. The method can be applied to a vehicle's onboard computer or a cloud server connected to the vehicle. The method includes the following steps:
[0048] S110 , obtaining an image generation instruction from a user, where the image generation instruction is used to instruct generation of a target image.
[0049] In this embodiment, the image generation method can be applied to a vehicle. The image generation instruction can be a voice instruction, a text instruction, or an image instruction. The voice instruction can be a user-speaking speech, the text instruction can be a user-entered text, and the image instruction can be a user-entered image.
[0050] For example, when the voice detection system in the vehicle is awakened, the voice detection system can be used to capture voice signals emitted by users inside or outside the vehicle, and then the collected voice signals can be analyzed using recognition technology to detect whether the voice signals contain voice commands. If voice commands are contained, the voice commands are obtained.
[0051] S120, identifying the image generation instruction, determining the content category of the target image based on the identification result, and obtaining content enrichment information corresponding to the content category;
[0052] In this embodiment, after receiving an image generation instruction, the image generation instruction can be recognized and converted into corresponding text content based on the recognition results. The text content can then be semantically understood. Specifically, the target image's prompt words can be filtered from the text content, and the content category of each item in the target image can be determined based on the prompt words. Content enrichment information associated with each content category can then be obtained, and this content enrichment information can be used to enrich and improve the prompt words.
[0053] Specifically, if the image generation instruction is a voice instruction, the voice instruction can be converted into text content corresponding to the image generation instruction. If the image generation instruction is a text instruction, the text instruction can be determined as the text content corresponding to the image generation instruction. If the image generation instruction is an image instruction, the description information of the image instruction can be obtained by extracting the features of the image instruction, and the description information of the image instruction can be determined as the text content corresponding to the image generation instruction.
[0054] As an optional embodiment, the identifying the image generation instruction and determining the content category of the target image based on the identification result includes:
[0055] Identifying the image generation instruction and obtaining text content corresponding to the image generation instruction;
[0056] Perform semantic understanding on the text content to obtain user intent; and extract prompt words for generating a target image from the user intent;
[0057] The content category of the target image is determined based on feature information in the prompt word.
[0058] In this embodiment, the content of the target image may include the object in the target image, the background in the target image, the image style of the target image, the main color of the target image, and the light and shadow effects of the target image.
[0059] Content with similar feature information can be determined to be of the same content category. Feature information refers to the attribute characteristics of the content generated by the prompt word, such as object attributes, style attributes, color attributes, etc. Specifically, if the content of the target image is an object within the target image, the content category of the target image is the object category, and the object category of the primary object within the target image is the subject category. For example, object categories can include people, scenery, animals, food, vehicles, etc. If the primary object within the prompt word is a cat or dog, the subject category of the prompt word is animals. If the primary object within the prompt word is coffee or bread, the subject category of the prompt word is food. If the content of the target image is the image style within the target image, the content category of the target image is the style category. Common style types can include realistic and artistic styles. Realistic styles can include realistic people, scenery, animals, architecture, etc., while artistic styles can include comics, oil paintings, ink paintings, line drawings, and other styles.
[0060] After determining the content category, the description words corresponding to the content category can be directly obtained as the content enrichment information, or the image generation model corresponding to the content category can be first determined, and then the keywords corresponding to the image generation model can be determined as the content enrichment information.
[0061] For example, if the content category is a subject category, the descriptive words corresponding to the subject category can be directly obtained as content enrichment information, or the basic image generation model corresponding to the subject category can be first determined, and then the subject keywords that are mapped to the basic image generation model can be queried as content enrichment information. If the content category is a style category, the descriptive words corresponding to the style category can be directly obtained as content enrichment information, or the specific object generation model corresponding to the style category can be first determined, and then the style keywords that are mapped to the specific object generation model can be queried as content enrichment information.
[0062] In the present embodiment, prompt words are used to describe the content of the target image that the user wants to generate. For example, the prompt words can be "generate a puppy" or "generate a big tree." The prompt words can include positive prompt words and negative prompt words. The positive prompt words mean the content that the target image is desired to include, while the negative prompt words indicate the content that the target image is not desired to include. For example, the positive prompt words can be "a cartoon-style Chinese girl running on the beach, with flying seagulls and a gorgeous rainbow behind her, the overall picture is poetic and picturesque," while the negative prompt words can be "pornographic, naked, ugly, deformed." Similarly, content enrichment information can also include positive enrichment information and negative enrichment information. The positive enrichment information is used to enrich the content of the positive prompt words, and the negative enrichment information is used to enrich the content of the negative prompt words.
[0063] For example, an ASR (Automatic Speech Recognition) algorithm can be used to convert image generation instructions into corresponding text form, i.e., text content. An LLM (Large Language Model) can also be used to perform semantic extraction on the text content in text form. This involves using this large language model to analyze the text content to extract semantic information from the text, and then determining the content category of the target image based on the semantic information, and obtaining content-rich information corresponding to the content category.
[0064] S130 , using the content enrichment information to perform text expansion on the prompt word corresponding to the image generation instruction to obtain a final text.
[0065] In this embodiment, the prompt word can be a relatively simple sentence extracted by the user based on the image generation instruction. In order to enrich the image content of the target image converted by the prompt word and improve the image quality, the prompt word can be expanded with the obtained content enrichment information to obtain a final text with richer content.
[0066] For example, the content-enriching information can be descriptive words related to the subject category of the target image. Descriptive words from a pre-set descriptive word library can be added to the prompt word to provide a detailed and extended description of the prompt word. For example, if the positive prompt word is "A cute, round-faced, chubby, anime-style tabby cat basking in the sun on a lawn covered with flowers," it can be expanded to produce the final text "A cute, round-faced, chubby, anime-style tabby cat basking in the sun on a lawn covered with flowers, honest and cute, with big eyes, charming scenery, and professional photography techniques."
[0067] As an optional embodiment, after using the content enrichment information to perform text expansion on the prompt word corresponding to the image generation instruction to obtain the final text, the method further includes:
[0068] The final text (i.e., the optimized prompt word) is reviewed according to a pre-established sample list including drawing texts; the sample list at least includes: drawing texts subject to conditional use and those not permitted to be used, and some of the drawing texts are provided with labels;
[0069] Using any prompt word in the final text as the current prompt word;
[0070] When the current prompt word matches the drawing text with a preset label in the sample list, adding the label of the drawing text to the current prompt word;
[0071] When the current prompt word matches a drawing text that is not allowed to be used in the sample list, deleting the current prompt word;
[0072] When the current prompt word matches the conditionally used drawing text in the sample list, the current prompt word is replaced.
[0073] In one embodiment, after obtaining the optimized prompt word, the method provided in the embodiment of the present application may further include: reviewing and editing the optimized prompt word to obtain the target prompt word. The editing includes but is not limited to: adding, replacing, deleting and / or correcting errors.
[0074] In one embodiment, editing includes: adding tags, deleting and / or replacing; accordingly, reviewing and editing the optimized prompt words may include: reviewing the optimized prompt words according to a pre-established sample list including drawing texts, and editing the optimized prompt words according to the review results. Among them, the hot-updated sample list is used to record different levels of drawing texts such as recommended use, conditionally restricted use, and not allowed use during the image generation process; drawing texts that are conditionally restricted use and not allowed use can be called negative sample texts, such as taboo texts involving illegal and irregular activities, and texts that are prohibited due to copyright management and other conditions. Recommended drawing texts can be called positive sample texts, such as label-type texts used to supplement detail modifiers such as attributes. Some of the above drawing texts can be pre-set with labels.
[0075] Based on the sample list including: drawing texts that are subject to conditional use and those that are not allowed to be used, the optimized prompt words can be edited according to the review results. Please refer to the following examples for details.
[0076] Any of the optimized prompt words is used as the current prompt word.
[0077] If the current prompt word matches a drawing text with a pre-set label in the sample list, the label of the drawing text is added to the current prompt word. For example, if the current prompt word is "dragon" and the user wants the Wensheng Diagram model to generate a Chinese dragon instead of a Western dragon by default, the drawing text describing a dragon in the sample list is labeled with the attribute "Chinese dragon". If the current prompt word matches a drawing text describing a dragon in the sample list, the label is added to the current prompt word, that is, the label "Chinese dragon" is added to the current prompt word "dragon".
[0078] If the current prompt word matches a drawing text that is not allowed in the sample list, the current prompt word will be deleted. If the current prompt word matches sensitive text such as nudity, pornography, violence, or politics in the sample list, the current prompt word will be deleted.
[0079] If the current prompt word matches a drawing text in the sample list that is subject to conditional use, the current prompt word is replaced. Regarding the understanding of conditional use, for example, for works that have applied for copyright registration, in order to avoid infringement, copyright-managed text can be included in the sample list. Copyright-managed text can include both text indicating that unauthorized use of the work is prohibited and text indicating that the work is authorized for use. It can be understood that the above copyright registration is only an example of conditional use, and there may be other situations in actual applications.
[0080] In one specific embodiment, the current prompt word is Mickey Mouse, and the matching conditionally restricted drawing text in the sample list is: "Mickey Mouse is an IP subject object whose use is prohibited without authorization." In this case, the current prompt word is replaced with a cartoon mouse without copyright restrictions. Alternatively, the current prompt word is Mickey Mouse, and the matching conditionally restricted drawing text in the sample list is: "Mickey Mouse is an IP subject object whose use is prohibited without authorization, and Jerry Mouse is an IP subject object whose use is authorized." In this case, the current prompt word is replaced with a Jerry Mouse with authorized use.
[0081] This embodiment can optimize prompt words in a more accurate and more conducive to model generation in real time and in a distributed manner through the sample list, and can also avoid the generation of illegal and irregular content to a certain extent.
[0082] S140: Convert the final text into a target image using an image generation model corresponding to the content category.
[0083] In this embodiment, after determining the content category, the image generation model corresponding to the content category can be queried based on the correspondence between the image generation model and the content category set in advance, and then the image generation model converts the final text into the target image.
[0084] As an optional embodiment, converting the final text into a target image using an image generation model corresponding to the content category includes:
[0085] Obtaining a basic image generation model and a specific object generation model corresponding to the content category;
[0086] Mounting the specific object generation model on the basic image generation model to obtain a fused image generation model, wherein the image generation model includes the basic image generation model, the specific object generation model, and the fused image generation model;
[0087] According to the task type corresponding to the image generation instruction, the final text is converted into a target image corresponding to target parameters using the fused image generation model, where the target parameters refer to a target size and / or target resolution corresponding to the task type.
[0088] In this embodiment, the image generation model may include a basic image generation model and a specific object generation model. The basic image generation model may be a base model in a stable diffusion image generation model, which may convert text content into a basic image, and the specific object generation model may be a low-rank adaptation of large language models (LoRA image generation model) in a stable diffusion image generation model, which may be used to generate images of certain specific styles or specific subjects. The specific object generation model cannot be used alone, and the specific object generation model needs to be mounted on the basic image generation model to assist the basic image generation model in realizing image generation.
[0089] After the specific object generation model is mounted on the basic image generation model to obtain the fused image generation model, a random noise image can be obtained, and the final text can be converted into a text embedding vector, which is then embedded into the fused image generation model loaded with the sampler to guide the image generation model to generate the first intermediate image.
[0090] In an image generation instruction, the user can not only express expectations for the image content of the target image, but also express expectations for the target parameters of the target image. Specifically, the target parameters of the target image can be specified by setting the task type of the target generation instruction. In other words, the correspondence between the task type and the image parameters can be set in advance. After determining the task type corresponding to the image generation instruction, the image parameters that have a corresponding relationship with the task type can be determined as the target parameters of the target image generated by the image generation instruction. Different task types can be distinguished based on the application scenario of the target image.
[0091] For example, a task type could be the vehicle's central control screen wallpaper, the vehicle's heads-up display wallpaper, or the vehicle's application background image. The target size of the central control wallpaper must match the central control screen's screen size, and the target resolution for the central control wallpaper can be 800x480. The target size of the heads-up wallpaper must match the size of the heads-up display, and the target resolution for the heads-up wallpaper can be 1024x600. The target size of the application's background image must match the size of the application's display frame, and the target resolution for the background image can be 1280x720.
[0092] After determining the final text and target parameters, the final text can be input into a pre-selected image generation model, and the image generation model can be used to convert the final text into a first intermediate image. The size and resolution of the first intermediate image are then adjusted according to the target parameters to obtain the final target image.
[0093] In this application, the user's image generation instruction is identified, and then the content category of the target image generated by the image generation instruction is determined based on the identification result. The prompt word corresponding to the image generation instruction is expanded using the content enrichment information corresponding to the content category to obtain the final text. The final text is then converted into the target image using the image generation model corresponding to the content category. In this way, the expanded final text can more clearly express the user's complete requirements for image content, and the image generation model corresponding to the content category can more accurately understand the user's complete requirements for image content, thereby obtaining a target image that better meets the user's needs and improves the quality of the target image generated in the vehicle.
[0094] As an optional embodiment, the identifying the image generation instruction and determining the content category of the target image based on the identification result includes:
[0095] Recognizing the image generation instruction, and determining a content phrase in the prompt word corresponding to the image generation instruction based on a recognition result, wherein the content phrase is used to instruct generation of a content element in the target image;
[0096] In the case where there are multiple content phrases in the prompt word, determining the content weight of the content element corresponding to each content phrase to obtain multiple content weights;
[0097] The content element with the highest content weight among the multiple content weights is determined as the main content of the target image, and the category of the main content is determined as the content category of the target image.
[0098] In this embodiment, the process of determining the content weight of the content element corresponding to each content phrase and obtaining multiple content weights includes: querying the content weight of the content element corresponding to each content phrase in a pre-set content weight table to obtain multiple content weights.
[0099] In this embodiment, the pre-set content weight table includes the content weight of each content element. The larger the content weight, the greater the influence of the content element corresponding to that content weight on the target image, that is, the more important the content element is in the target image. Therefore, the content element with the highest content weight can be determined as the main content of the target image.
[0100] Furthermore, since the content of the target image includes multiple dimensions, such as objects, background, image style, and color, the content category also includes multiple dimensions, such as subject category, background category, style category, color category, etc. The content element with the highest content weight in each dimension can be determined as the main content of the target image, and then the category of the main content in each dimension can be determined as the content category of the target image.
[0101] Specifically, the prompt word can be segmented to obtain multiple segmentation phrases, and then the semantics of each segmentation phrase can be identified. The semantic recognition results of each segmentation phrase can be used to determine whether each segmentation phrase indicates that a corresponding content element is generated in the target image. If the segmentation phrase indicates that a corresponding content element is generated in the target image, the segmentation phrase is determined to be the content phrase in the prompt word, and the dimension to which the content element belongs is further determined based on the semantic recognition results. For example, if the segmentation phrase "kitten" is present in the prompt word, semantic recognition can be performed on "kitten", and "kitten" can be determined to be a content phrase indicating that a content element is generated in the target image, and the dimension to which the kitten belongs can be further determined to be an object; if the segmentation phrase "animation" is present in the prompt word, semantic recognition can be performed on "animation", and animation can be determined to be a content phrase indicating that a content element is generated in the target image, and the dimension to which animation belongs can be further determined to be an image style.
[0102] For example, if the content category is the subject category, the content weight table is an object weight table. Object weights represent the degree of influence of the corresponding primary object on the target image. If the prompt word contains multiple objects to be generated, the object weights of each object to be generated can be queried in a pre-set object weight table to obtain multiple object weights. The object to be generated with the highest object weight among the multiple object weights is determined as the subject object of the target image, and the category of the subject object is determined as the subject category of the target image.
[0103] Therefore, if the prompt word includes multiple objects to be generated, the object weight of each object to be generated can be queried in a pre-defined object weight table, and then the object to be generated with the highest object weight is determined as the main object of the target image, and the category of the main object is the main category of the target image.
[0104] The object weights for each object to be generated in the object weight table can be pre-set based on the user's preferences for each object category. For example, the subject categories can be people, scenery, animals, food, vehicles, etc. If the object in the prompt word is a cat or dog, the object category of the prompt word is animals. If the object in the prompt word is coffee or bread, the object category of the prompt word is food. If the user prefers animals over food, the object weight of animals can be set to be greater than the object weight of food.
[0105] In addition, the object weight of each object to be generated in the object weight table can also be determined based on the user's language habits. For example, if the user is accustomed to saying the object to be generated with a higher object weight first in the prompt word, and then saying the object to be generated with a lower object weight, then if there are multiple objects to be generated in the prompt word, then the production object with the earliest word order position can be determined as the main object with the highest object weight. For example, in the case where the prompt word is "Generate me an image with a cat and bread". Among them, the category of "cat" is animal, and the category of "bread" is food. Since the word order position of "cat" is before "bread", it can be assumed that the object weight of "cat" is higher than "bread", that is, "cat" is determined as the main object, and the main category is animal.
[0106] In this way, the subject objects and subject categories of the subject objects in the prompt words can be accurately screened out.
[0107] For example, if the content category is a style category, then the content weight table is also a style weight table. Style weights are used to indicate the degree of influence of the corresponding style on the target image. If the prompt word includes multiple image styles, the style weights of each image style are queried in a pre-set style weight table to obtain multiple style weights. The image style with the highest style weight among the multiple style weights is determined as the primary style of the target image, and the category of the primary style is determined as the style category of the target image.
[0108] Then, in this embodiment, the pre-set style weight table includes the style weight of each image style. The larger the style weight, the greater the influence of the image style corresponding to the style weight on the target image, that is, the more important the image style is in the target image.
[0109] Similarly, the style weights of each image style in the style weight table can be set in advance based on the user's preferences for each style category. If the user prefers anime style to realistic style, and realistic style to abstract style, the style weight of anime style can be set to be greater than that of realistic style, and the style weight of realistic style can be set to be greater than that of abstract style.
[0110] In addition, the style weights of each image style in the style weight table can also be determined based on the user's language habits. The style weight of the image style directly expressed in the prompt word can be set to be greater than the style of the image style indirectly implied. For example, the prompt word can be "Help me generate an anime-style starry sky like Van Gogh." Among them, the directly expressed image style is "anime style." Similarly, the "oil painting style" indirectly implied in the prompt word can be obtained based on the participle phrase "Van Gogh." However, since the anime style is directly expressed and the oil painting style is indirectly implied, the style weight of the anime style is greater than that of the oil painting style.
[0111] Therefore, if the prompt word includes multiple image styles, the style weight of each image style can be queried in a pre-defined style weight table, and then the image style with the highest style weight is determined as the main style of the target image. The category of the main style is then the style category of the target image.
[0112] For example, image style generally refers to the visual appearance and stylistic features of an image. Common style types can be divided into two categories: realistic style and artistic style. Among them, the realistic style can include realistic characters, landscapes, animals, buildings, etc. The realistic style emphasizes the true expression of colors and strives to restore the true colors and light and shadow effects of the characters. Therefore, it has high requirements for image quality and is very demanding on the details of the image. The artistic style can include comics, oil paintings, ink paintings, line drawings and other styles, which focus more on the display of painting style and have lower requirements for image details than the realistic style. The weight of the realistic style can be greater than that of the artistic style.
[0113] In this way, the main styles in the prompt words and the style categories of the main styles can be accurately screened out.
[0114] As an optional embodiment, the identifying the image generation instruction and determining the content category of the target image based on the identification result includes:
[0115] Recognizing the image generation instruction, and determining a prompt word corresponding to the image generation instruction based on a recognition result;
[0116] Conducting a security review on the prompt word to determine whether the prompt word meets the preset text security specification conditions;
[0117] In a case where the prompt word meets the text safety specification condition, the content category of the target image is determined based on feature information of the prompt word.
[0118] In this embodiment, the image generation instruction can be converted into a corresponding text form, i.e., text content, through semantic recognition, and then prompt words related to the target image can be extracted from the text content. The content of the prompt words can be reviewed to determine whether the prompt words meet the pre-set text safety specification conditions. Only when the prompt words meet the text safety specification conditions can the content category of the target image be further determined and the prompt words can be expanded textually.
[0119] Specifically, a sensitive vocabulary database can be constructed in advance, which contains multiple keywords related to terrorism, pornography, politics, hot public opinion, etc., and then the content of the prompt word and the sensitive vocabulary database are matched with keywords. Once a sensitive word is matched in the prompt word, it is considered that the prompt word does not meet the preset text safety specification conditions, and the image generation instruction needs to be directly filtered out; if no sensitive word is matched in the prompt word, it can be considered that the prompt word meets the preset text safety specification conditions.
[0120] In this way, it can be ensured that only highly secure image generation instructions are executed, thereby ensuring the security of the target image.
[0121] As an optional embodiment, the image generation model includes a basic image generation model and a specific object generation model, the content category includes a subject category and a style category, and obtaining content enrichment information corresponding to the content category includes:
[0122] Obtaining a basic image generation model corresponding to the subject category of the target image and a specific object generation model corresponding to the style category of the target image;
[0123] Searching a pre-set subject mapping table for subject keywords that have a mapping relationship with the basic image generation model, and determining the subject keywords as content enrichment information corresponding to the subject category;
[0124] A style keyword having a mapping relationship with the specific object generation model is searched in a preset style mapping table, and the style keyword is determined as content enrichment information corresponding to the style category.
[0125] In this embodiment, the image generation model can include a base image generation model and a specific object generation model. Each base image generation model is used to implement a linguistic image task for one or more categories of subjects, and each specific object generation model is used to implement a linguistic image task for one or more categories of image styles. Therefore, the base image generation model corresponding to the subject category and the specific object generation model corresponding to the style category can be determined based on the subject category.
[0126] For the image generation model, a subject keyword mapping table associated with the basic image generation model can be pre-defined. In the subject keyword mapping table, each basic image generation model has a subject keyword with a mapping relationship with it. After determining the basic image generation model, the subject keyword with a mapping relationship of the basic image generation model can be retrieved and determined from the subject keyword mapping table, and these subject keywords can be added to the prompt words as content enrichment information.
[0127] For example, if the subject of the target image is a kitten, then the subject category is animal, and the determined basic image generation model is the basic image generation model corresponding to animals. Therefore, the subject keyword mapping table associated with the basic image generation model can be searched for subject keywords corresponding to "kitten," such as "furry, cute, lively," and so on. These terms can be used as content enrichment information. This enriches the content of the prompt words and, by extension, the content of the target image, making the target image more detailed and more accurately reflecting the user's needs.
[0128] For example, since each basic image generation model corresponds to at least one subject type and each specific object generation model corresponds to a style type, the subject keyword having a mapping relationship with the basic image generation model is used to describe the subject type corresponding to the basic image generation model, and the style keyword having a mapping relationship with the specific object generation model is used to describe the image style type corresponding to the specific object generation model.
[0129] Similarly, a style keyword mapping table associated with a specific object generation model can be pre-defined. In the style keyword mapping table, each specific object generation model has a style keyword with a mapping relationship with it. After determining the specific object generation model, the style keyword with a mapping relationship for the specific object generation model can be retrieved and determined from the style keyword mapping table, and these style keywords can be added to the prompt words as content-enriching information.
[0130] The embodiment of the present application adds keywords corresponding to the model to the prompt words, so that the text content of the prompt words is richer, and it is easier for the model to better understand the content category corresponding to the text content of the input model, thereby improving the accuracy of the generated image.
[0131] As an optional embodiment, obtaining content enrichment information corresponding to the content category includes: determining at least one candidate art style type that is different from the style category and matches the subject category; and determining vocabulary representing the candidate art style type as content enrichment information matching the style category.
[0132] For example, for the target subject (i.e., subject category) of the tiger shown in FIG2 , in addition to being able to adapt to the extracted target style type (i.e., subject category), other artistic style types can also be adapted. Based on this, this embodiment can supplement the artistic style with richer and more detailed vocabulary, thereby determining multiple candidate artistic style types such as realistic style, comic style, watercolor style, etc. that are different from the target style type and match the tiger, and determining the vocabulary representing each of the above candidate artistic style types as rich vocabulary that matches the target style type. Accordingly, the above-mentioned rich vocabulary representing the candidate artistic style types is used as a supplement to the initial prompt words, and is added to the initial prompt words to obtain the optimized prompt words. Furthermore, when using the prompt words optimized in terms of artistic style to generate images, a variety of different artistic styles can be adapted for the same target subject, resulting in rich and diverse artistic effects.
[0133] As an optional embodiment, before determining the content category of the target image based on the recognition result, the method further includes:
[0134] When detecting that a special word exists in the prompt word corresponding to the image generation instruction, searching for a common word that has a corresponding relationship with the special word;
[0135] The common words are used to replace the corresponding special words in the prompt words.
[0136] In this embodiment, special vocabulary includes regional slang, colloquialisms, classical Chinese poetry, proverbs, internet jargon, technical terms, and scientific terms. These special vocabulary may be commonly used linguistic expressions within a region or community, or they may be obscure industry terms, often with certain regional characteristics and limitations. Image generation models generally cannot understand these special vocabulary.
[0137] A special vocabulary mapping table may be pre-set. In the special vocabulary mapping table, each special vocabulary has a common vocabulary with a mapping relationship. The common vocabulary is an easy-to-understand and popular explanation of the special vocabulary.
[0138] Therefore, when it is detected that there is a special word in the prompt word, a general word having a mapping relationship with the special word can be searched in the special word mapping table, and then the corresponding special word in the prompt word is replaced with the general word.
[0139] For example, in a special vocabulary mapping table, special vocabulary can be pre-classified into multiple dimensions, such as regional special vocabulary, industry special vocabulary, and community special vocabulary. For example, "slang, colloquialisms, ancient poetry, and proverbs" can be identified as regional special vocabulary, while "professional terminology and scientific nouns" can be identified as industry special vocabulary. In this way, the special vocabulary table can be first searched for the dimension to which the special vocabulary belongs, and then the common vocabulary to which the special vocabulary has a mapping relationship can be further searched.
[0140] For example, when the prompt word is "A cute tabby cat with a chubby face and an anime style is basking in the sun on a lawn full of flowers", "hutouhunao" is a Chinese slang or colloquialism, that is, a special word. Therefore, the general word "round face, chubby" that has a first mapping relationship with "hutouhunao" can be queried, and the prompt word can be converted to "A cute tabby cat with a round face and chubby face and anime style is basking in the sun on a lawn full of flowers".
[0141] The embodiment of the present application converts obscure special words in the text content to be input into the model, so that the model can better understand the text content of the input model, thereby improving the accuracy of the generated image.
[0142] As an optional embodiment, converting the final text into a target image using an image generation model corresponding to the content category includes:
[0143] Obtaining a basic image generation model and a specific object generation model corresponding to the content category;
[0144] Mounting the specific object generation model on the basic image generation model to obtain a fused image generation model, wherein the image generation model includes the basic image generation model, the specific object generation model, and the fused image generation model;
[0145] According to the task type corresponding to the image generation instruction, the final text is converted into a target image corresponding to target parameters using the fused image generation model, where the target parameters refer to a target size and / or target resolution corresponding to the task type.
[0146] In this embodiment, the specific object generation model cannot be used alone. The specific object generation model needs to be mounted on the basic image generation model to assist the basic image generation model in achieving image generation.
[0147] As an optional embodiment, the process of converting the final text into a target image corresponding to the target parameters by using the fused image generation model includes:
[0148] Using the final text to guide the fused image generation model to convert the randomly generated noise image into a first intermediate image;
[0149] performing image editing on the first intermediate image to obtain a second intermediate image, wherein an image size of the second intermediate image matches a target size corresponding to the task type;
[0150] Based on the target resolution corresponding to the task type, image details of the second intermediate image are restored to obtain the target image.
[0151] After the specific object generation model is mounted on the basic image generation model to obtain the fused image generation model, a random noise image can be obtained, and the final text can be converted into a text embedding vector. The text embedding vector is embedded in the fused image generation model loaded with the sampler to guide the image generation model to perform multiple iterative denoising on the noise image to obtain the first intermediate image.
[0152] The first intermediate image can then be further edited based on the target format and target size specified in the predefined target parameters. Specifically, because the image format of the first intermediate image output by the image generation model is often different from the target format and the image size is also different from the target size, the first intermediate image must be edited to convert its format to the target format and adjust its size to the target size, thereby generating the second intermediate image.
[0153] For example, the target image can be a wallpaper displayed full-screen on the vehicle's central control display, or the background image of the vehicle's on-board control program. In order for the target image to adapt to the size and aspect ratio of the display or on-board control program display frame, the target image's target size needs to be specified in advance.
[0154] Furthermore, because the image generation model requires significant power consumption, and the vehicle's controller has limited computing power available to allocate to it, the second intermediate image generated by the model often suffers from poor resolution and detail. To address this, semantic segmentation can be performed on the second intermediate image, and the main region of the second intermediate image can be determined based on the results. Image detail restoration can then be performed within this region to ensure that the image resolution matches the target resolution specified in the target parameters, thereby obtaining the final target image.
[0155] For example, semantic segmentation can label pixels in an image as belonging to specific categories, thereby identifying and distinguishing different objects, structures, or regions in the second intermediate image. For example, the object located in the center of the second intermediate image and covering the largest area can be determined as the subject of the second intermediate image. The area containing the subject is then considered the subject area of the second intermediate image. Detail restoration is then performed only on the subject area.
[0156] For example, the second intermediate image may be semantically segmented into multiple regions using a SAM (Segment Anything Model) segmentation model, and the region with a central position and the largest coverage area among the multiple regions may be determined as the main region.
[0157] As an optional embodiment, converting the final text into a first intermediate image using an image generation model includes:
[0158] Get a noisy image;
[0159] Encoding the final text to obtain a text embedding vector;
[0160] De-noising the noisy image using the image generation model embedded with the text embedding vector to obtain a third latent feature image;
[0161] The third latent feature image is decoded to obtain the first intermediate image.
[0162] In this embodiment, in the process of converting text into an image, a randomly generated noise image can be obtained, and the text content of the prompt word can be mapped to a high-dimensional vector space to obtain a text embedding vector. After obtaining the text embedding vector, the text embedding vector and the noise image can be input into the fusion model together, and the text embedding vector can be used to guide the image generation model to perform denoising on the noise image to obtain a third latent feature image. The third latent feature image can then be decoded by using the autovariation encoder image generation model to obtain a target image that can be recognized by the naked eye.
[0163] In this embodiment, the conversion from the prompt word to the target image can be completed accurately.
[0164] As an optional embodiment, converting the final text into a target image using an image generation model corresponding to the content category includes:
[0165] Conducting a security review on the final text to determine whether the final text complies with preset text security specification conditions;
[0166] If the final text meets the text security specification condition, converting the final text into the target image using the image generation model corresponding to the content category;
[0167] If the final text contains sensitive words that do not meet the text security specification conditions, the sensitive words in the final text are deleted, and the final text with the sensitive words deleted is converted into the target image using the image generation model corresponding to the content category.
[0168] In this embodiment, after the image generation instruction is converted into the final text, the content of the final text can be reviewed to determine whether the final text meets the pre-set text security standard conditions. Only if the final text meets the text security standard conditions can the final text be converted into the target image. If the final text does not meet the text security standard conditions, sensitive words in the final text need to be deleted and the final text with the sensitive words deleted needs to be converted into the security standard conditions.
[0169] Specifically, a sensitive vocabulary lexicon can be constructed in advance, which contains multiple keywords related to terrorism, pornography, politics, hot public opinion, etc., and then the content of the final text is matched with the sensitive vocabulary lexicon for keywords. Once a sensitive word is matched in the final text, it is considered that the final text does not meet the preset text security specification conditions, and the sensitive word in the final text needs to be deleted; if no sensitive word is matched in the final text, it can be considered that the final text meets the preset text security specification conditions.
[0170] In this way, the privacy and security of the generated images can be guaranteed.
[0171] As an optional embodiment, the repairing of image details of the second intermediate image to obtain the target image includes:
[0172] Performing super-resolution processing on the second intermediate image using a preset super-resolution model to obtain an image to be restored;
[0173] Perform detail restoration on the main area of the image to be restored to obtain the target image.
[0174] In this embodiment, after obtaining, the trained super-resolution model is used to perform super-resolution processing on the entire low-resolution image to preliminarily improve the overall resolution of the low-resolution image and obtain the image to be repaired.
[0175] For example, the super-resolution model can be an Enhanced Super-Resolution Generative Adversarial Network (ESRGAN), which is used to perform 2X super-resolution on low-resolution images. Specifically, ESRGAN uses deep learning technology and a generative adversarial network to improve the spatial resolution of low-resolution images. 2X super-resolution can process low-resolution images and generate an image with twice the resolution of the low-resolution image in the horizontal and vertical directions, thereby improving the visual quality and details of the image. For example, the resolution of the low-resolution image is 512*512, and the resolution of the image to be restored after super-resolution processing is 1024*1024.
[0176] As an optional embodiment, the super-resolution model can be a residual in residual dense block (RRDB) module without batch normalization (BN). For example, the RRDB module may include three interconnected dense connection blocks (Dense Block), each of which consists of five convolutional layers. These convolutional layers may have different filters and feature map depths for learning representations of different levels of the image. In this way, the super-resolution model has better generalization.
[0177] In an embodiment of the present application, a low-resolution image (a second intermediate image) is obtained, wherein the low-resolution image is converted from a prompt word (i.e., an image generation instruction) input by a user; the low-resolution image is super-resolved using a pre-set super-resolution model to obtain an image to be repaired; and the main area of the image to be repaired is repaired in detail to obtain a target image. In other words, the prompt word input by the user can be first converted into a low-resolution image with relatively coarse details, and then the low-resolution image is super-resolved to obtain the image to be repaired, and the main area of the image to be repaired is repaired in detail to improve the quality of the image. In this way, compared with the prior art that consumes a lot of computing power to directly convert the prompt word into a picture with rich overall details, the prompt word can first be converted into a lower-quality image through lower computing power, and then the image is simply super-resolved, and only the main part of the image is repaired in detail, which can effectively reduce the cost of image generation while ensuring image quality.
[0178] In this embodiment, the details of the image to be restored after super-resolution are still relatively rough. The main area of the image to be restored can be determined in the image to be restored, and then the missing details in the main area can be restored or repaired, thereby improving the richness of the details in the main area. The main area is the main part of the image.
[0179] As an optional embodiment, performing detail restoration on the main area of the image to be restored to obtain the target image includes:
[0180] Performing semantic segmentation on the image to be repaired to determine a main area of the image to be repaired where details need to be repaired;
[0181] Adding noise to the main area of the image to be repaired to obtain an image to be processed;
[0182] According to the description information of the image to be repaired, noise reduction processing is performed on the main area of the image to be processed where noise is added, so as to obtain a target image after detail repair.
[0183] In this embodiment, semantic segmentation can label pixels in an image as belonging to specific categories, thereby identifying and distinguishing different objects, structures, or regions in the image to be repaired. Specifically, the region in the center of the image to be repaired and with the largest coverage area can be determined as the main region of the image to be repaired.
[0184] For example, the image to be repaired can be semantically segmented into multiple regions using a SAM (Segment Anything Model) segmentation model, and the region with the center position and the largest coverage area among the multiple regions is determined as the main region.
[0185] After determining the main area of the image to be repaired, a pre-set random number seed can be obtained, and the random number seed can be converted into an initialized Gaussian noise. Then, the Gaussian noise is added to each pixel in the main area to obtain the image to be processed with blurred noise in the main area.
[0186] After adding noise to the main area of the image to be processed, the noisy main area can be subjected to denoising to complete the image restoration. Specifically, the noisy image to be processed can be input into the image-generated model. The image-generated model can predict the noise on the image to be processed and use the loaded sampler to denoise the predicted noise, thereby obtaining the target image after detail restoration.
[0187] For example, the graph-based model may be a base model in a stable diffusion model.
[0188] Furthermore, the image description can be used to guide the image generation model in predicting and removing noise from the image to be processed. The image description can characterize the image content. Specifically, computer vision extraction techniques can be used to extract features of the image to be restored, and then these features can be converted into natural language descriptions to obtain the image description. Alternatively, user-entered prompts can be used directly as the image description.
[0189] In an embodiment of the present application, an image to be repaired is obtained, and semantic segmentation is performed on the image to be repaired, and the main area of the image to be repaired is determined based on the result of the semantic segmentation; noise is added to the main area of the image to be repaired; and noise reduction is performed on the main area of the image to be repaired to obtain a target image. In other words, the low-resolution image can be first converted into an image to be repaired with relatively coarse details, and then the main area of the image to be repaired can be determined by semantic segmentation of the image to be repaired, and only the main area can be noised and denoised to enrich the details of the main area in the image and improve the quality of the image. In this way, compared with the prior art that consumes a lot of computing power to directly convert the prompt words into a picture with rich overall details, the prompt words can first be converted into a lower-quality image to be repaired through lower computing power, and then only the main parts of the image to be repaired are repaired for details, which can effectively reduce the cost of image generation while ensuring image quality.
[0190] As an optional embodiment, performing semantic segmentation on the image to be repaired to determine a main area of the image to be repaired where details are to be repaired includes:
[0191] Performing semantic segmentation on the image to be repaired, and identifying at least one target object in the image to be repaired;
[0192] Determine the target object containing the largest number of pixels among the at least one target object as the main object;
[0193] The area where the main object is located is determined as the main area to be repaired in detail in the image to be repaired.
[0194] In this embodiment, semantic segmentation can be performed on the image to be restored, assigning a semantic label to each pixel in the image to be restored, such as "person," "vehicle," or "road." Each semantic label represents an object category. All connected pixels with the same semantic label can be identified as the same target object.
[0195] If there is only one target object in the image to be repaired, the target object can be directly determined as the main object in the image to be repaired; if there are multiple target objects in the image to be repaired, the target object with the largest number of pixels among the multiple target objects can be determined as the main object.
[0196] Then, the area where the main object is located can be determined as the main area in the image to be repaired. In this way, the main area can be accurately and quickly determined in the image to be repaired.
[0197] As an optional embodiment, the process of using the final text to guide the fusion image generation model to convert the randomly generated noise image into the first intermediate image includes:
[0198] An image to be processed and a first prompt word are obtained (i.e., the noise image is obtained and the first prompt word is obtained from the final text), wherein the first prompt word is used to determine the redrawing content. The global features of the image to be processed are extracted through a pre-constructed image control network, and the area to be processed of the image to be processed is determined, wherein the area to be processed includes a redrawing area and a non-redrawing area, the redrawing area refers to an expanded area of the image to be processed and / or an internal editing area of the image to be processed, and the non-redrawing area refers to at least a portion of the area in the image to be processed that does not need to be redrawn (i.e., the global features of the noise image are extracted through a pre-constructed image control network, and the area to be processed of the noise image is determined, wherein the area to be processed includes a redrawing area and a non-redrawing area, the redrawing area refers to an expanded area of the noise image and / or an internal editing area of the noise image, and the non-redrawing area refers to at least a portion of the area in the noise image that does not need to be redrawn). Redraw the area to be processed based on the first prompt word and the global features to generate a first image (i.e., redraw the area to be processed based on the first prompt word and the global features to generate a first intermediate image). Extract image information of the image to be processed, and redraw the first image based on the image information to generate a second image (the process of editing the first intermediate image to obtain the second intermediate image includes extracting image information of the noise image and redrawing the first intermediate image based on the image information to generate a second intermediate image). The image information includes edge information and / or segmentation information.
[0199] It is understandable that the image to be processed (noise image) and the first prompt word are obtained. The image to be processed can be understood as the original image, and the first prompt word refers to the content to be redrawn on the original image. There are two specific requirements for generating a redrawn image based on the original image: image expansion and internal image editing. Among them, image expansion is based on the original image, and the surrounding expansion area is generated according to the content of the original image itself, that is, the edge of the image is expanded; internal image editing is to select the redraw area in the original image according to the user's drawing idea, and redraw the redraw area, which can realize functions such as watermark removal and eraser.
[0200] It is understandable that the global features of the image to be processed are extracted through a pre-built image control network (ControlNet). The global features can also be understood as global information. ControlNet is a simple transfer learning method. The extracted global features can be used in the subsequent image-to-image process, that is, the information of the original image, such as the depth map, segmentation map, key points and other data, is used in the image-to-image process to control the newly generated image. Among them, the redrawing area refers to the expanded area of the image to be processed and / or the internal editing area of the image to be processed, and the non-redrawing area refers to at least part of the area in the image to be processed that does not need to be redrawn. Determine the redrawing area of the area to be processed. For image expansion requirements, the redrawing area can be an expanded area based on the expansion of the image to be processed. For the internal editing area of the image, the redrawing area can be a certain editing area in the image to be processed. For both requirements, the redrawing area includes the expanded area and a certain editing area. The editing area refers to generating new content in the area. At the same time, other areas in the image to be processed except a certain editing area can be understood as non-redrawing areas. After determining the redrawing area, the area to be processed is constructed based on the redrawing area and the non-redrawing area of the image to be processed. The non-redrawing area refers to at least a portion of the area in the image to be processed that does not need to be redrawn. The size of the non-redrawing area can be the same as the size of the image to be processed, that is, the entire image to be processed is the non-redrawing area, or it can be smaller than the size of the image to be processed, that is, part of the image to be processed is the non-redrawing area. The area to be processed can be understood as a mask area. For example, the edge of the non-redrawing area is expanded. In this case, the picture content of the area other than the non-redrawing area in the image to be processed may be changed.
[0201] Understandably, optionally, after obtaining the global features, an image-to-image method is used to redraw the area to be processed according to the first prompt word and the global features to generate a first image, which specifically includes the following contents.
[0202] It is understandable that a large model of Wenshengtu (Stable Diffusion 1.5, SD1.5 or Stable Diffusion XL, SDXL) can be used to control the redrawing content according to the first prompt word and global features / global information, and perform the first redraw to generate the first image. If the first prompt word is blank text, the image expansion can be achieved based on the global information. If the first prompt word is input text, the internal editing of the image can be achieved based on the first prompt word, and the image expansion can be achieved based on the global information. In other words, in one redrawing process, multiple redrawing functions can be achieved, including internal image editing, image expansion, and single redrawing. Other achievable redrawing functions and function types are not limited here. It is understandable that the first redrawing of the area to be processed according to the first prompt word and global features generates a first image with a relatively coherent picture and insignificant differences in picture content. Specifically, the Refiner Inpaint module in the SDXL model can be used for the first redrawing. If the first redrawing requirement is image expansion, the first prompt word can be blank text or input text without the original prompt word. The original prompt word refers to the description word related to the original screen content of the image to be processed, and the input text refers to the description word related to the newly added content of the image to be processed.
[0203] It is understandable that image information of the image to be processed is extracted through the image control network (ControlNet). The image information can be edge information and / or segmentation information. Subsequently, the image to be processed is redrawn a second time using the image generation method based on the image information to generate a second image. The second redrawing can be understood as a process of optimizing the first image generated by the first redrawing. It is understandable that the picture in the generated first image may still have a certain color difference and a clear sense of boundary at the edge. Therefore, the first image can be redrawn a second time based on the edge information and / or segmentation information, that is, the first image is optimized to eliminate the picture difference and the area of discontinuous edge.
[0204] As an optional embodiment, the process of obtaining the noise image and the first prompt word includes:
[0205] Acquire a noise image and extract a description text of the image in the noise image; use the description text and / or input text as the first prompt word, wherein the input text refers to the new content to be redrawn on the noise image; or
[0206] Acquire prompt text, and generate the noise image based on the prompt text using a text-to-image method; and use one or more texts among the prompt text, the input text, and blank text as the first prompt word.
[0207] It is understandable that if the image to be processed is directly obtained, the text description of the picture in the image to be processed is extracted through the neural network model to obtain the description text, wherein the neural network model can be a CLIP (Contrastive Language-Image Pre-Training) model or a multimodal model (Bootstrapping Language-Image Pre-training, BLIP). In this case, the first prompt word is the description text and / or input text, wherein the input text specifically refers to the new content to be redrawn on the image to be processed, for example, cats, dogs, etc. to be redrawn on the original image. Alternatively, if the prompt text (Prompt, or original prompt word) is obtained, the text image generation model is used to generate the image to be processed according to the prompt text. In this case, the first prompt word can be at least one of the prompt text, input text, and blank text. Blank text means that the first prompt word can be empty, and the surrounding area can be expanded according to the content of the picture itself.
[0208] As an optional embodiment, the image size includes an image length and an image width, the target size includes a target length and a target width, and the image editing the first intermediate image to obtain the second intermediate image includes:
[0209] If the image length of the first intermediate image is less than the target length, or the image width of the first intermediate image is less than the target width, cropping the first intermediate image to obtain an image to be edge-expanded, wherein the image length of the image to be edge-expanded is less than or equal to the target length, and the image width of the edge-expanded image is less than or equal to the target width;
[0210] Performing edge expansion on the image to be expanded using blank pixels to obtain an expanded image, wherein the image length of the expanded image is equal to the target length, and the image width of the expanded image is equal to the target width;
[0211] The expanded area of the expanded image is redrawn to obtain the second intermediate image, wherein the expanded area is the area filled with the blank pixels.
[0212] In this embodiment, if the image length of the first intermediate image is smaller than the target length, or the image width of the first intermediate image is smaller than the target width, it means that the image of the target size cannot be obtained by simply cropping the first intermediate image.
[0213] If the image length of the first intermediate image is less than the target length and the image width of the first intermediate image is less than the target width, the first intermediate image can be directly used as the image to be expanded. If the image length of the first intermediate image is greater than the target length, or the image width of the first intermediate image is greater than the target width, the first intermediate image needs to be cropped to obtain an image to be expanded whose image length is less than or equal to the target length and whose image width is less than or equal to the target width.
[0214] Specifically, the subject in the first intermediate image can be identified through image segmentation or target detection, and based on the results of target detection or image segmentation, the cropping area containing the subject can be determined, and the subject can be located in the middle of the cropping area as much as possible, and the length of the cropping area is smaller than the target length, and the width is smaller than the target width. Then, the cropping area can be cropped out in the first intermediate image, and the cropping area can be determined as the image to be expanded.
[0215] Among them, the target length can be the length of the display area corresponding to the inner screen of the vehicle-mounted screen, and the target width can be the width of the display area corresponding to the inner screen of the vehicle-mounted screen.
[0216] Subsequently, blank pixels with no content can be applied to the edges of the image to be expanded, thus expanding the edges of the image to the target size. The expanded area outside the image to be expanded is composed of blank pixels with no content. To improve image quality, the blank pixels in the expanded area can be redrawn according to user requirements to obtain a second intermediate image that matches the target size.
[0217] In this embodiment, by cropping and expanding the image, the image output by the image generation model can meet the target size required by the target image.
[0218] As an optional embodiment, redrawing the expanded area of the expanded image to obtain the second intermediate image includes:
[0219] Obtaining description information of the first intermediate image;
[0220] performing masking processing on the expanded image using a grayscale mask image;
[0221] adding noise to the expanded area according to the pixel value of each pixel point in the expanded area after the masking process;
[0222] performing noise reduction processing on the expanded area after the mask processing based on the description information of the first intermediate image to obtain a third intermediate image;
[0223] A transition region between the expanded region and the initial region in the third intermediate image is smoothed to obtain the second intermediate image, where the initial region is a region outside the transition region in the expanded image.
[0224] In this embodiment, the description information of the first intermediate image can represent the image content of the first intermediate image. Specifically, computer vision extraction technology can be used to extract features of the first intermediate image, and then these features can be converted into a natural language description to obtain the description information of the first intermediate image. Alternatively, the resulting text can be directly used as the description information of the first intermediate image.
[0225] Then, after determining the expanded area of the expanded image, a grayscale mask image can be masked on the expanded image. The portion of the grayscale mask image that masks the initial area is a black area with a pixel value of 0, and the pixel value of the portion of the grayscale mask image that masks the expanded area is not 0.
[0226] Then, by obtaining a preset random number seed, the random number seed is converted into an initialized Gaussian noise, and then the Gaussian noise is added to the expanded image. Since the pixel value of the initial area is 0, the initial area cannot be noised. Noise can be added to each pixel in the expanded area based on the pixel value of each pixel in the expanded area. The larger the pixel value, the more noise is added, which means the redrawing amplitude is larger.
[0227] After adding noise to the expanded area in the first intermediate image, the expanded area with the added noise can be subjected to noise reduction processing to complete the redrawing of the expanded area. Specifically, the description information of the expanded image with the added noise and the first intermediate image can be input into the image-generated model. The description information can guide the image-generated model to predict the noise on the expanded image, and the predicted noise can be denoised using the loaded sampler to complete the redrawing of the expanded image to obtain a third intermediate image. The transition area between the expanded area and the initial area in the third intermediate image is smoothed to obtain a second intermediate image.
[0228] For example, the graph-based model may be a base model in a stable diffusion model.
[0229] In this embodiment, the expanded area consisting of blank pixels can be denoised and redrawn with only a small amount of computing power, so that the image generated by the model can match the required target size.
[0230] As an optional embodiment, the step of smoothing the transition region between the expanded region and the initial region in the third intermediate image to obtain the second intermediate image includes:
[0231] adding noise to the transition region in the third intermediate image;
[0232] Noise reduction processing is performed on the transition region in the third intermediate image based on the description information of the first intermediate image to obtain the second intermediate image.
[0233] In this embodiment, after the expanded area of the first intermediate image is redrawn to obtain the third intermediate image, the visual connection or transition between the redrawn area and the unredrawn area in the third intermediate image may not be natural or smooth, and may have obvious edges, discontinuous transitions, or differences in color, brightness, etc.
[0234] By adding noise and denoising the transition area in the third intermediate image, the details of the transition area in the third intermediate image can be restored, so that the connection of the entire image is smooth and natural, thereby obtaining the second intermediate image.
[0235] For example, the transition area in the third intermediate image may be restored by performing denoising and noise addition on the entire third intermediate image.
[0236] In this embodiment, by adding noise and reducing noise on a portion of the redrawn image, the details of the image can be restored, so that the connection of the entire image is smooth and natural, thereby improving the image quality.
[0237] As an optional embodiment, the editing the first intermediate image to obtain the second intermediate image includes:
[0238] When the image length of the first intermediate image is greater than the target length and the image width of the first intermediate image is greater than the target width, the first intermediate image is cropped according to the target size to obtain the second intermediate image.
[0239] In this embodiment, if the length of the first intermediate image is greater than the target length and the width of the first intermediate image is greater than the target width, the first intermediate image output by the image generation model can be directly cropped so that the length of the cropped image is the target length and the width of the cropped image is the target width. Thus, the second intermediate image is directly obtained through cropping.
[0240] For example, the subject in the first intermediate image can be identified through image segmentation or target detection, and based on the results of target detection or image segmentation, a cropping area containing the subject can be determined, and the subject can be located as close as possible to the middle of the cropping area, and the size of the cropping area matches the target size. Then, the cropping area can be cropped out in the first intermediate image, and the cropping area can be determined as the second intermediate image.
[0241] In an embodiment of the present application, a larger image output by the image generation model can be cropped so that the cropped image meets the target size specified by the target image.
[0242] As an optional embodiment, the repairing of image details of the second intermediate image to obtain the target image includes:
[0243] Obtaining description information of the second intermediate image;
[0244] adding noise to a subject area of the second intermediate image;
[0245] Based on the description information of the second intermediate image, noise reduction processing is performed on the main body area with noise added in the second intermediate image to obtain the target image.
[0246] In this embodiment, the description information of the second intermediate image can represent the image content of the second intermediate image. Specifically, computer vision extraction technology can be used to extract features of the second intermediate image, and then these features can be converted into a natural language description to obtain the description information of the second intermediate image. Alternatively, the final text can be directly used as the description information of the second intermediate image.
[0247] After determining the main body area of the second intermediate image, a preset random number seed can be obtained, and the random number seed can be used to convert it into an initialized Gaussian noise, and then the Gaussian noise can be added to each pixel in the main body area.
[0248] After adding noise to the main area of the second intermediate image, the noisy main area can be subjected to denoising to enrich the image details. Specifically, the noisy second intermediate image can be input into the image-generated model. The image-generated model can predict the noise on the second intermediate image and denoise the predicted noise using the loaded sampler, thereby obtaining a target image with restored details.
[0249] For example, the graph-based model may be a base model in a stable diffusion model.
[0250] In this embodiment, after determining the main body area of the second intermediate image, only the main body area can be denoised and reduced, thereby enriching the details of the main body in the image with a faster processing speed through smaller computing power and power consumption, thereby improving the image quality.
[0251] As an optional embodiment, adding noise to the main area of the second intermediate image includes:
[0252] masking a non-subject area of the second intermediate image using a black mask image with a pixel value of 0, wherein the non-subject area is an area outside the subject area of the second intermediate image;
[0253] adding noise to the second intermediate image after masking;
[0254] The performing noise reduction processing on the main region of the second intermediate image containing noise based on the description information of the second intermediate image to obtain the target image includes:
[0255] encoding the second intermediate image after the masking process to obtain a first latent feature image;
[0256] performing noise reduction processing on the first latent feature image based on the description information of the second intermediate image to obtain a second latent feature image;
[0257] The second latent feature image is decoded to obtain the target image.
[0258] In this embodiment, after determining the subject region in the second intermediate image, the region outside the subject region can be determined as a non-subject region. Because the subject region in the second intermediate image is the region containing the subject, and the subject is the object that needs to be highlighted in the image, the subject region has higher requirements for image quality and image detail, while the non-subject region has relatively lower requirements for image quality and image detail.
[0259] To save computational power, only the main area in the second intermediate image needs to be repaired, while the non-main area does not need to be repaired. Therefore, a black mask image with a pixel value of 0 can be used to mask the non-main area, that is, all pixel values in the non-main area are set to 0.
[0260] After masking the non-subject area, when adding noise to the second intermediate image, the masked non-subject area will not be affected by the noise. That is to say, adding noise to the masked second intermediate image will only add noise to the main part of the second intermediate image.
[0261] By using the above method, it is possible to quickly and accurately add noise only to the main area in the second intermediate image.
[0262] By adding noise to the masked second intermediate image, a blurred image with noise is obtained. The blurred area represents the main area of the image. Subsequently, the noisy second intermediate image can be encoded using an autoencoder. This maps the noisy second intermediate image into the latent space, generating the first latent feature image.
[0263] Subsequently, the first latent feature image can be input into the base model in the stable diffusion model. Since the non-subject area has been masked, after loading the appropriate sampler in the base model, only the noise in the subject area can be predicted, and based on the predicted result, the subject area in the first latent feature image is denoised to obtain the second feature image.
[0264] After obtaining the second feature image, an autovariational encoder may be used to decode the second feature image to obtain a target image visible to the naked eye.
[0265] In this embodiment, the areas outside the main area that requires detail restoration and super-resolution can be masked so that the non-main area does not participate in the noise addition and noise reduction process, and only a small amount of computing power is required to complete the detail enrichment of the specified area.
[0266] As an optional embodiment, before repairing the image details of the second intermediate image to obtain the target image, the method further includes:
[0267] Using interpolation technology, the second intermediate image is expanded along the length direction and the width direction respectively;
[0268] The second intermediate image after image expansion is enhanced by a deblurring algorithm.
[0269] In this embodiment, since the operation of the image generation model requires a large amount of computing power, the resolution of the intermediate image output by the image generation model is usually low. The intermediate image can be super-resolution processed using a trained super-resolution model to improve the overall resolution of the intermediate image and obtain a clearer image.
[0270] For example, the super-resolution model can be an Enhanced Super-Resolution Generative Adversarial Network (ESRGAN), which performs 2X super-resolution on the intermediate image. Specifically, ESRGAN uses deep learning techniques and a generative adversarial network to improve the spatial resolution of the intermediate image. 2X super-resolution processes the second intermediate image to generate an image with twice the resolution in both horizontal and vertical directions, thereby improving the visual quality and detail of the image.
[0271] As an optional embodiment, the super-resolution model can be a residual in residual dense block (RRDB) module without batch normalization (BN). For example, the RRDB module may include three interconnected dense connection blocks (Dense Block), each of which consists of five convolutional layers. These convolutional layers may have different filters and feature map depths for learning representations of different levels of the image. In this way, the super-resolution model has better generalization.
[0272] As an optional embodiment, after repairing the image details of the second intermediate image to obtain the target image, the method further includes:
[0273] When the target image meets the preset image safety standard conditions, the target image is stored.
[0274] In this embodiment, after the final target image is generated, the target image needs to be reviewed to determine whether the target image meets the pre-set image security specification conditions. Only if the target image meets the image security specification conditions can the target image be stored according to the specified path; otherwise, the target image is filtered.
[0275] Specifically, a hash function can be used in advance to generate a hash value of an image and establish a hash database of sensitive content. The hash database contains multiple sensitive hash data involving terrorism, politics, pornography, nudity, hot public opinion, religion, specific IPs or characters, and then the hash database is used to match the target image. Once there is sensitive hash data matching the target image in the hash database, it is considered that the target image does not meet the preset image security specification conditions; if there is no sensitive hash data matching the target image in the hash database, it can be considered that the target image meets the preset image security specification conditions.
[0276] In this way, the privacy and security of the generated images can be guaranteed.
[0277] FIG3 is a flow chart of a method for optimizing prompt words for text-generated images provided in an embodiment of the present application. This method can be performed by a device for optimizing prompt words for text-generated images, which can be implemented in software and / or hardware, such as an electronic device or a server. The electronic device may include a vehicle-mounted processor, a mobile phone, a tablet computer, a desktop computer, a laptop computer, or other devices with communication functions. The server may be a cloud server or a server cluster, or other device with storage and computing functions. It should be noted that the following embodiments are explained exemplarily using electronic devices as the execution subject.
[0278] As shown in FIG3 , the method for optimizing prompt words for generating images from text includes the following steps:
[0279] S310: Obtain the initial prompt word input by the user.
[0280] In one embodiment, after the user inputs text into the Stable Diffusion model, the text input by the user, i.e., the initial prompt word, is obtained. The Stable Diffusion model is a model that diffuses in a latent space and includes: a text model, a UNet structure, and an encoder-decoder. The Stable Diffusion model is, for example, a Stable Diffusion model of version 1.5 or 2.0, or Stable Diffusion XL (SDXL for short), which is not limited here. The focus of this embodiment is on the optimization of the prompt word, so the above-mentioned Stable Diffusion model will not be described in detail.
[0281] S320: Extracting a target subject and a target style type based on the initial prompt word.
[0282] In some embodiments, the initial prompt word can be input into a pre-trained first text convolutional neural network (Text CNN); the target subject and target style type matched by the initial prompt word are extracted through the first text convolutional neural network.
[0283] The first text convolutional neural network may include: a convolutional module with a residual structure, a fully connected module, a first output module, and a second output module. Specifically, the convolutional module may be a six-layer one-dimensional convolution using a residual structure; the convolutional module is followed by a fully connected module, which includes three fully connected layers; and the fully connected module is followed by two branches: a first output module and a second output module. The implementation process for extracting the target subject and target style type matching the initial prompt word using the first text convolutional neural network can be referred to as follows.
[0284] The contextual text features of the initial prompt word are extracted through the convolution module, and the contextual text features are input into the fully connected module; the contextual text features are used to represent the semantic parsing information of the context of the current initial prompt word, the logical relationship between the initial prompt word and its context, and other information in the complete text input by the user. This embodiment uses a convolution module with a residual structure to extract the contextual text features of the initial prompt word. Since the residual structure can better utilize the contextual information by introducing jump connections, it can improve the expressive power of the contextual text features and thus improve the performance of the task. At the same time, the residual structure directly passes the information of the previous layer to the subsequent layer through jump connections, which can alleviate the gradient disappearance problem, making it easier for the gradient to propagate in the network and making it easier for the network to converge.
[0285] The context text features are sorted in a preset order through the fully connected module to obtain a text feature vector, and the text feature vector is input into the first output module and the second output module.
[0286] The first output module classifies the initial prompt word into subjects based on the text feature vector, obtaining multiple candidate subjects and the weight of each candidate subject. The weights of each candidate subject are then smoothed, and the target subject is determined based on the smoothed weights. Specifically, the first output module classifies the initial prompt word into subjects based on the text feature vector, obtaining multiple candidate subjects and the weight of each candidate subject. The subjects can be pre-set categories, such as people, animals, plants, mountains, and water, with these weights ranging from 0 to 1. The candidate subjects and their weights are then post-processed, including smoothing, to obtain the target subject that the initial prompt word ultimately matches.
[0287] The second output module classifies the initial prompt word by artistic style based on the text feature vector, obtaining multiple candidate style types and scores for each candidate style type. The scores for each candidate style type are then smoothed, and a target style type is determined based on the smoothed scores. Specifically, the second output module classifies the initial prompt word by artistic style based on the text feature vector, obtaining multiple candidate style types and scores for each candidate style type. The artistic styles can be pre-set categories, such as realistic, comic, or oil painting, with these scores ranging from 0 to 1. Post-processing, such as smoothing, is performed on the candidate style types and their weights to determine the target style type that the initial prompt word ultimately matches.
[0288] The first text convolutional neural network in this embodiment can not only accurately extract the target subject and target style type that match the initial prompt word, but also, due to the simplicity of the network design, the processing speed is fast, thereby improving the extraction efficiency of the target subject and target style type.
[0289] Considering the user's language habits, the initial prompt words entered by the user may contain special words that are difficult for machines to recognize, such as slang and common sayings.
[0290] Based on this, after the above step S320, the method provided in this embodiment may further include:
[0291] Match the initial prompt words with a preset text mapping table; wherein, the text mapping table is used to record the mapping relationship between specified words and standard words, and the specified words and the standard words are texts with the same semantics but different expressions;
[0292] When the initial prompt words match the target specified words in the text mapping table, map the target specified words in the initial prompt words to target standard words according to the mapping relationship, and obtain the initial prompt words after mapping.
[0293] In a specific embodiment, the specified words are generally words that are not easily understood by machines preset by users, such as slang, common sayings, ancient Chinese poems, proverbs, abbreviations, and Internet terms, etc.; this embodiment can determine at least one standard word with the same semantics as the specified word but different expressions according to the semantics expressed by the specified word. The standard word is a word with a more standard expression and conforming to the cognition of machines and most people; for example: the specified word "chubby" can be mapped to the following standard word: "round face, plump". Establish the mapping relationship between the specified word and the standard word, and record the specified word and the standard word with the mapping relationship into the text mapping table.
[0294] In the actual application of prompt word optimization, match the initial prompt words with a preset text mapping table. When the initial prompt words match the target specified words in the text mapping table, map the target specified words in the initial prompt words to target standard words. Exemplarily, the initial prompt words entered by the user are "A cute tabby cat with an anime style and a chubby appearance is sunbathing on a lawn full of flowers", and the "chubby" in it is a slang and common saying, which belongs to the specified words recorded in the text mapping table. Then, according to the mapping relationship, use the target standard word "round face, plump" to replace the target specified word "chubby", and convert the above initial prompt words into "A cute tabby cat with an anime style, a round face, and a plump appearance is sunbathing on a lawn full of flowers".
[0295] This embodiment uses the text mapping table to map some target specified words with special expressions and difficult to understand into target standard words with standard expressions and low understanding difficulty, which can reduce the difficulty of prompt word processing, improve the accuracy of prompt words, help the text-to-image model better understand the prompt words, avoid understanding deviations, and thus can generate the desired image more accurately.
[0296] S330: Determine a rich vocabulary that matches the target subject and / or target style type; wherein the rich vocabulary includes: vocabulary used to supplement the content of the initial prompt word.
[0297] In some embodiments, a text optimization framework can be derived based on different versions of text graph models such as Stable Diffusion and SDXL. This framework organizes the subject matter and artistic style types, providing multiple, supplementary, and more refined enriched vocabularies for both the prompt words representing the subject matter and the prompt words representing the artistic style type. Within this text optimization framework, enriched vocabularies matching the target subject matter and / or target style type are determined, thereby optimizing and supplementing the initial prompt words in terms of the subject matter and / or artistic style through these enriched vocabularies.
[0298] S340: Supplement the enriched vocabulary into the initial prompt word to obtain an optimized prompt word.
[0299] In this embodiment, the initial prompt words are supplemented with rich vocabulary that matches the target subject, and / or rich vocabulary that matches the target style type, so that the optimized prompt words can express the subject and artistic style more fully and detailedly, thereby achieving optimized processing in terms of the subject and artistic style.
[0300] For better understanding, the following embodiments describe the above steps S330 and S340 in detail.
[0301] In one embodiment, determining a rich vocabulary that matches the target style type may include: determining at least one candidate art style type that is different from the target style type and matches the target subject; and determining a vocabulary representing the candidate art style type as the rich vocabulary that matches the target style type.
[0302] For example, for the target subject of the tiger shown in Figure 2, in addition to being able to adapt to the extracted target style type, other artistic style types can also be adapted. Based on this, this embodiment can supplement the artistic style with a richer and more detailed vocabulary. Thus, multiple candidate artistic style types such as realistic style, comic style, watercolor style, etc. that are different from the target style type and match the tiger are determined, and the vocabulary representing each of the above candidate artistic style types is determined as a rich vocabulary that matches the target style type. Accordingly, the above rich vocabulary representing the candidate artistic style types is used as a supplement to the initial prompt words and is added to the initial prompt words to obtain the optimized prompt words. Furthermore, when using the prompt words optimized in terms of artistic style to generate images, multiple different artistic styles can be adapted for the same target subject, producing rich and diverse artistic effects.
[0303] In another embodiment, determining a rich vocabulary that matches the target style type may include: under the specified target style type, determining modification information used to describe the target subject, the modification information including: properties, states, characteristics and / or attributes; and determining the vocabulary representing the modification information as a rich vocabulary that matches the target subject.
[0304] Under certain target style types, the target subject's detailed modification information can be supplemented with richer and more detailed vocabulary. For example, for a realistic tiger, more detailed modification information can be determined to describe the tiger, such as: whether the tiger is a cub or an adult, whether the tiger is lying or standing, the tiger's fur color and stripes, whether the image containing the tiger is set to high definition and colorful, etc.; the vocabulary representing the above modification information is determined as a rich vocabulary that matches the target subject. In another example, for the target subject of a person, rich vocabulary representing modification information such as gender, clothing, and hair color can be determined. Accordingly, the rich vocabulary representing the modification information is used as a supplement to the initial prompt words and added to the initial prompt words to obtain optimized prompt words. Furthermore, when the optimized prompt words based on the target subject's detailed modification information are used to generate an image, the details of the image can be presented more accurately. At the same time, this embodiment can ensure that the added rich vocabulary can match the current target style type (such as realistic style) by supplementing the target subject with rich vocabulary that describes and modifies information under the specified target style type, avoiding the appearance of abstract, distorted and other words that are obviously contrary to the realistic style, high resolution and other words that are obviously contrary to the pixel painting style, etc.
[0305] It is understandable that the above embodiments can be combined with each other to perform more comprehensive, detailed and accurate optimization processing on the initial prompt word.
[0306] In another embodiment, the rich vocabulary also includes: vocabulary associated with a preset text graph model; determining the rich vocabulary matching the target subject and / or the target style type may include: determining a first text graph model matching the target subject and / or a second text graph model matching the target style type; determining the vocabulary associated with the first text graph model as the rich vocabulary matching the target subject; and determining the vocabulary associated with the second text graph model as the rich vocabulary matching the target style type.
[0307] This embodiment trains different models for specific subjects (such as certain vehicle brands) and specific artistic styles (such as Chinese artistic styles) to improve the performance of generated images. For example, for some subjects, the Lora (Low-Rank Adaptation, low-rank adaptation of large language models) model is matched as the first text-to-image model, and for some artistic styles, the Dreambooth model is matched as the second text-to-image generation model. After extracting the target subject and target style type, the matching Lora model will be mounted on the target subject, and the matching Dreambooth model will be mounted on the target style type.
[0308] Based on this, a first text-graph model matching the target subject is determined. This first text-graph model, for example, includes the aforementioned Lora model. Vocabulary associated with the first text-graph model is determined as a rich vocabulary matching the target subject, and this rich vocabulary is added to the initial prompt word. The resulting optimized prompt word is adapted to the first text-graph model.
[0309] Determine a second text-image model that matches the target style, such as the Dreambooth model described above. Determine the vocabulary associated with the second text-image generation module as a rich vocabulary that matches the target style, and add the rich vocabulary to the initial prompt words. The resulting optimized prompt words are adapted to the second text-image model.
[0310] In actual applications, some Lora models and Dreambooth models require specific vocabulary to activate and validate the model due to their inherent technical characteristics. Therefore, the initial prompt words need to be optimized and supplemented to ensure that the optimized prompt words are compatible with the model. For example, the initial prompt word is "Aa," which represents the model of a certain car. However, the Lora model cannot directly use "Aa" to generate an image of the corresponding vehicle. In this case, based on the Lora model's preset vocabulary "car," "Aa" is supplemented with the word "car" to obtain an effective prompt word that is compatible with the Lora model: "Aa, car." This prompt word is compatible with the Lora model and can be validated by the Lora model, thereby generating an image of the vehicle model Aa.
[0311] Furthermore, the Dreambooth model can generally be used directly for common artistic styles such as realism, oil painting, and comics. However, for specific and less common artistic styles such as the Dunhuang murals, the Dreambooth model generally requires additional vocabulary to be used. Based on this, this embodiment pre-configures the Dreambooth model with vocabulary that can activate these specific artistic styles, identifying them as enriched vocabulary that matches the target style. This enriched vocabulary is then optimized and supplemented with the initial prompt words, ensuring that the optimized prompt words are compatible with the Dreambooth model. Consequently, the Dreambooth model can use these optimized prompt words to generate images in the Dunhuang mural style.
[0312] According to the above embodiments, it is possible to optimize prompt words representing specific subjects and specific artistic styles, thereby increasing the applicable scope of the prompt words.
[0313] According to the above embodiment, after obtaining the optimized prompt word in step S340, the method provided in the embodiment of the present application may further include reviewing and editing the optimized prompt word to obtain the target prompt word. Editing includes but is not limited to adding, replacing, deleting, and / or correcting errors.
[0314] In one embodiment, editing includes: adding tags, deleting and / or replacing; accordingly, reviewing and editing the optimized prompt words may include: reviewing the optimized prompt words according to a pre-established sample list including drawing texts, and editing the optimized prompt words according to the review results. Among them, the hot-updated sample list is used to record different levels of drawing texts such as recommended use, conditionally restricted use, and not allowed use during the image generation process; drawing texts that are conditionally restricted use and not allowed use can be called negative sample texts, such as taboo texts involving illegal and irregular activities, and texts that are prohibited due to copyright management and other conditions. Recommended drawing texts can be called positive sample texts, such as label-type texts used to supplement detail modifiers such as attributes. Some of the above drawing texts can be pre-set with labels.
[0315] Based on the sample list including: drawing texts that are subject to conditional use and those that are not allowed to be used, the optimized prompt words can be edited according to the review results. Please refer to the following examples for details.
[0316] Any of the optimized prompt words is used as the current prompt word.
[0317] If the current prompt word matches a drawing text with a pre-set label in the sample list, the label of the drawing text is added to the current prompt word. For example, if the current prompt word is "dragon" and the user wants the Wensheng Diagram model to generate a Chinese dragon instead of a Western dragon by default, the drawing text describing a dragon in the sample list is labeled with the attribute "Chinese dragon". If the current prompt word matches a drawing text describing a dragon in the sample list, the label is added to the current prompt word, that is, the label "Chinese dragon" is added to the current prompt word "dragon".
[0318] If the current prompt word matches a drawing text that is not allowed in the sample list, the current prompt word will be deleted. If the current prompt word matches sensitive text such as nudity, pornography, violence, or politics in the sample list, the current prompt word will be deleted.
[0319] If the current prompt word matches a drawing text in the sample list that is subject to conditional use, the current prompt word is replaced. Regarding the understanding of conditional use, for example, for works that have applied for copyright registration, in order to avoid infringement, copyright-managed text can be included in the sample list. Copyright-managed text can include both text indicating that unauthorized use of the work is prohibited and text indicating that the work is authorized for use. It can be understood that the above copyright registration is only an example of conditional use, and there may be other situations in actual applications.
[0320] In one specific embodiment, the current prompt word is Mickey Mouse, and the matching conditionally restricted drawing text in the sample list is: "Mickey Mouse is an IP subject object whose use is prohibited without authorization." In this case, the current prompt word is replaced with a cartoon mouse without copyright restrictions. Alternatively, the current prompt word is Mickey Mouse, and the matching conditionally restricted drawing text in the sample list is: "Mickey Mouse is an IP subject object whose use is prohibited without authorization, and Jerry Mouse is an IP subject object whose use is authorized." In this case, the current prompt word is replaced with a Jerry Mouse with authorized use.
[0321] This embodiment can optimize prompt words in a more accurate and more conducive to model generation in real time and in a distributed manner through the sample list, and can also avoid the generation of illegal and irregular content to a certain extent.
[0322] In one embodiment, editing includes error correction; accordingly, reviewing and editing the optimized prompt words may include: inputting the optimized prompt words into a pre-trained second text convolutional neural network; and reviewing and removing erroneous conflicting prompt words in the optimized prompt words through the second text convolutional neural network.
[0323] For example, generating a pixel-style image requires a prompt word of "low resolution", while the globally optimized prompt words include prompt words such as "high resolution" and "high detail". Obviously, the "low resolution" required for the pixel style is a conflicting prompt word with "high resolution" and "high detail". Based on this, for conflicting prompt words, this embodiment outputs a value for each prompt word based on the second text convolutional neural network to review whether it is a conflicting word that needs to be deleted, and performs text vocabulary correction based on the output result, eliminating the incorrect conflicting prompt words in the optimized prompt words. Eliminating the conflicting prompt words can effectively avoid the problem of confusion caused by the text-based graph model.
[0324] In another embodiment, for error correction sentence editing, reviewing and editing the optimized prompt words can also be achieved in the following manner:
[0325] The optimized prompt words are input into the pre-trained Seq2Seq model; the context feature vector of the optimized prompt words is extracted through the encoder of the Seq2Seq model; the context feature vector is decoded by the decoder of the Seq2Seq model using the attention mechanism model, and the incorrect conflicting prompt words in the optimized prompt words are reviewed and eliminated based on the decoding results.
[0326] This embodiment uses a Seq2Seq model to learn to remove incorrect conflicting prompt words from the optimized prompt words. The Seq2Seq model consists of two parts: an encoder and a decoder. The encoder is used to obtain high-level features of the optimized prompt words, namely the context feature vector. The decoder uses an attention mechanism model, specifically an LSTM recurrent neural network model with attention mechanism. During decoding, it can generate corresponding representations based on the context feature vector, thereby more accurately reviewing and removing incorrect conflicting prompt words, achieving a higher accuracy rate.
[0327] Based on the above embodiments, this embodiment provides a prompt word optimization method for generating an image from text as shown in FIG4 , which specifically includes the following steps:
[0328] (1) Obtain the initial prompt word input by the user;
[0329] (2) extracting the target subject and target style type that match the initial prompt word through the first text convolutional neural network;
[0330] (3) determining at least one candidate artistic style type that is different from the target artistic style type and matches the target subject, and determining a vocabulary representing the candidate artistic style type as a rich vocabulary that matches the target artistic style type;
[0331] (4) determining, under the specified target style type, modification information for describing the target subject; and determining a vocabulary representing the modification information as a rich vocabulary that matches the target subject;
[0332] (5) Identify the vocabulary associated with the Lora model as a rich vocabulary that matches the target subject;
[0333] (6) Identify the vocabulary associated with the Dreambooth model as a rich vocabulary that matches the target style type;
[0334] (7) Supplementing the above enriched vocabulary into the initial prompt words to obtain optimized prompt words;
[0335] (8) Review the optimized prompt words according to the pre-established sample list and perform editing such as adding tags, deleting and replacing;
[0336] (9) The second text convolutional neural network is used to review and eliminate incorrect conflicting prompt words in the optimized prompt words.
[0337] The final target prompt word is obtained through the above steps.
[0338] In summary, the method for optimizing prompt words for text-generated images provided in the embodiment of the present application includes: first, obtaining the initial prompt words input by the user; second, extracting the target subject and target style type based on the initial prompt words; then determining the rich vocabulary that matches the target subject and / or the target style type; wherein the rich vocabulary includes: vocabulary for supplementing the content of the initial prompt words; and, supplementing the rich vocabulary to the initial prompt words to obtain the optimized prompt words. This technical solution will extract the target subject and target style type from the initial prompt words input by the user, and then perform enriched vocabulary supplementation and optimization on the prompt words in terms of the subject and style type of the image, so that the optimized prompt words can express more fully and detailed in terms of the subject and artistic style; and then, improve the accuracy of the prompt words by editing such as adding tags, deleting and replacing. Therefore, this solution can improve the richness, refinement and accuracy of the optimized target prompt words.
[0339] Furthermore, this application can provide an engineering solution for optimizing prompt words for different versions of the Stable Diffusion model and SDXL. The initial prompt words entered by the user are optimized in terms of subject and artistic style classification, and through review and editing such as adding tags, deletion and replacement, a certain degree of interception is performed on illegal and prohibited entries. The initial prompt words are optimized and supplemented according to the vocabulary associated with the literary graph model, and special optimization of prompt words can be performed for the generation of Chinese-style artistic styles and specific themes (such as a certain brand of car).
[0340] Figure 5 shows a flow chart of an image optimization method provided by one embodiment of the present application. The method can be applied to a vehicle's onboard computer or a cloud server connected to the vehicle. The method includes the following steps:
[0341] S510, acquiring a low-resolution image, wherein the low-resolution image is converted from a prompt word input by a user;
[0342] In this embodiment, the text-based graph model can be used to convert the prompt word input by the user into a low-resolution image, and the prompt word is used to describe the content of the low-resolution image that the user wants to generate. For example, the prompt word can be "generate a puppy" or "generate a tree."
[0343] Furthermore, the input prompt words can include positive prompt words and negative prompt words. Positive prompt words indicate desired content in the target image, while negative prompt words indicate undesirable content in the target image. For example, a positive prompt word could be "A cartoon-style Chinese girl running on the beach, with flying seagulls and a gorgeous rainbow behind her, the overall picture is poetic and picturesque," while a negative prompt word could be "pornographic, nude, ugly, deformed."
[0344] For example, the text graph model can be a base model in a stable diffusion model. A sampler with a faster iteration speed can be loaded into the base model, and the base model can convert the text content input into the model into a low-resolution image with less details and lower resolution at a faster speed.
[0345] S520, performing super-resolution processing on the low-resolution image using a preset super-resolution model to obtain an image to be restored;
[0346] In this embodiment, after obtaining, the trained super-resolution model is used to perform super-resolution processing on the entire low-resolution image to preliminarily improve the overall resolution of the low-resolution image and obtain the image to be repaired.
[0347] For example, the super-resolution model can be an Enhanced Super-Resolution Generative Adversarial Network (ESRGAN), which is used to perform 2X super-resolution on low-resolution images. Specifically, ESRGAN uses deep learning technology and a generative adversarial network to improve the spatial resolution of low-resolution images. 2X super-resolution can process low-resolution images and generate an image with twice the resolution of the low-resolution image in the horizontal and vertical directions, thereby improving the visual quality and details of the image. For example, the resolution of the low-resolution image is 512*512, and the resolution of the image to be restored after super-resolution processing is 1024*1024.
[0348] As an optional embodiment, the super-resolution model can be a residual in residual dense block (RRDB) module without batch normalization (BN). For example, the RRDB module may include three interconnected dense connection blocks (Dense Block), each of which consists of five convolutional layers. These convolutional layers may have different filters and feature map depths for learning representations of different levels of the image. In this way, the super-resolution model has better generalization.
[0349] S530: Perform detail restoration on the main area of the image to be restored to obtain a target image.
[0350] In this embodiment, the details of the image to be restored after super-resolution are still relatively rough. The main area of the image to be restored can be determined in the image to be restored, and then the missing details in the main area can be restored or repaired, thereby improving the richness of the details in the main area. The main area is the main part of the image.
[0351] In an embodiment of the present application, a low-resolution image is obtained, wherein the low-resolution image is converted from a prompt word input by a user; the low-resolution image is super-resolved using a pre-set super-resolution model to obtain an image to be repaired; and the main area of the image to be repaired is repaired to obtain a target image. In other words, the prompt word input by the user can first be converted into a low-resolution image with relatively coarse details, and then the low-resolution image can be super-resolved to obtain the image to be repaired, and the main area of the image to be repaired is repaired to improve the quality of the image. In this way, compared with the prior art that consumes a lot of computing power to directly convert the prompt word into an overall picture with rich details, the prompt word can first be converted into a lower-quality image using lower computing power, and then the image can be simply super-resolved, and only the main part of the image can be repaired. This can effectively reduce the cost of image generation while ensuring image quality.
[0352] As an optional embodiment, before performing super-resolution processing on the low-resolution image using a preset super-resolution model, the method further includes:
[0353] Obtaining a super-resolution model to be trained and a real image set, wherein the real image set includes multiple real images with the same image resolution;
[0354] The super-resolution model is trained using the real image set until the loss function of the super-resolution model converges.
[0355] In this embodiment, before the super-resolution model is applied to the super-resolution of low-resolution images, the model parameters of the super-resolution model need to be initialized first, and then the super-resolution model can be trained in sequence using multiple real images with the same resolution until the loss function of the super-resolution model converges to a certain value or stabilizes. The super-resolution model training can be considered completed.
[0356] As an optional embodiment, the training of the super-resolution model using the real image set until the loss function of the super-resolution model converges includes:
[0357] downsampling the real images in the real image set to obtain downsampled images;
[0358] Performing super-resolution processing on the downsampled image using the super-resolution model to obtain a super-resolution image, wherein the image resolution of the super-resolution image is the same as the image resolution of the real image;
[0359] Determining a predicted true probability that the super-resolved image is the true image;
[0360] Based on the predicted true probability, model parameters of the super-resolution model are adjusted until the predicted true probability is greater than or equal to a preset target probability threshold, and the loss function of the super-resolution model is determined to be converged.
[0361] In this embodiment, the super-resolution model can be a residual-in-residual dense block without batch normalization. During the training of the super-resolution model, a real image set consisting of multiple real images of the same resolution can be obtained, and then a fake image corresponding to each real image, i.e., a super-resolution image, can be obtained. The super-resolution image and the real image are then used to train the super-resolution model.
[0362] Specifically, taking the first real image as an example, the first real image can be downsampled to obtain a first downsampled image corresponding to the first real image. The resolution of the first downsampled image is lower than that of the first real image, but the first real image and the first downsampled image match in content.
[0363] Then, the super-resolution model can be used to super-resolve the first downsampled image to obtain a first super-resolved image. The first super-resolved image is a false image predicted from the first downsampled image to the first real image, and the resolution of the first super-resolved image is the same as that of the first real image. Subsequently, the first super-resolved image can be compared with the corresponding first real image to calculate the predicted true probability that the first super-resolved image is the real image. The discriminator loss and generator loss in the super-resolved model are calculated based on this predicted true probability, and the probability that the super-resolved image is the real image is increased through iterative training, thereby reducing the discriminator loss and generator loss and improving the performance of the model. Until the discriminator loss and generator loss converge, the super-resolved model is considered to have completed training.
[0364] Specifically, the true probability of the super-resolution model can be calculated by formula (1): Ra (x r ,x f )=σ(C(x r )-E xf [C(x f )]) (1)
[0365] Among them, x r is the first real image, x f The predicted first super-resolution image, C(·) is the original output of the discriminator before activation, E x (·) predicts the expectation of the distribution to which the super-resolved image belongs, and σ is a scaling factor.
[0366] As an optional embodiment, performing detail restoration on the main area of the image to be restored to obtain a target image includes:
[0367] Performing semantic segmentation on the image to be repaired to determine a main area of the image to be repaired where details need to be repaired;
[0368] Adding noise to the main area of the image to be repaired to obtain an image to be processed;
[0369] According to the description information of the image to be repaired, noise reduction processing is performed on the main area of the image to be processed where noise is added, so as to obtain a target image after detail repair.
[0370] In this embodiment, semantic segmentation can label pixels in an image as belonging to specific categories, thereby identifying and distinguishing different objects, structures, or regions in the image to be repaired. Specifically, the region in the center of the image to be repaired and with the largest coverage area can be determined as the main region of the image to be repaired.
[0371] For example, the image to be repaired can be semantically segmented into multiple regions using a SAM (Segment Anything Model) segmentation model, and the region with the center position and the largest coverage area among the multiple regions is determined as the main region.
[0372] After determining the main area of the image to be repaired, a pre-set random number seed can be obtained, and the random number seed can be converted into an initialized Gaussian noise. Then, the Gaussian noise is added to each pixel in the main area to obtain the image to be processed with blurred noise in the main area.
[0373] After adding noise to the main area of the image to be processed, the noisy main area can be subjected to denoising to complete the image restoration. Specifically, the noisy image to be processed can be input into the image-generated model. The image-generated model can predict the noise on the image to be processed and use the loaded sampler to denoise the predicted noise, thereby obtaining the target image after detail restoration.
[0374] For example, the graph-based model may be a base model in a stable diffusion model.
[0375] Furthermore, the image description can be used to guide the image generation model in predicting and removing noise from the image to be processed. The image description can characterize the image content. Specifically, computer vision extraction techniques can be used to extract features of the image to be restored, and then these features can be converted into natural language descriptions to obtain the image description. Alternatively, user-entered prompts can be used directly as the image description.
[0376] In an embodiment of the present application, an image to be repaired is obtained, and semantic segmentation is performed on the image to be repaired, and the main area of the image to be repaired is determined based on the result of the semantic segmentation; noise is added to the main area of the image to be repaired; and noise reduction is performed on the main area of the image to be repaired to obtain a target image. In other words, the low-resolution image can be first converted into an image to be repaired with relatively coarse details, and then the main area of the image to be repaired can be determined by semantic segmentation of the image to be repaired, and only the main area can be noised and denoised to enrich the details of the main area in the image and improve the quality of the image. In this way, compared with the prior art that consumes a lot of computing power to directly convert the prompt words into a picture with rich overall details, the prompt words can first be converted into a lower-quality image to be repaired through lower computing power, and then only the main parts of the image to be repaired are repaired for details, which can effectively reduce the cost of image generation while ensuring image quality.
[0377] As an optional embodiment, performing semantic segmentation on the image to be repaired to determine a main area of the image to be repaired where details are to be repaired includes:
[0378] Performing semantic segmentation on the image to be repaired, and identifying at least one target object in the image to be repaired;
[0379] Determine the target object containing the largest number of pixels among the at least one target object as the main object;
[0380] The area where the main object is located is determined as the main area to be repaired in detail in the image to be repaired.
[0381] In this embodiment, semantic segmentation can be performed on the image to be restored, assigning a semantic label to each pixel in the image to be restored, such as "person," "vehicle," or "road." Each semantic label represents an object category. All connected pixels with the same semantic label can be identified as the same target object.
[0382] If there is only one target object in the image to be repaired, the target object can be directly determined as the main object in the image to be repaired; if there are multiple target objects in the image to be repaired, the target object with the largest number of pixels among the multiple target objects can be determined as the main object.
[0383] Then, the area where the main object is located can be determined as the main area in the image to be repaired. In this way, the main area can be accurately and quickly determined in the image to be repaired.
[0384] As an optional embodiment, adding noise to the main area of the image to be repaired to obtain the image to be processed includes:
[0385] Determining a non-subject area in the image to be repaired, wherein the non-subject area is an area outside the main area in the image to be repaired;
[0386] Adjusting the pixel values of the pixels in the non-subject area of the image to be repaired to 0;
[0387] Obtain randomly initialized Gaussian noise, and superimpose the Gaussian noise on the main area of the image to be repaired after the pixel values are adjusted to obtain the image to be processed.
[0388] In this embodiment, after determining the main area of the image to be restored, the area outside the main area of the image to be restored can be determined as the non-main area. Since the main area of the image to be restored is the area where the main object is located, and the main object is the object that needs to be highlighted in the image, the main area has higher requirements for image quality and image details, while the non-main area has relatively lower requirements for image quality and image details.
[0389] To save computational power, only the main area of the image to be repaired needs to be repaired, while the non-main area does not need to be repaired. Therefore, the non-main area can be masked, and the pixel values of the main area are not adjusted, and the values of all pixels in the non-main area are set to 0.
[0390] After adjusting the pixel values of the non-subject area, when adding noise to the image to be repaired, the masked non-subject area will not be affected by the noise. In other words, adding noise to the masked image to be repaired will only add noise to the main part of the image to be repaired.
[0391] Through the above method, it is possible to quickly and accurately add noise only to the main area of the image to be repaired.
[0392] As an optional embodiment, performing noise reduction processing on a main region of the image to be processed where noise is added according to the description information of the image to be repaired to obtain a target image after detail repair includes:
[0393] Encoding the image to be processed to obtain an initial potential feature image;
[0394] performing multiple iterations of denoising on the initial latent feature image according to the description information of the image to be repaired, to obtain a repaired latent feature image;
[0395] The restored latent feature image is decoded to obtain the target image.
[0396] In this embodiment, after adding noise to the masked image to be inpainted, a blurred image with noise is obtained, namely the processed image. The blurred area in the processed image is the main area of the image. Subsequently, the noisy image to be processed can be encoded using an autoencoder, that is, the noisy image to be processed is mapped into the latent space to generate an initial latent feature image.
[0397] Subsequently, the initial latent feature image can be input into the base model in the stable diffusion model. Since the non-subject area has been masked, after loading the appropriate sampler in the base model, only the noise in the subject area can be predicted. Based on the predicted result, the subject area in the initial latent feature image is denoised to obtain the repaired latent feature image.
[0398] After obtaining the restored latent feature image, the restored latent feature image can be decoded using an autovariation encoder to obtain a target image visible to the naked eye.
[0399] In this embodiment, the pixel values of the areas outside the main area that requires detail restoration and super-resolution can be reset to zero so that the non-main area does not participate in the noise addition and noise reduction process, and only a small amount of computing power is required to complete the detail enrichment of the specified area.
[0400] As an optional embodiment, performing multiple iterations of denoising on the initial latent feature image according to the description information of the image to be restored to obtain the restored latent feature image includes:
[0401] Encoding the description information of the image to be repaired to obtain a text embedding vector;
[0402] The text embedding vector is embedded in an image generation model, and the image generation model embedded with the text embedding vector is used to perform multiple iterative noise reduction processes on the initial latent feature image to obtain the repaired latent feature image.
[0403] In this embodiment, before performing detail restoration on the image to be restored, a pre-trained image description generation model can be obtained. The image description generation model has been trained using a large-scale image corpus and can generate relatively accurate image description information.
[0404] After the image to be repaired is input into the image description generation model, it outputs a textual content that describes the main content of the image to be repaired. This description can include information such as objects, scenes, and actions. Therefore, the output textual content can be determined as the description information of the image to be repaired.
[0405] The text content of the descriptive information can be mapped to a high-dimensional vector space to obtain a text embedding vector. Then, the initial latent feature image and the text embedding vector can be input into the image generation model. The text embedding vector guides the image generation model to perform multiple iterations of denoising on the initial latent feature image to obtain the repaired latent feature image.
[0406] In this embodiment, the description information obtained by describing the image to be repaired can be used to guide accurate detail repair of the main area in the image to be repaired.
[0407] Figure 6 is a flow chart of an image generation method provided by an embodiment of the present application, which is applied to a server or terminal. In one possible application scenario, the server redraws the image to be processed for the first time and sends the generated first image to the terminal. The terminal redraws the received first image for the second time to generate a second image with a better redrawing effect. In another possible application scenario, the terminal or server performs the first redrawing of the image to be processed and the second redrawing of the first image on its own. Other possible application scenarios are not limited here. The following embodiment is described in detail using the server executing the image generation method as an example, specifically including the following steps S610 to S640 as shown in Figure 6:
[0408] S610: Obtain an image to be processed and a first prompt word.
[0409] The first prompt word is used to determine the redraw content.
[0410] It is understandable that the image to be processed and the first prompt word are obtained. The image to be processed can be understood as the original image, and the first prompt word refers to the content to be redrawn on the original image. Generating a redrawn image based on the original image specifically requires two requirements: image expansion and internal image editing. Image expansion is based on the original image, and the surrounding expansion area is generated according to the content of the original image itself, that is, the edge of the image is expanded; internal image editing is to select the redraw area in the original image according to the user's drawing ideas and redraw the redraw area, which can realize functions such as watermark removal and eraser.
[0411] Optionally, the above step S610 of obtaining the image to be processed and the first prompt word can be implemented by the following steps:
[0412] Obtain an image to be processed and extract descriptive text of the image in the image to be processed; use the descriptive text and / or input text as the first prompt word, wherein the input text refers to new content to be redrawn on the image to be processed; or obtain prompt text and generate the image to be processed based on the prompt text using a text-to-image method; use one or more texts among the prompt text, the input text, and blank text as the first prompt word.
[0413] It is understandable that if the image to be processed is directly obtained, the text description of the picture in the image to be processed is extracted through the neural network model to obtain the description text, wherein the neural network model can be a CLIP (Contrastive Language-Image Pre-Training) model or a multimodal model (Bootstrapping Language-Image Pre-training, BLIP). In this case, the first prompt word is the description text and / or input text, wherein the input text specifically refers to the new content to be redrawn on the image to be processed, for example, cats, dogs, etc. to be redrawn on the original image. Alternatively, if the prompt text (Prompt, or original prompt word) is obtained, the text image generation model is used to generate the image to be processed according to the prompt text. In this case, the first prompt word can be at least one of the prompt text, input text, and blank text. Blank text means that the first prompt word can be empty, and the surrounding area can be expanded according to the content of the picture itself.
[0414] S620: extracting global features of the image to be processed through a pre-built image control network, and determining a region to be processed of the image to be processed.
[0415] Among them, the area to be processed includes a redrawing area and a non-redrawing area, the redrawing area refers to the expanded area of the image to be processed and / or the internal editing area of the image to be processed, and the non-redrawing area refers to at least a part of the area in the image to be processed that does not need to be redrawn.
[0416] It is understandable that, based on the above S610, the global features of the image to be processed are extracted through a pre-built image control network (ControlNet). The global features can also be understood as global information. ControlNet is a simple transfer learning method. The extracted global features can be used in the subsequent image-to-image process, that is, the information of the original image, such as the depth map, segmentation map, key points and other data, is used in the image-to-image process to control the newly generated image. Among them, the redrawing area refers to the expanded area of the image to be processed and / or the internal editing area of the image to be processed, and the non-redrawing area refers to at least part of the area in the image to be processed that does not need to be redrawn. Determine the redrawing area of the area to be processed. For image expansion requirements, the redrawing area can be an expanded area based on the expansion of the image to be processed. For the internal editing area of the image, the redrawing area can be a certain editing area in the image to be processed. For both requirements, the redrawing area includes the expanded area and a certain editing area. The editing area refers to generating new content in the area. At the same time, other areas in the image to be processed except a certain editing area can be understood as non-redrawing areas. After determining the redrawing area, the area to be processed is constructed based on the redrawing area and the non-redrawing area of the image to be processed. The non-redrawing area refers to at least a portion of the area in the image to be processed that does not need to be redrawn. The size of the non-redrawing area can be the same as the size of the image to be processed, that is, the entire image to be processed is the non-redrawing area, or it can be smaller than the size of the image to be processed, that is, part of the image to be processed is the non-redrawing area. The area to be processed can be understood as a mask area. For example, the edge of the non-redrawing area is expanded. In this case, the picture content of the area other than the non-redrawing area in the image to be processed may be changed.
[0417] S630: Redraw the area to be processed according to the first prompt word and the global feature to generate a first image.
[0418] It is understandable that, based on the above S620, optionally, after obtaining the global features, an image-to-image method is used to redraw the area to be processed according to the first prompt word and the global features to generate a first image, which specifically includes the following contents.
[0419] It is understood that, based on the above-mentioned S910, a large model of Wenshengtu (Stable Diffusion 1.5, SD1.5 or Stable Diffusion XL, SDXL) is used to control the redrawing content based on the first prompt word and global features / global information, and perform a first redraw to generate a first image. If the first prompt word is blank text, image expansion can be achieved based on the global information. If the first prompt word is input text, internal image editing can be achieved based on the first prompt word, and image expansion can be achieved based on the global information. In other words, in a single redrawing process, multiple redrawing functions can be achieved, including internal image editing, image expansion, and a single redrawing function. Other achievable redrawing functions and function types are not limited here. It is understood that the first redrawing of the processed area based on the first prompt word and global features generates a first image with a relatively coherent picture and insignificant differences in picture content. Specifically, the Refiner Inpaint module in the SDXL model can be used for the first redrawing. If the first redrawing requirement is image expansion, the first prompt word can be blank text or input text without the original prompt word. The original prompt word refers to the description word related to the original screen content of the image to be processed, and the input text refers to the description word related to the newly added content of the image to be processed.
[0420] S640: Extract image information of the image to be processed, and redraw the first image according to the image information to generate a second image.
[0421] The image information includes edge information and / or segmentation information.
[0422] It is understandable that, based on the above S630, the image information of the image to be processed is extracted through the image control network (ControlNet), and the image information may be edge information and / or segmentation information. Subsequently, the image to be processed is redrawn a second time based on the image information using the image-to-image method to generate a second image. The second redrawing can be understood as a process of optimizing the first image generated by the first redrawing. It is understandable that the picture in the generated first image may still have a certain color difference and a clear sense of boundary at the edge. Therefore, the first image can be redrawn a second time based on the edge information and / or segmentation information, that is, the first image is optimized to eliminate the picture difference and the area with discontinuous edges.
[0423] Optionally, the step of redrawing the first image according to the image information to generate the second image in S640 may be implemented by the following steps:
[0424] The proportion of the image to be processed to the first image is calculated; if the proportion is less than a preset threshold, the first image is redrawn according to the image information and the prompt text of the image to be processed to generate a second image.
[0425] It is understandable that the ratio of the image to be processed to the first image is calculated. If the ratio is less than the preset threshold, that is, if the main body of the picture occupies less of the frame, the original prompt word (prompt text) is added during the second redrawing process to avoid more picture differences caused by the second redrawing.
[0426] Optionally, the step of redrawing the first image according to the image information to generate the second image in S640 may be implemented by the following steps:
[0427] The first image is redrawn according to the image information to generate a third image of a first size; the third image is redrawn according to the extracted image information and redrawing parameters of the third image to generate a second image of a second size; wherein the first size is smaller than the second size and larger than the size of the image to be processed.
[0428] It is understood that during the process of redrawing the image to be processed to generate a second image of a second size, a third image of the first size can be generated to further improve the generation effect. The third image can be understood as an intermediate image, with the first size being smaller than the second size and larger than the size of the image to be processed. The number of third images is not limited. Specifically: the first image is redrawn a second time based on the image information to generate the third image; the third image is redrawn a third time based on the image information and redrawing parameters to generate the second image, wherein the redrawing parameters may include information such as a third prompt word and the size of the generated image. The third prompt word can be at least one of input text, blank text, prompt text, and descriptive text. The image information can be information about the image to be processed or information about the third image. In one possible scenario, the first redrawing implements the image expansion function, the first prompt word is empty text, the second redrawing implements the image internal editing function, the third prompt word is input text, such as cat, the third redrawing performs the redrawing optimization process, and the second prompt word is empty text. In another possible scenario, the first redrawing implements the image expansion function, and the second and third redrawings perform the optimization process. Other possible generation processes are not limited here and can be determined according to user needs, that is, the drawing process of adjusting the image screen such as image expansion and / or image editing is based on global information, and the optimization process of redrawing the image is based on image information.
[0429] For example, refer to Figure 7, which is a structural diagram of an image generation method provided by an embodiment of the present application, based on the image expansion requirements implemented by SD1.5. Specifically, a 512*512 original image is obtained, and the original image is preliminarily expanded, for example, the original image is vertically expanded to 512*1024, and then the 512*1024 expanded image is expanded to determine the redrawing area. As shown in Figure 7, the redrawing area is the area expanded from 512*1024 to 1024*1024. The redrawing area and the target area where the original image is located are combined to obtain a 1024*1024 area to be processed. The 1024*1024 area to be processed is used as the input for the first redrawing to generate a 1024*1024 first image. Subsequently, the 1024*1024 first image is redrawn (optimized) for a second time to generate a 1024*1024 second image with a coherent picture.
[0430] For example, see Figure 8, which is a structural diagram of another image generation method provided by an embodiment of the present application. Based on the image expansion requirements implemented by SDXL, after obtaining the original prompt word Prompt (prompt text), the SDXL text image function is used to generate a 1024*1024 size original image according to Prompt, and Prompt is used as the first prompt word for the first redraw. In the case that the subject occupies a small portion of the frame, Prompt can also be added to the second prompt word for the second redraw to avoid the difference caused by redrawing. Alternatively, the original image of 1024*1024 size is obtained. After obtaining the initial image, the description text of the original image is obtained through CLIP or BLIP, and the description text is used as the first prompt word for the first redrawing. Subsequently, the redrawing area of the image to be processed is determined, and the redrawing area and the target area where the original image is located are spliced to obtain the original area. As shown in Figure 8, two 512*1024 redrawing areas and a 1024*1024 target area are spliced to obtain a 2024*1024 area to be processed. Subsequently, the 2024*1024 area to be processed is redrawn for the first time to generate the first image, and then the first image is optimized to generate the second image.
[0431] The image generation method provided in the embodiment of the present application performs a first redrawing of the image to be processed using prompt words to generate a first image, and then performs a second redrawing of the first image based on the image information of the image to be processed. The second redrawing is a process of continuing to optimize the first image based on edge information and / or segmentation information. While realizing redrawing-related functions such as image expansion and image element elimination, it optimizes the redrawing areas with discontinuous screen edge connections. Through the process of multiple redrawings, the redrawing accuracy is effectively improved, the screen differences are reduced, and a second image with better generation effect is obtained.
[0432] Based on the above embodiment, FIG9 is a schematic diagram of a detailed process of S630 provided in an embodiment of the present application. Optionally, redrawing the area to be processed based on the first prompt word and the global feature to generate a first image specifically includes the following steps S910 to S920 as shown in FIG9 :
[0433] S910: Set the redrawing amplitude of the redrawing area to a first amplitude value, and set the redrawing amplitude of the remaining areas of the to-be-processed area except the redrawing area to a second amplitude value.
[0434] The first amplitude value is greater than the second amplitude value, and the amplitude value of each area in the remaining areas gradually increases as it approaches the redrawing area, and the amplitude value of each area in the redrawing area gradually decreases as it approaches the remaining areas.
[0435] S920: Based on the first prompt word and the global feature, redraw the redraw area at the first amplitude value, and redraw the remaining area at the second amplitude value to generate a first image.
[0436] It can be understood that after obtaining the area to be processed, the redrawing amplitude of the redrawing area is set to the first amplitude value, and the redrawing amplitude of the remaining areas in the area to be processed except the redrawing area is set to the second amplitude value. When the redrawing area is an expanded area of the image to be processed, the remaining area is the target area. When the redrawing area is an edited area in the image to be processed, the target area is the complete area of the image to be processed, that is, the redrawing area is within the target area, so that the size of the combined area to be processed is the same as the size of the target area / image to be processed. In this case, the remaining area can be understood as the remaining area in the target area except the redrawing area, wherein the first amplitude value is greater than the second amplitude value. Preferably, the redrawing amplitude of the remaining areas of the Mask is controlled to be 0 to keep the image to be processed unchanged. The remaining areas are the areas where the image to be processed is located. The closer to the remaining areas, the smaller the redrawing amplitude, that is, the screen content of the remaining areas remains almost unchanged, and the closer to the redrawing area, the larger the redrawing amplitude.
[0437] The image generation method provided in the embodiment of the present application changes the redrawing position corresponding to the image to be processed and the redrawing amplitude corresponding to different positions through Mask. Different grayscale values of Mask can be used to blur the edges of the picture at the edges to perform accurate redrawing and picture continuity, thereby ensuring the redrawing effect.
[0438] Based on the above embodiment, FIG10 is a schematic diagram of a detailed process of S640 provided in an embodiment of the present application. Optionally, extracting image information of the image to be processed and redrawing the first image according to the image information to generate a second image specifically includes the following steps S101 to S103 as shown in FIG10 :
[0439] S101 . Extracting image information of the image to be processed through a pre-built image control network.
[0440] The image information includes edge information and / or segmentation information.
[0441] It can be understood that ControlNet is used to obtain image information of the image to be processed, and the image information includes edge information (Canny) or segmentation information (Segmentation).
[0442] S102: Perform image enhancement processing on the first image to obtain an enhanced image.
[0443] It is understandable that before optimizing the first image, in order to reduce the sense of style of the picture, the first image can be preprocessed. For example, the first image can be processed by image sharpening and other enhancement methods to enhance the sense of edge of the picture, so that more details can be retained after the subsequent image is generated. The specific enhancement method is not limited.
[0444] S103: Redraw the enhanced image according to the image information, and perform a preset number of samplings during the redrawing process to generate a second image.
[0445] It can be understood that based on the above S101 and S102, the image generation function of SDXL or SD1.5 is used, the second prompt word is set to empty, the redrawing amplitude of the redrawing area is set to 0.2, the enhanced image is optimized according to the image information, and the sampling number of the sampler is set to a preset number during the redrawing process to generate a second image with a coherent picture.
[0446] The preset number includes a first value and a second value, wherein the first value is smaller than the second value.
[0447] Optionally, the enhanced image is redrawn according to the image information and the second prompt word in S103, and a preset number of samplings are performed during the redrawing process to generate the second image. This can be specifically achieved by the following steps:
[0448] Obtain a second prompt word, wherein the second prompt word refers to blank text or prompt text of the image to be processed; redraw the enhanced image based on the image information and the second prompt word until the current sampling number obtained in real time during the redrawing process reaches the first value, and generate redrawing data; continue to redraw based on the redrawing data based on the second prompt word until the current sampling number obtained in real time during the redrawing process reaches the second value, and generate a second image.
[0449] It can be understood that, taking the preset number of 20 sampling steps as an example, image information is used to control screen generation and generate redrawing data starting from step 0. When the sampling step reaches the first value (50%), that is, step 10, the image information is stopped for control. From step 11, image information is not used for control, and automatic redrawing continues based on the redrawing data until the sampling step reaches step 20, thereby obtaining a second image with a coherent screen.
[0450] Preferably, the parameters for redrawing based on the image generation function of SDXL are set as follows: setting the Mask Blur to 15 can obtain a better edge blur fusion effect; when the first image is generated by the first redrawing, the masked area (Masked Content) selects the description text or the prompt text of the image to be processed as the first prompt word, and the prompt text can be at least part of the picture content in the image to be processed, that is, the description text and the prompt text can be consistent or different, and the DPM++2M Karras sampler is used for sampling, 2M refers to the second-order sampling scheme, Karras is a noise schedule, the main performance of which is that the noise step size will be smaller near the end, which helps to improve the quality of the image returned; in addition, the image generated during the redrawing process can only be in the latent space, and there is no need to generate an image visible to the human eye. This method helps to improve the redrawing speed. Then, a global optimization is performed on the first image that has been redrawn in the latent space. In the second redrawing process, the sampler remains unchanged, 20 sampling steps are used, the CFG Scale parameter is set to 7, and the redrawing amplitude is set between 0.3 and 0.4. The generated image is decoded into the original space using the SDXL image generation function to obtain a unified second image.
[0451] The image processing method provided in the embodiment of the present application reduces the sense of screen segmentation and retains more image details by enhancing the first image. Subsequently, while optimizing the enhanced image, a staged sampling method is used to control screen generation through image information in the first sampling stage and automatically generate the second image in the second sampling stage, thereby obtaining a second image with a coherent screen.
[0452] Based on the image generation method provided in the above embodiment, the present application also provides a specific implementation of an image generation device. Please refer to the following embodiment.
[0453] First, referring to FIG11 , the image generation device 200 provided in this embodiment of the application includes the following parts:
[0454] An acquisition part 201 is configured to acquire an image generation instruction from a user, where the image generation instruction is used to instruct generation of a target image;
[0455] The recognition part 202 is configured to recognize the image generation instruction, determine the content category of the target image based on the recognition result, and obtain content enrichment information corresponding to the content category;
[0456] An expansion part 203 is configured to perform text expansion on the prompt word corresponding to the image generation instruction using the content enrichment information to obtain a final text;
[0457] The conversion part 204 is configured to convert the final text into a target image using an image generation model corresponding to the content category.
[0458] The device can identify the user's image generation instruction, then determine the content category of the target image indicated by the image generation instruction based on the recognition result. It then uses the content enrichment information corresponding to the content category to expand the prompt word corresponding to the image generation instruction to obtain the final text. The final text is then converted into the target image using the image generation model corresponding to the content category. In this way, the expanded final text can more clearly express the user's complete requirements for image content, and the image generation model corresponding to the content category can more accurately understand the user's complete requirements for image content, thereby obtaining a target image that better meets the user's needs and improves the quality of the target image generated in the vehicle.
[0459] As an implementation of the present application, the identification part 202 may further include:
[0460] a first recognition unit configured to recognize the image generation instruction and obtain text content corresponding to the image generation instruction;
[0461] An understanding unit configured to perform semantic understanding on the text content to obtain user intention; and extract prompt words for generating a target image from the user intention;
[0462] The first determining unit is configured to determine the content category of the target image based on feature information in the prompt word.
[0463] As an implementation of the present application, the content category includes a subject category, and the identification part 202 may further include:
[0464] a second determining unit configured to determine a content phrase in the prompt word based on the user's intention, the content phrase being used to indicate the generation of a content element in the target image; and, if there are multiple content phrases in the prompt word, determine a content weight of a content element corresponding to each content phrase to obtain multiple content weights;
[0465] The third determining unit is configured to determine the content element with the highest content weight among the multiple content weights as the main content of the target image, and determine the category of the main content as the content category of the target image.
[0466] As an implementation of the present application, the identification part 202 may further include:
[0467] a second recognition unit configured to recognize the image generation instruction and determine a prompt word corresponding to the image generation instruction based on a recognition result;
[0468] a first review unit configured to perform a security review on the prompt word to determine whether the prompt word meets a preset text security specification condition;
[0469] The fourth determining unit is configured to determine the content category of the target image based on feature information of the prompt word when the prompt word meets the text safety specification condition.
[0470] As an implementation of the present application, the image generation model includes a basic image generation model and a specific object generation model. The acquisition part 201 may further include:
[0471] A first acquisition unit is configured to acquire a basic image generation model corresponding to a subject category of the target image and a specific object generation model corresponding to a style category of the target image;
[0472] a second query unit configured to query a preset subject mapping table for subject keywords that have a mapping relationship with the basic image generation model, and determine the subject keywords as content enrichment information corresponding to the subject category;
[0473] The third query unit is configured to query a preset style mapping table for style keywords that have a mapping relationship with the specific object generation model, and determine the style keywords as content enrichment information corresponding to the style category.
[0474] As an implementation of the present application, the obtaining portion 201 may further include:
[0475] The fifth determining unit is configured to determine at least one candidate art style type that is different from the style category and matches the subject category; and determine a vocabulary representing the candidate art style type as content enrichment information matching the style category.
[0476] As an implementation of the present application, the image generating device 200 may further include:
[0477] A query part configured to query common words that have a corresponding relationship with the special words when detecting that the special words exist in the final text;
[0478] The replacement part is configured to use the common vocabulary to replace the corresponding special vocabulary in the final text.
[0479] As an implementation of the present application, the conversion portion 204 may further include:
[0480] a second acquiring unit configured to acquire a basic image generation model and a specific object generation model corresponding to the content category;
[0481] a mounting unit configured to mount the specific object generation model onto the basic image generation model to obtain a fused image generation model, wherein the image generation model includes the basic image generation model, the specific object generation model, and the fused image generation model;
[0482] The first conversion unit is configured to convert the final text into a target image corresponding to a target parameter using the fused image generation model according to the task type corresponding to the image generation instruction, where the target parameter refers to a target size and / or target resolution corresponding to the task type.
[0483] As an implementation of the present application, the first conversion unit may further include:
[0484] a first conversion subunit, configured to use the final text to guide the fused image generation model to convert the randomly generated noise image into a first intermediate image;
[0485] an editing subunit configured to perform image editing on the first intermediate image to obtain a second intermediate image, wherein an image size of the second intermediate image matches a target size corresponding to the task type;
[0486] The restoration subunit is configured to restore the image details of the second intermediate image based on the target resolution corresponding to the task type to obtain the target image.
[0487] As an implementation of the present application, the conversion portion 204 may further include:
[0488] A second review unit is configured to perform a security review on the final text to determine whether the final text meets the preset text security specification conditions;
[0489] a second conversion unit configured to convert the final text into the target image using the image generation model corresponding to the content category if the final text meets the text security specification condition;
[0490] The third conversion unit is configured to delete the sensitive words in the final text if there are sensitive words that do not meet the text security specification conditions in the final text, and use the image generation model corresponding to the content category to convert the final text with the sensitive words deleted into the target image.
[0491] As an implementation of the present application, the image size includes the image length and the image width, the target size includes the target length and the target width, and the editing subunit may further include:
[0492] a first cropping subunit configured to, when an image length of the first intermediate image is less than the target length or an image width of the first intermediate image is less than the target width, crop the first intermediate image to obtain an image to be edge-expanded, wherein the image length of the image to be edge-expanded is less than or equal to the target length, and the image width of the edge-expanded image is less than or equal to the target width;
[0493] an image expansion subunit configured to expand the edge of the image to be expanded using blank pixels to obtain an expanded image, wherein the image length of the expanded image is equal to the target length, and the image width of the expanded image is equal to the target width;
[0494] The redrawing subunit is configured to redraw the expanded area of the expanded image to obtain the second intermediate image, wherein the expanded area is the area filled with the blank pixels.
[0495] As an implementation of the present application, the redraw subunit may also be configured as follows:
[0496] Obtaining description information of the first intermediate image;
[0497] performing masking processing on the expanded image using a grayscale mask image;
[0498] adding noise to the expanded area according to the pixel value of each pixel point in the expanded area after the masking process;
[0499] performing noise reduction processing on the expanded area after the mask processing based on the description information of the first intermediate image to obtain a third intermediate image;
[0500] A transition region between the expanded region and the initial region in the third intermediate image is smoothed to obtain the second intermediate image, wherein the initial region is a region outside the transition region in the expanded image.
[0501] As an implementation of the present application, the above-mentioned repair subunit may further include:
[0502] an acquiring subunit, configured to acquire description information of the second intermediate image;
[0503] a noise adding subunit, configured to add noise to a main area of the second intermediate image;
[0504] The denoising subunit is configured to perform denoising on a main region of the second intermediate image to which noise is added based on the description information of the second intermediate image to obtain the target image.
[0505] As an implementation of the present application, the image generating device 200 may further include:
[0506] an interpolation part configured to expand the second intermediate image along the length direction and the width direction respectively;
[0507] The deblurring part is configured to perform image enhancement on the second intermediate image after image expansion through a deblurring algorithm.
[0508] As an implementation of the present application, the image generating device 200 may further include:
[0509] a third acquiring unit configured to acquire the noise image and acquire a first prompt word from the final text; the first prompt word is used to determine the redrawing content;
[0510] a sixth determining unit, configured to extract global features of the noise image through a pre-built image control network, and determine a region to be processed of the noise image, wherein the region to be processed includes a redrawing region and a non-redrawing region, the redrawing region refers to an expanded region of the noise image and / or an internal edited region of the noise image, and the non-redrawing region refers to at least a portion of the noise image that does not need to be redrawn;
[0511] The redrawing unit is configured to redraw the area to be processed according to the first prompt word and the global feature to generate a first intermediate image;
[0512] Correspondingly, the redrawing unit is configured to extract image information of the noise image and redraw the first intermediate image according to the image information to generate a second intermediate image, wherein the image information includes edge information and / or segmentation information.
[0513] As an implementation of the present application, the third acquisition unit is configured to acquire a noise image and extract a descriptive text of the picture in the noise image; use the descriptive text and / or input text as the first prompt word, wherein the input text refers to the new content to be redrawn on the noise image; or, acquire prompt text and generate the noise image based on the prompt text using a text-based image method; use one or more texts among the prompt text, the input text and blank text as the first prompt word.
[0514] As an implementation method of the present application, the restoration subunit is configured to use a pre-set super-resolution model to perform super-resolution processing on the second intermediate image to obtain an image to be restored; perform semantic segmentation on the image to be restored to determine the main area of the image to be restored to be restored in detail; add noise to the main area of the image to be restored to obtain an image to be processed; and perform noise reduction processing on the main area of the image to be processed to which noise is added based on the description information of the image to be restored to obtain a target image after detail restoration.
[0515] As an implementation method of the present application, the repair subunit is configured to perform semantic segmentation on the image to be repaired, identify at least one target object in the image to be repaired; determine the target object containing the largest number of pixels among the at least one target object as the main object; and determine the area where the main object is located as the main area to be repaired in detail in the image to be repaired.
[0516] As an implementation of the present application, the image generating device 200 may further include:
[0517] a third review unit configured to review the final text according to a pre-established sample list including drawing texts; the sample list at least includes: drawing texts subject to conditional use and not permitted to be used, some of the drawing texts being provided with labels;
[0518] a sixth determining unit, configured to use any prompt word in the final text as a current prompt word;
[0519] An adding unit configured to add a label of the drawing text to the current prompt word when the current prompt word matches the drawing text with a preset label in the sample list;
[0520] a deleting unit configured to delete the current prompt word if the current prompt word matches a drawing text that is not allowed to be used in the sample list;
[0521] The replacement unit is configured to replace the current prompt word when the current prompt word matches the drawing text in the sample list that is subject to conditional use.
[0522] The image generation device provided in the embodiment of the present invention can implement each step in the above method embodiment, and will not be described again here to avoid repetition.
[0523] The present application also provides an image generation system, comprising:
[0524] The server is configured to obtain an image generation instruction from a user, where the image generation instruction is used to instruct generation of a target image;
[0525] The application side is configured to identify the image generation instruction, determine the content category of the target image based on the identification result, and obtain content enrichment information corresponding to the content category;
[0526] The application end is further configured to use the content enrichment information to perform text expansion on the prompt word corresponding to the image generation instruction to obtain a final text;
[0527] The application end is further configured to convert the final text into a target image corresponding to a target parameter according to a task type corresponding to the image generation instruction, where the target parameter refers to a target size and / or target resolution corresponding to the task type.
[0528] In some embodiments, the server is further configured to:
[0529] Receive a user's voice command and send the voice command to the application end;
[0530] Receiving a final text converted by the application end from the voice command;
[0531] displaying the final text on a display screen of the vehicle;
[0532] Sending the final text to the application end, and receiving a target image generated by the application end based on the final text;
[0533] The target image is displayed on a display screen of the vehicle.
[0534] In the image generation system, the server is located on the vehicle, while the application can be located in the cloud or on the vehicle. The two are connected via an intermediary layer. The application can deploy algorithms related to large models, while the server supports the interface display and interaction of the vehicle's in-cabin screen.
[0535] Specifically, the server can include a voice program and a drawing program. The voice program receives the user's voice command, passes it through the middle layer, and sends the voice command to the ASR algorithm of the application side for processing to obtain prompt words, and then sends the processed prompt words to the application side for display on the screen.
[0536] After that, the voice program sends the final text to the LLM algorithm on the application side. The LLM algorithm processes the final text, converts the prompt words into the final text, and sends the final text to the server side.
[0537] The voice program on the server side forwards the final text to the drawing program, calling up the drawing program. After obtaining the final text, the drawing program calls the text-based image-related algorithm on the application side to draw and obtain the target image. After content review, the target image is returned to the drawing program for display.
[0538] The image generation system provided by the embodiment of the present invention can implement each step in the above method embodiment, and will not be described again here to avoid repetition.
[0539] FIG12 shows a schematic diagram of the hardware structure of the image generating device provided in an embodiment of the present application.
[0540] The image generating device may include a processor 301 and a memory 302 storing computer program instructions.
[0541] Specifically, the processor 301 may include a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.
[0542] The memory 302 may include a large capacity memory for data or instructions. By way of example and not limitation, the memory 302 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory 302 may include removable or non-removable (or fixed) media. Where appropriate, the memory 302 may be inside or outside the integrated gateway disaster recovery device. In a specific embodiment, the memory 302 is a non-volatile solid-state memory.
[0543] The memory may include read-only memory (ROM), random access memory (RAM), magnetic disk storage media devices, optical storage media devices, flash memory devices, electrical, optical or other physical / tangible memory storage devices. Thus, generally, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to an aspect of the present application.
[0544] The processor 301 reads and executes computer program instructions stored in the memory 302 to implement any one of the image generation methods in the above embodiments.
[0545] In one example, the image generation device may further include a communication interface 303 and a bus 310. As shown in FIG12, the processor 301, the memory 302, and the communication interface 303 are connected via the bus 310 and communicate with each other.
[0546] The communication interface 303 is mainly used to implement communication between various modules, devices, units and / or equipment in the embodiments of the present application.
[0547] Bus 310 includes hardware, software or both, and the parts of image generation device are coupled to each other.For example, and not limitation, bus can include accelerated graphics port (AGP) or other graphics bus, enhanced industry standard architecture (EISA) bus, front side bus (FSB), hypertransport (HT) interconnection, industry standard architecture (ISA) bus, infinite bandwidth interconnection, low pin count (LPC) bus, memory bus, micro channel architecture (MCA) bus, peripheral component interconnection (PCI) bus, PCI-Express (PCI-X) bus, serial advanced technology attachment (SATA) bus, video electronics standard association local (VLB) bus or other suitable bus or two or more of these combinations.In suitable cases, bus 310 can include one or more buses.Although the present application embodiment describes and shows specific bus, the application considers any suitable bus or interconnection.
[0548] The image generation device may be based on the above embodiments, thereby realizing the combination of the above image generation method and apparatus.
[0549] In addition, in combination with the image generation method in the above embodiment, the embodiment of the present application may provide a computer storage medium for implementation. The computer storage medium stores computer program instructions; when the computer program instructions are executed by the processor, any one of the image generation methods in the above embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, they will not be described here. Among them, the above-mentioned computer-readable storage medium may include a non-transitory computer-readable storage medium, such as a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc., which is not limited here.
[0550] In addition, an embodiment of the present application also provides a vehicle, including computer program instructions, which, when executed by a processor, can implement the steps and corresponding contents of the aforementioned method embodiment.
[0551] It should be understood that the present application is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, a detailed description of known methods is omitted here. In the above embodiments, several specific steps are described and illustrated as examples. However, the method process of the present application is not limited to the specific steps described and illustrated. Those skilled in the art can make various changes, modifications, and additions, or change the order of the steps after understanding the spirit of the present application.
[0552] The functional blocks shown in the above block diagram can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of the present application are programs or code segments that are used to perform the required tasks. Programs or code segments can be stored in machine-readable media, or transmitted on a transmission medium or a communication link by a data signal carried in a carrier wave. "Machine-readable media" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROMs, flash memories, erasable ROMs (EROMs), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (RF) links, etc. The code segments can be downloaded via computer networks such as the Internet, intranets, etc.
[0553] It should also be noted that the exemplary embodiments mentioned in this application describe some methods or systems based on a series of steps or devices. However, this application is not limited to the order of the above steps. In other words, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.
[0554] Aspects of the present application are described above with reference to the flowcharts and / or block diagrams of the methods, devices and vehicles according to the embodiments of the present application. It should be understood that each box in the flowchart and / or block diagram and the combination of the boxes in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine so that these instructions executed by the processor of the computer or other programmable data processing device enable the implementation of the functions / actions specified in one or more boxes in the flowchart and / or block diagram. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor or a field programmable logic circuit. It is also understood that each box in the block diagram and / or flowchart and the combination of the boxes in the block diagram and / or flowchart can also be implemented by dedicated hardware that performs the specified function or action, or can be implemented by a combination of dedicated hardware and computer instructions.
[0555] The above is only a specific implementation method of the present application. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, modules and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. It should be understood that the scope of protection of the present application is not limited to this. Any technician familiar with this technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in this application, and these modifications or replacements should be included in the scope of protection of this application. Industrial Applicability
[0556] The embodiments of the present application provide an image generation method, apparatus, system, device, medium, and vehicle. The image generation method includes: obtaining a user's image generation instruction, the image generation instruction being used to instruct the generation of a target image; identifying the image generation instruction, determining the content category of the target image based on the identification result, and obtaining content-enrichment information corresponding to the content category; using the content-enrichment information to perform text expansion on the prompt word corresponding to the image generation instruction to obtain a final text; and using an image generation model corresponding to the content category to convert the final text into a target image. The implementation scheme using the above method can more clearly express the user's complete requirements for image content through the expanded final text, and more accurately understand the user's complete requirements for image content through the image generation model corresponding to the content category, thereby obtaining a target image that better meets the user's requirements and improves the quality of the target image generated in the vehicle.
Claims
1. A method for generating an image, the method comprising: Acquire an image generation instruction from a user, where the image generation instruction is used to instruct generation of a target image; Identify the image generation instruction, determine the content category of the target image based on the identification result, and obtain content enrichment information corresponding to the content category; Using the content enrichment information, text expansion is performed on the prompt word corresponding to the image generation instruction to obtain a final text; The final text is converted into a target image using an image generation model corresponding to the content category.
2. The image generation method according to claim 1, wherein: The step of identifying the image generation instruction and determining the content category of the target image based on the identification result includes: Identify the image generation instruction and obtain text content corresponding to the image generation instruction; Perform semantic understanding on the text content to obtain the user's intention; and extract prompt words for generating a target image from the user's intention; The content category of the target image is determined based on the feature information in the prompt word.
3. The image generation method according to claim 1, wherein: The step of identifying the image generation instruction and determining the content category of the target image based on the identification result includes: Recognize the image generation instruction, and determine a content phrase in the prompt word corresponding to the image generation instruction based on the recognition result, wherein the content phrase is used to indicate the generation of content elements in the target image; In the case where there are multiple content phrases in the prompt word, determining the content weights of the content elements corresponding to the content phrases to obtain multiple content weights; The content element with the highest content weight among the multiple content weights is determined as the main content of the target image, and the category of the main content is determined as the content category of the target image.
4. The image generation method according to claim 1, wherein: The step of identifying the image generation instruction and determining the content category of the target image based on the identification result includes: Recognize the image generation instruction, and determine the prompt word corresponding to the image generation instruction based on the recognition result; Conducting a security review on the prompt word to determine whether the prompt word meets the preset text security specification conditions; In the case that the prompt word meets the text safety specification condition, the content category of the target image is determined based on the feature information of the prompt word.
5. The image generation method according to any one of claims 1 to 4, wherein: The content category includes a subject category and a style category, and the obtaining of content enrichment information corresponding to the content category includes: Acquire a basic image generation model corresponding to the subject category of the target image and a specific object generation model corresponding to the style category of the target image; Searching a pre-set subject mapping table for subject keywords that have a mapping relationship with the basic image generation model, and determining the subject keywords as content enrichment information corresponding to the subject category; A style keyword having a mapping relationship with the specific object generation model is searched in a preset style mapping table, and the style keyword is determined as content enrichment information corresponding to the style category.
6. The image generation method according to any one of claims 1 to 4, wherein: The obtaining of the content enrichment information corresponding to the content category includes: determining at least one candidate art style type that is different from the style category and matches the subject category; The vocabulary representing the candidate artistic style type is determined as content-enriched information matching the style category.
7. The image generation method according to any one of claims 1 to 6, wherein: Before determining the content category of the target image based on the recognition result, the method further includes: When it is detected that there is a special word in the prompt word corresponding to the image generation instruction, searching for a common word that has a corresponding relationship with the special word; The common words are used to replace the corresponding special words in the prompt words.
8. The image generation method according to any one of claims 1 to 7, wherein: The converting the final text into a target image by using the image generation model corresponding to the content category includes: Conducting a security review on the final text to determine whether the final text meets the preset text security specification conditions; If the final text meets the text security specification condition, converting the final text into the target image using the image generation model corresponding to the content category; If the final text contains sensitive words that do not meet the text security specification conditions, the sensitive words in the final text are deleted, and the final text with the sensitive words deleted is converted into the target image using the image generation model corresponding to the content category.
9. The image generation method according to any one of claims 1 to 7, wherein: The converting the final text into a target image by using the image generation model corresponding to the content category includes: Acquire a basic image generation model and a specific object generation model corresponding to the content category; Mounting the specific object generation model onto the basic image generation model to obtain a fused image generation model, wherein the image generation model includes the basic image generation model, the specific object generation model and the fused image generation model; According to the task type corresponding to the image generation instruction, the final text is converted into a target image corresponding to a target parameter using the fused image generation model, where the target parameter refers to a target size and / or a target resolution corresponding to the task type.
10. The image generation method according to claim 9, wherein: The step of converting the final text into a target image corresponding to a target parameter by using the fused image generation model includes: Using the final text to guide the fused image generation model to convert the randomly generated noise image into a first intermediate image; Performing image editing on the first intermediate image to obtain a second intermediate image, wherein an image size of the second intermediate image matches a target size corresponding to the task type; Based on the target resolution corresponding to the task type, image details of the second intermediate image are restored to obtain the target image.
11. The image generation method according to claim 10, wherein: The image size includes an image length and an image width, the target size includes a target length and a target width, and the image editing the first intermediate image to obtain a second intermediate image includes: When the image length of the first intermediate image is less than the target length, or the image width of the first intermediate image is less than the target width, the first intermediate image is cropped to obtain an image to be edge expanded, the image length of the image to be edge expanded is less than or equal to the target length, and the image width of the edge expanded image is less than or equal to the target width; Performing edge expansion on the image to be expanded by using blank pixels to obtain an expanded image, wherein the image length of the expanded image is equal to the target length, and the image width of the expanded image is equal to the target width; The expanded area of the expanded image is redrawn to obtain the second intermediate image, wherein the expanded area is an area filled with the blank pixels.
12. The image generation method according to claim 11, wherein: The step of redrawing the expanded area of the expanded image to obtain the second intermediate image includes: Obtaining description information of the first intermediate image; Using a grayscale mask image to perform mask processing on the expanded image; adding noise to the expanded area according to the pixel value of each pixel point in the expanded area after mask processing; Performing noise reduction processing on the expanded area after mask processing based on the description information of the first intermediate image to obtain a third intermediate image; A transition area between the extended area and the initial area in the third intermediate image is smoothed to obtain the second intermediate image, wherein the initial area is an area outside the transition area in the extended image.
13. The image generation method according to any one of claims 10 to 12, wherein: The repairing of the image details of the second intermediate image to obtain the target image includes: Obtaining description information of the second intermediate image; adding noise to a subject area of the second intermediate image; Based on the description information of the second intermediate image, noise reduction processing is performed on the main area with noise added in the second intermediate image to obtain the target image.
14. The image generation method according to any one of claims 10 to 13, wherein: Before repairing the image details of the second intermediate image to obtain the target image, the method further includes: Using interpolation technology, the second intermediate image is expanded along the length direction and the width direction respectively; The second intermediate image after image expansion is enhanced by a deblurring algorithm.
15. The image generation method according to any one of claims 10 to 14, wherein: The step of using the final text to guide the fusion image generation model to convert the randomly generated noise image into a first intermediate image includes: Acquire the noise image, and acquire a first prompt word from the final text; the first prompt word is used to determine the redrawing content; Extracting global features of the noise image through a pre-constructed image control network, and determining a region to be processed of the noise image, wherein the region to be processed includes a redrawing region and a non-redrawing region, the redrawing region refers to an expanded region of the noise image and / or an internal editing region of the noise image, and the non-redrawing region refers to at least a portion of the region in the noise image that does not need to be redrawn; Redrawing the area to be processed according to the first prompt word and the global feature to generate a first intermediate image; Accordingly, the editing the first intermediate image to obtain the second intermediate image includes: Image information of the noise image is extracted, and the first intermediate image is redrawn according to the image information to generate a second intermediate image, wherein the image information includes edge information and / or segmentation information.
16. The image generation method according to claim 15, characterized in that: The step of acquiring the noise image and the first prompt word comprises: Acquire a noise image, and extract a description text of a picture in the noise image; use the description text and / or input text as the first prompt word, wherein the input text refers to the new content to be redrawn on the noise image; or, Acquire prompt text, and generate the noise image based on the prompt text by adopting a text-generated image method; and use one or more texts among the prompt text, the input text and blank text as the first prompt word.
17. The image generation method according to any one of claims 10 to 16, wherein: The repairing of the image details of the second intermediate image to obtain the target image includes: Performing super-resolution processing on the second intermediate image using a preset super-resolution model to obtain an image to be restored; Performing semantic segmentation on the image to be repaired to determine a main area of the image to be repaired where details are to be repaired; Adding noise to the main area of the image to be repaired to obtain an image to be processed; According to the description information of the image to be repaired, a noise reduction process is performed on the main area of the image to be processed where noise is added, so as to obtain a target image after detail repair.
18. The image generation method according to claim 17, characterized in that: The performing semantic segmentation on the image to be repaired to determine a main area of the image to be repaired where details are to be repaired includes: Performing semantic segmentation on the image to be restored, and identifying at least one target object in the image to be restored; Determine the target object containing the largest number of pixels in the at least one target object as the main object; The area where the main object is located is determined as the main area to be repaired in detail in the image to be repaired.
19. The image generation method according to any one of claims 1 to 18, wherein: After the prompt word corresponding to the image generation instruction is expanded using the content enrichment information to obtain a final text, the method further includes: The final text is reviewed according to a pre-established sample list including drawing texts; the sample list at least includes: drawing texts that are subject to conditional use and those that are not allowed to be used, and some of the drawing texts are provided with labels; Using any prompt word in the final text as the current prompt word; When the current prompt word matches the drawing text with a preset label in the sample list, adding the label of the drawing text to the current prompt word; When the current prompt word matches the drawing text that is not allowed to be used in the sample list, deleting the current prompt word; When the current prompt word matches the drawing text in the sample list that is used subject to conditional restrictions, the current prompt word is replaced.
20. An image optimization method, wherein: The method comprises: Acquire a low-resolution image, wherein the low-resolution image is converted from a prompt word input by a user; Performing super-resolution processing on the low-resolution image using a preset super-resolution model to obtain an image to be repaired; Perform detail restoration on the main area of the image to be restored to obtain a target image.
21. A method for generating an image, wherein: The method comprises: Acquire an image to be processed and a first prompt word, wherein the first prompt word is used to determine the redrawing content; Extracting global features of the image to be processed through a pre-constructed image control network, and determining a region to be processed of the image to be processed, wherein the region to be processed includes a redrawing region and a non-redrawing region, the redrawing region refers to an expanded region of the image to be processed and / or an internal editing region of the image to be processed, and the non-redrawing region refers to at least a portion of the region of the image to be processed that does not need to be redrawn; Redrawing the area to be processed according to the first prompt word and the global feature to generate a first image; Image information of the image to be processed is extracted, and the first image is redrawn according to the image information to generate a second image, wherein the image information includes edge information and / or segmentation information.
22. A prompt word optimization method for generating an image from text, wherein: The method comprises: Get the initial prompt word entered by the user; Extracting a target subject and a target style type based on the initial prompt word; Determine a rich vocabulary that matches the target subject and / or the target style type; wherein the rich vocabulary includes: vocabulary used to supplement the content of the initial prompt word; The enriched vocabulary is added to the initial prompt word to obtain an optimized prompt word.
23. An image generating device, the device comprising: An acquisition part, configured to acquire an image generation instruction of a user, wherein the image generation instruction is used to instruct generation of a target image; an identification part configured to identify the image generation instruction, determine the content category of the target image based on the identification result, and obtain content enrichment information corresponding to the content category; An expansion part is configured to perform text expansion on the prompt word corresponding to the image generation instruction using the content enrichment information to obtain a final text; The conversion part is configured to convert the final text into a target image using an image generation model corresponding to the content category.
24. An image generation system, the system comprising: The server is configured to obtain an image generation instruction from a user, where the image generation instruction is used to instruct generation of a target image; The application end is configured to identify the image generation instruction, determine the content category of the target image based on the identification result, and obtain content enrichment information corresponding to the content category; The application end is further configured to perform text expansion on the prompt word corresponding to the image generation instruction using the content enrichment information to obtain a final text; The application end is further configured to convert the final text into a target image using an image generation model corresponding to the content category.
25. An image generating device, the device comprising: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, the image generating method according to any one of claims 1 to 22 is implemented.
26. A computer storage medium having computer program instructions stored thereon, wherein the computer program instructions, when executed by a processor, implement the image generation method according to any one of claims 1 to 22.
27. A vehicle, the vehicle comprising at least one of the following: The image generating device as claimed in claim 23; The image generation system as claimed in claim 24; The image generating device as claimed in claim 25; The computer storage medium of claim 26.
Citation Information
Patent Citations
Video conference-oriented text region repairing system and method
CN114240791A
Image generation method and device, electronic equipment and computer readable storage medium
CN116580127A
Processing system and method for generating image based on user text cue word
CN116680425A
Image generation method and device, electronic equipment and storage medium
CN116797684A
Rapid image generation method and device, medium and equipment
CN117094881A
Cited By
AIGC-based photographed image generation method and system, and photographing device
CN120897043A
Training method of generative model for generating visual IP picture material and visual IP picture generation method
CN120931765A
Video content forgetting method and system based on double-layer optimization, and medium
CN121174016A
Scene editing method and system based on anchor point diagram structure and three-dimensional Gaussian sputtering
CN121414963A
Scene editing method and system based on anchor point graph structure and three-dimensional gaussian splatting
CN121414963B