Image generation method, device, system and equipment, medium and vehicle
By identifying the user's image generation instructions, determining the content category and text expansion, and converting images using the corresponding image generation model, the problem of low image quality in the vehicle is solved and higher quality image generation is achieved.
Patent Information
- Application Number
- CN202311756165.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-19
- Publication Date
- 2025-06-20
AI Technical Summary
In the prior art, the image quality generated in the vehicle is low and cannot effectively match the user's image generation instructions, resulting in the generated target image content that does not meet the user's needs.
By obtaining user's image generation instructions, identifying the instruction content, determining the content category of the target image, and using the content enrichment information of the category to expand the prompt words, and finally converting the expanded text into the target image using the corresponding image generation model.
Through the combination of text expansion and image generation model, users' complete needs for image content can be expressed more clearly, the quality of the generated target images can be improved, and the image can be more in line with users' expectations.
Smart Images

Figure CN120182972A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of artificial intelligence technology, and particularly relates to an image generation method, device, system, equipment, medium and vehicle. Background Art
[0002] Text-to-image is a computer generation task aimed at converting image generation instructions or natural language text into corresponding images. In this task, the image generation model in the computer needs to understand the user's image generation instructions and generate images that match the image generation instructions.
[0003] In the related art, algorithms related to the image generation model can be transplanted into the vehicle's controller. Thus, the user can instruct the controller to generate a target image through simple image generation instructions in the vehicle's cabin, and display the target image on the vehicle's display screen or application. However, in the related art, the prompt words obtained by converting the user's image generation instructions can be input into a pre-set image generation model, and the image generation model is used to convert the prompt words into images that match the prompt words. Since the image generation instructions proposed by the user are often relatively simple and cannot clearly express the user's complete requirements for the target image, it is easy to cause the content of the generated target image to not match the content requirements indicated by the user in the image generation instructions. Therefore, the content of the target image generated in the related art often does not meet the user's requirements, resulting in low-quality generated images. Summary of the Invention
[0004] Embodiments of this application provide an image generation method, device, system, equipment, medium and vehicle, which can solve the problem of low quality of images generated in vehicles.
[0005] In a first aspect, embodiments of this application provide an image generation method, the method including:
[0006] Obtain a user's image generation instruction, where the image generation instruction is used to indicate the generation of a target image;
[0007] Identify the image generation instruction, determine the content category of the target image based on the identification result, and obtain content enrichment information corresponding to the content category;
[0008] Use the content enrichment information to expand the text of the prompt words corresponding to the image generation instruction to obtain a final text;
[0009] Use the image generation model corresponding to the content category to convert the final text into a target image.
[0010] In some embodiments, the recognition of the image generation instruction and determining the content category of the target image based on the recognition result includes:
[0011] Recognize the image generation instruction to obtain the text content corresponding to the image generation instruction;
[0012] Perform semantic understanding on the text content to obtain the user intention; and extract the prompt words for generating the target image from the user intention;
[0013] Determine the content category of the target image based on the feature information in the prompt words.
[0014] In some embodiments, the recognition of the image generation instruction and determining the content category of the target image based on the recognition result includes:
[0015] Recognize the image generation instruction and determine the content phrase in the prompt word corresponding to the image generation instruction based on the recognition result, where the content phrase is used to indicate the content element to be generated in the target image;
[0016] In the case where there are multiple content phrases in the prompt word, query the content weights of the content elements corresponding to each content phrase in a preset content weight table to obtain multiple content weights;
[0017] Determine the content element with the highest content weight among the multiple content weights as the main content of the target image, and determine the category of the main content as the content category of the target image.
[0018] In some embodiments, the recognition of the image generation instruction and determining the content category of the target image based on the recognition result includes:
[0019] Recognize the image generation instruction and determine the prompt word corresponding to the image generation instruction based on the recognition result;
[0020] Perform a security review on the prompt word to determine whether the prompt word meets the preset text security specification conditions;
[0021] In the case where the prompt word meets the text security specification conditions, determine the content category of the target image based on the feature information of the prompt word.
[0022] In some embodiments, the image generation model includes a basic image generation model and a specific object generation model, the content category includes a subject category and a style category, and obtaining the content enrichment information corresponding to the content category includes:
[0023] Obtain the basic image generation model corresponding to the main category of the target image, and the specific object generation model corresponding to the style category of the target image;
[0024] Query in the pre-set main mapping table for the main keywords that have a mapping relationship with the basic image generation model, and determine the main keywords as the content-rich information corresponding to the main category;
[0025] Query in the pre-set style mapping table for the style keywords that have a mapping relationship with the specific object generation model, and determine the style keywords as the content-rich information corresponding to the style category.
[0026] In some embodiments, before determining the content category of the target image based on the recognition result, the method further includes:
[0027] When it is detected that there are special words in the prompt words corresponding to the image generation instruction, query for the general words that have a corresponding relationship with the special words;
[0028] Use the general words to replace the corresponding special words in the prompt words.
[0029] In some embodiments, the converting the final text into a target image by using the image generation model corresponding to the content category includes:
[0030] Conduct a security review on the final text to determine whether the final text meets the preset text security specification conditions;
[0031] If the final text meets the text security specification conditions, use the image generation model corresponding to the content category to convert the final text into the target image;
[0032] If there are sensitive words in the final text that do not meet the text security specification conditions, delete the sensitive words in the final text, and use the image generation model corresponding to the content category to convert the final text after deleting the sensitive words into the target image.
[0033] In some embodiments, the converting the final text into a target image by using the image generation model corresponding to the content category includes:
[0034] Obtain the basic image generation model and the specific object generation model corresponding to the content category;
[0035] Mount the specific object generation model onto the basic image generation model to obtain a fused image generation model, where the image generation model includes the basic image generation model, the specific object generation model, and the fused image generation model;
[0036] According to the task type corresponding to the image generation instruction, use the fusion image generation model to convert the final text into a target image corresponding to the target parameters, where the target parameters refer to the target size and / or target resolution corresponding to the task type.
[0037] In some embodiments, the converting the final text into a target image corresponding to the target parameters by using the fusion image generation model includes:
[0038] Use the final text to guide the fusion image generation model to convert a randomly generated noise image into a first intermediate image;
[0039] Perform image editing on the first intermediate image to obtain a second intermediate image, where the image size of the second intermediate image matches the target size corresponding to the task type;
[0040] Based on the target resolution corresponding to the task type, repair the image details of the second intermediate image to obtain the target image.
[0041] In some embodiments, the image size includes the image length and the image width, the target size includes the target length and the target width, and the performing image editing on the first intermediate image to obtain a second intermediate image includes:
[0042] In the case where the image length of the first intermediate image is less than the target length or the image width of the first intermediate image is less than the target width, crop the first intermediate image to obtain an image to be border-expanded, where the image length of the image to be border-expanded is less than or equal to the target length, and the image width of the border-expanded image is less than or equal to the target width;
[0043] Use blank pixels to perform edge expansion on the image to be border-expanded to obtain an expanded image, where the image length of the expanded image is equal to the target length, and the image width of the expanded image is equal to the target width;
[0044] Redraw the expanded area of the expanded image to obtain the second intermediate image, where the expanded area is the area filled with the blank pixels.
[0045] In some embodiments, the redrawing the expanded area of the expanded image to obtain the second intermediate image includes:
[0046] Obtain the description information of the first intermediate image;
[0047] Use a grayscale mask image to perform mask processing on the expanded image;
[0048] Add noise to the expanded area according to the pixel values of each pixel point in the expanded area after mask processing;
[0049] Perform noise reduction processing on the expanded area after mask processing based on the description information of the first intermediate image to obtain a third intermediate image;
[0050] Smooth the transition area between the expanded area and the initial area in the third intermediate image to obtain the second intermediate image, where the initial area is the area outside the transition area in the expanded image.
[0051] In some embodiments, the repairing the image details of the second intermediate image to obtain the target image includes:
[0052] Obtain the description information of the second intermediate image;
[0053] Add noise to the main area of the second intermediate image;
[0054] Perform noise reduction processing on the main area with added noise in the second intermediate image based on the description information of the second intermediate image to obtain the target image.
[0055] In some embodiments, before the repairing the image details of the second intermediate image to obtain the target image, the method further includes:
[0056] Adopt an interpolation technique to perform image expansion on the second intermediate image along the length direction and the width direction respectively;
[0057] Perform image enhancement on the second intermediate image after image expansion through a deblurring algorithm.
[0058] In a second aspect, an image generation device provided by an embodiment of the present application includes:
[0059] An acquisition module, configured to acquire an image generation instruction of a user, where the image generation instruction is used to indicate generating a target image;
[0060] An identification module, configured to identify the image generation instruction, determine the content category of the target image based on the identification result, and acquire content rich information corresponding to the content category;
[0061] An expansion module, configured to expand the prompt word corresponding to the image generation instruction by using the content rich information to obtain a final text;
[0062] A conversion module, configured to convert the final text into a target image by using an image generation model corresponding to the content category.
[0063] In a third aspect, an embodiment of the present application provides an image generation system, which includes:
[0064] A server for obtaining an image generation instruction of a user, where the image generation instruction is used to indicate the generation of a target image;
[0065] An application terminal for recognizing the image generation instruction, determining the content category of the target image based on the recognition result, and obtaining content enrichment information corresponding to the content category;
[0066] The application terminal is further configured to use the content enrichment information to expand the prompt words corresponding to the image generation instruction to obtain a final text;
[0067] The application terminal is further configured to use the image generation model corresponding to the content category to convert the final text into a target image.
[0068] In some embodiments, the server is further configured to:
[0069] Receive a voice instruction of the user and send the voice instruction to the application terminal;
[0070] Receive the final text converted from the voice instruction by the application terminal;
[0071] Display the final text on the display screen of the vehicle;
[0072] Send the final text to the application terminal and receive the target image generated by the application terminal based on the final text;
[0073] Display the target image on the display screen of the vehicle.
[0074] In a fourth aspect, an embodiment of the present application provides an image generation device, which includes: a processor and a memory storing computer program instructions;
[0075] When the processor executes the computer program instructions, the above-mentioned image generation method is implemented.
[0076] In a fifth aspect, an embodiment of the present application provides a computer storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the above-mentioned image generation method is implemented.
[0077] In a sixth aspect, an embodiment of the present application provides a vehicle, which includes the above-mentioned image generation device, image generation system, image generation device, and storage medium.
[0078] In this application, by recognizing the user's image generation instruction, then determining the content category of the target image indicated by the recognized result for the image generation instruction, and using the content enrichment information corresponding to the content category to expand the prompt word corresponding to the image generation instruction to obtain the final text, and then using the image generation model corresponding to the content category to convert the final text into the target image. In this way, the complete requirements of the user for the image content can be more clearly expressed through the expanded final text, and the complete requirements of the user for the image content can be more accurately understood through the image generation model corresponding to the content category, so as to obtain a target image that better meets the user's needs and improve the quality of the target image generated in the vehicle. BRIEF DESCRIPTION OF THE DRAWINGS
[0079] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required to be used in the embodiments of the present application will be briefly introduced below. Obviously, the following described drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0080] Figure 1 It is a schematic flowchart of an image generation method provided by an embodiment of the present application;
[0081] Figure 2 It is a schematic structural diagram of an image generation device provided by an embodiment of the present application;
[0082] Figure 3 It is a schematic hardware structure diagram of an image generation device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0083] The features and exemplary embodiments of various aspects of the present application will be described in detail below. In order to make the purpose, technical solutions and advantages of the present application more clear, the present application will be further described in detail below in combination with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described here are only intended to explain the present application, rather than to limit the present application. For those skilled in the art, the present application can be implemented without some of these specific details. The following description of the embodiments is only to provide a better understanding of the present application by showing examples of the present application.
[0084] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0085] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments may be combined with each other. The embodiments will be described in detail below with reference to the accompanying drawings.
[0086] Specifically, to solve the problems of the prior art, embodiments of the present application provide an image generation method, apparatus, system, device, medium and vehicle. First, the image generation method provided by the embodiments of the present application will be introduced below.
[0087] Figure 1 The flowchart of the image generation method provided by an embodiment of the present application is shown. This method can be applied to the in-vehicle computer of a vehicle or a cloud server communicatively connected to the vehicle. The method includes the following steps:
[0088] S110, obtain an image generation instruction of a user, where the image generation instruction is used to indicate generating a target image.
[0089] In this embodiment, the image generation method can be applied to a vehicle. The image generation instruction can be a voice instruction, a text instruction or an image instruction. The voice instruction can be a piece of voice spoken by the user, the text instruction can be a piece of text input by the user, and the image instruction can be an image input by the user.
[0090] Exemplarily, when the voice detection system in the vehicle is awakened, the voice detection system can be used to capture the voice signal emitted by the user inside or outside the vehicle, and then, using recognition technology, analyze the collected voice signal to detect whether the voice signal contains a voice instruction. If the voice signal contains a voice instruction, obtain the voice instruction.
[0091] S120, recognize the image generation instruction, determine the content category of the target image based on the recognition result, and obtain the content enrichment information corresponding to the content category;
[0092] In this embodiment, after receiving an image generation instruction, the image generation instruction can be recognized, and based on the recognition result, the image generation instruction can be converted into corresponding text content. Then, semantic understanding is performed on the text content, that is, the prompt words for the target image can be filtered out from the text content, and the content categories of each item in the target image can be determined based on the prompt words. Then, content enrichment information associated with each content category can be obtained, and the content enrichment information can enrich and improve the prompt words.
[0093] Specifically, when the image generation instruction is a voice instruction, the voice instruction can be converted into the text content corresponding to the image generation instruction. When the image generation instruction is a text instruction, the text instruction can also be determined as the text content corresponding to the image generation instruction. When the image generation instruction is an image instruction, the description information of the image instruction can be obtained by extracting the features of the image instruction, and the description information of the image instruction can be determined as the text content corresponding to the image generation instruction.
[0094] As an optional embodiment, the recognition of the image generation instruction and the determination of the content category of the target image based on the recognition result include:
[0095] Recognize the image generation instruction to obtain the text content corresponding to the image generation instruction;
[0096] Perform semantic understanding on the text content to obtain the user intention; and extract the prompt words for generating the target image from the user intention;
[0097] Determine the content category of the target image based on the feature information in the prompt words.
[0098] In this embodiment, the content of the target image can include the objects in the target image, the background in the target image, the image style of the target image, the main color of the target image, and the lighting effect of the target image, etc.
[0099] Contents with similar feature information can be determined as the same content category. Feature information refers to the attribute features of the content generated as indicated by the prompt. Such as object attributes, style attributes, color attributes, etc. Specifically, when the content of the target image is an object in the target image, the content category of the target image is the object category, and the object category of the main object in the target image is the main category. For example, the object category can be people, scenery, animals, food, transportation vehicles, etc. When the main object in the prompt is a cat or a dog, the main category of the prompt is animals. When the main object in the prompt is coffee or bread, the main category of the prompt is food. When the content of the target image is the image style in the target image, the content category of the target image is the style category. Common style types can include two major categories: realistic styles and artistic styles. Among them, realistic styles can further include realistic people, scenery, animals, buildings, etc., and artistic styles can include styles such as comics, oil paintings, ink paintings, line drawings, etc.
[0100] After determining the content category, the corresponding descriptor of the content category can be directly obtained as the content enrichment information, or the image generation model corresponding to the content category can be determined first, and then the keywords corresponding to the image generation model can be determined as the content enrichment information.
[0101] For example, when the content category is the main category, the corresponding descriptor of the main category can be directly obtained as the content enrichment information, or the basic image generation model corresponding to the main category can be determined first, and then the main keywords that have a mapping relationship with the basic image generation model can be queried as the content enrichment information. When the content category is the style category, the corresponding descriptor of the style category can be directly obtained as the content enrichment information, or the specific object generation model corresponding to the style category can be determined first, and then the style keywords that have a mapping relationship with the specific object generation model can be queried as the content enrichment information.
[0102] In this embodiment, the prompt is used to describe the content of the target image that the user wants to generate. For example, the prompt can be "generate a little dog" or "generate a big tree". And the prompt can include a positive prompt and a negative prompt. The meaning of the positive prompt is the content that is expected to be included in the target image, while the negative prompt indicates the content that is not expected to be included in the target image. For example, the positive prompt can be "a Chinese girl in anime style running on the beach, with flying seagulls and a gorgeous rainbow behind, and the overall picture is poetic", and the negative prompt can be "pornography, nudity, ugliness, deformity". Similarly, the content enrichment information can also include positive enrichment information and negative enrichment information. Among them, the positive enrichment information is used to enrich the content of the positive prompt, and the negative enrichment information is used to enrich the content of the negative prompt.
[0103] Exemplarily, the image generation instruction can be converted into the corresponding text form, i.e., the text content, by using the ASR (Automatic Speech Recognition) algorithm. The LLM (Large Language Model) can also be used to extract the semantics of the text content in text form, that is, to analyze the text content by using this large language model to extract the semantic information of the text, and then determine the content category of the target image based on the semantic information, and obtain the content-rich information corresponding to the content category.
[0104] S130. Use the content-rich information to expand the prompt word corresponding to the image generation instruction to obtain the final text.
[0105] In this embodiment, the prompt word can be a relatively simple sentence refined by the user based on the image generation instruction. In order to enrich the image content of the target image converted from the prompt word and improve the image quality, the obtained content-rich information can be used to expand the prompt word to obtain the final text with richer content.
[0106] Exemplarily, the content-rich information can be a descriptive word related to the main category of the target image. The descriptive word in the pre-set descriptive word library can be added to the prompt word to describe and extend the prompt word in detail; for example, when the positive prompt word is "a chubby and cute tabby cat with a round face in anime style is sunning itself on the lawn full of flowers", it can be expanded to obtain the final text "a chubby and cute tabby cat with a round face in anime style is sunning itself on the lawn full of flowers, simple and lovely, with big eyes, charming scenery, and professional photography techniques".
[0107] S140. Use the image generation model corresponding to the content category to convert the final text into a target image.
[0108] In this embodiment, after determining the content category, the image generation model corresponding to the content category can be queried based on the corresponding relationship between the pre-set image generation model and the content category, and then the image generation model converts the final text into a target image.
[0109] As an optional embodiment, the step of using the image generation model corresponding to the content category to convert the final text into a target image includes:
[0110] Obtain the basic image generation model and the specific object generation model corresponding to the content category;
[0111] Mount the specific object generation model onto the basic image generation model to obtain a fused image generation model, where the image generation model includes the basic image generation model, the specific object generation model, and the fused image generation model;
[0112] According to the task type corresponding to the image generation instruction, use the fused image generation model to convert the final text into a target image corresponding to the target parameters, where the target parameters refer to the target size and / or target resolution corresponding to the task type.
[0113] In this embodiment, the image generation model may include a basic image generation model and a specific object generation model. Among them, the basic image generation model may be the base model in the latent diffusion image generation model (Stable Diffusion model), which can convert text content into a basic image. The specific object generation model may be the Low-Rank Adaptation of Large Language Models (LoRA image generation model) in the latent diffusion image generation model, which is used to generate images of certain specific styles or specific subjects. Among them, the specific object generation model cannot be used alone and needs to be mounted on the basic image generation model to assist the basic image generation model in realizing image generation.
[0114] After mounting the specific object generation model onto the basic image generation model to obtain a fused image generation model, a random noise image can be obtained, and the final text can be converted into a text embedding vector. The text embedding vector is embedded into the fused image generation model loaded with a sampler to guide the image generation model to generate a first intermediate image.
[0115] In the image generation instruction, the user can not only put forward expectations for the image content in the target image, but also put forward expectations for the target parameters of the target image. Specifically, the target parameters of the target image can be specified by setting the task type of the target generation instruction. That is to say, the corresponding relationship between the task type and the image parameters can be set in advance. After determining the task type corresponding to the image generation instruction, the image parameters corresponding to this task type can be determined as the target parameters of the target image indicated by this image generation instruction, where different task types can be distinguished based on the application scenario of the target image.
[0116] For example, the task type can be the center control wallpaper of the vehicle center control screen, the head-up wallpaper of the vehicle head-up display, or the background image of an application in the vehicle. The target size corresponding to the center control wallpaper needs to match the screen size of the center control screen, and the target resolution corresponding to the center control wallpaper can be 800x480; the target size corresponding to the head-up wallpaper needs to match the size of the head-up display, and the target resolution corresponding to the head-up wallpaper can be 1024x600. The target size corresponding to the background image of the application needs to match the size of the display frame of the application, and the target resolution corresponding to the background image can be 1280x720.
[0117] After determining the final text and the target parameters, the final text can be input into a pre-selected image generation model. The image generation model is used to convert the final text into a first intermediate image, and then the size and resolution of the first intermediate image are adjusted according to the target parameters to obtain the final target image.
[0118] In this application, by identifying the user's image generation instruction, then determining the content category of the target image indicated by the image generation instruction based on the recognition result, and using the content enrichment information corresponding to the content category to expand the prompt word corresponding to the image generation instruction to obtain the final text, and then using the image generation model corresponding to the content category to convert the final text into the target image. In this way, the complete requirements of the user for the image content can be more clearly expressed through the expanded final text, and the complete requirements of the user for the image content can be more accurately understood through the image generation model corresponding to the content category, so as to obtain a target image that better meets the user's needs and improve the quality of the target image generated in the vehicle.
[0119] As an optional embodiment, the recognition of the image generation instruction and the determination of the content category of the target image based on the recognition result include:
[0120] Recognize the image generation instruction, and determine the content phrase in the prompt word corresponding to the image generation instruction based on the recognition result. The content phrase is used to indicate the content element to be generated in the target image;
[0121] In the case where there are multiple content phrases in the prompt word, query the content weights of the content elements corresponding to each content phrase in a pre-set content weight table to obtain multiple content weights;
[0122] Determine the content element with the highest content weight among the multiple content weights as the main content of the target image, and determine the category of the main content as the content category of the target image.
[0123] In this embodiment, the pre-set content weight table includes the content weights of each content element. The greater the content weight, the greater the influence of the content element corresponding to the content weight on the target image, that is, the more important the content element is in the target image. Therefore, the content element with the highest content weight can be determined as the main content of the target image.
[0124] In addition, since the content in the target image includes multiple dimensions, such as objects, backgrounds, image styles, and colors, the content categories accordingly also include multiple dimensions, such as subject categories, background categories, style categories, color categories, etc. The content element with the highest content weight in each dimension can be determined as the main content of the target image, and then the category of the main content in each dimension can be determined as the content category of the target image.
[0125] Specifically, the prompt can be segmented to obtain multiple segmented phrases, and then the semantics of each segmented phrase can be recognized. Based on the semantic recognition results of each segmented phrase, it can be determined whether each segmented phrase indicates the generation of a corresponding content element in the target image. If the segmented phrase indicates the generation of a corresponding content element in the target image, the segmented phrase can be determined as the content phrase in the prompt, and further, based on the semantic recognition results, the dimension to which the content element belongs can be determined. For example, if there is a segmented phrase "kitten" in the prompt, the semantics of "kitten" can be recognized to determine that "kitten" is a content phrase indicating the generation of a content element in the target image, and further determine that the dimension to which the kitten belongs is an object; if there is a segmented phrase "animation" in the prompt, the semantics of "animation" can be recognized to determine that animation is a content phrase indicating the generation of a content element in the target image, and further determine that the dimension to which animation belongs is the image style.
[0126] Exemplarily, when the content category is the subject category, the content weight table is the object weight table, and the object weight is used to represent the influence degree of the corresponding main object on the target image. Then, in the case where there are multiple objects to be generated in the prompt, the object weights of each object to be generated can be queried in the pre-set object weight table to obtain multiple object weights; the object to be generated with the highest object weight among the multiple object weights can be determined as the main object of the target image, and the category of the main object can be determined as the subject category of the target image.
[0127] Therefore, if the prompt includes multiple objects to be generated, the object weights of each object to be generated can be queried in the pre-defined object weight table, and then the object to be generated with the highest object weight can be determined as the main object of the target image, and the category of the main object is the subject category of the target image.
[0128] Among them, the object weights of each object to be generated in the object weight table can be set in advance by the user's preference for various types of objects. For example, the main object categories can be people, landscapes, animals, food, transportation vehicles, etc. When the object in the prompt is a cat or a dog, the object category of the prompt is an animal. When the object in the prompt is coffee or bread, the object category of the prompt is food. If the user's preference for animals is greater than that for food, the object weight of animals can be set to be greater than the object weight of food.
[0129] In addition, the object weights of each object to be generated in the object weight table can also be determined based on the user's language habits. For example, if the user is used to stating the object to be generated with a higher object weight first and then the object to be generated with a lower object weight in the prompt, then if there are multiple objects to be generated in the prompt, the object to be generated with the highest object weight can be determined as the main object by its order of appearance. For example, when the prompt is "Generate an image of a cat and bread for me". Among them, the category of "cat" is an animal, and the category of "bread" is food. Since the order of appearance of "cat" is before that of "bread", it can be defaulted that the object weight of "cat" is higher than that of "bread", that is, "cat" is determined as the main object, and the main object category is an animal.
[0130] In this way, the main object in the prompt and the main object category of the main object can be accurately screened out.
[0131] For example, when the content category is the style category, the content weight table is the style weight table, and the style weight is used to represent the influence degree of its corresponding style on the target image. Then, in the case where the prompt includes multiple image styles, query the style weights of each image style in the pre-set style weight table to obtain multiple style weights; determine the image style with the highest style weight among the multiple style weights as the main style of the target image, and determine the category of the main style as the style category of the target image.
[0132] Then, in this embodiment, the pre-set style weight table includes the style weights of each image style. The larger the style weight, the greater the influence of the image style corresponding to the style weight on the target image, that is, the more important the image style is in the target image.
[0133] Similarly, the style weights of each image style in the style weight table can be set in advance by the user's preference for various types of styles. If the user's preference for the anime style is greater than that for the realistic style, and the preference for the realistic style is greater than that for the abstract style, the style weight of the anime style can be set to be greater than the style weight of the realistic style, and the style weight of the realistic style can be set to be greater than the style weight of the abstract style.
[0134] In addition, the style weights of each image style in the style weight table can also be determined based on the user's language habits. The style weight of the image style directly expressed in the prompt can be set to be greater than that of the image style indirectly implied. For example, the prompt can be "Help me generate an anime-style starry sky like Van Gogh's". Among them, the directly expressed image style is "anime style", and the "oil painting style" indirectly implied in the prompt can also be obtained based on the segmented phrase "Van Gogh". However, since the anime style is directly expressed and the oil painting style is indirectly implied, the style weight of the anime style is greater than that of the oil painting style.
[0135] Therefore, if the prompt includes multiple image styles, the style weights of each image style can be queried in the pre-defined style weight table, and then the image style with the highest style weight can be determined as the main style of the target image, and the category of the main style is the style category of the target image.
[0136] Exemplarily, the image style generally refers to the visual appearance and style characteristics of the image. Common style types can include two major categories: realistic styles and artistic styles. Among them, realistic styles can further include realistic figures, landscapes, animals, buildings, etc. Realistic styles emphasize real color expression and strive to restore the real colors and light and shadow effects of the figures. Therefore, they have higher requirements for image quality and greatly test the details of the image. Artistic styles can include styles such as comics, oil paintings, ink paintings, line drawings, etc. They focus more on the display of the painting style and have lower requirements for the details of the image than realistic styles. The weight of realistic styles can be greater than that of artistic styles.
[0137] In this way, the main style in the prompt and the style category of the main style can be accurately screened out.
[0138] As an optional embodiment, the recognition of the image generation instruction and the determination of the content category of the target image based on the recognition result include:
[0139] Recognize the image generation instruction, and determine the prompt corresponding to the image generation instruction based on the recognition result;
[0140] Conduct a security review of the prompt to determine whether the prompt meets the preset text security specification conditions;
[0141] When the prompt meets the text security specification conditions, determine the content category of the target image based on the feature information of the prompt.
[0142] In this embodiment, the image generation instruction can be converted into the corresponding text form, that is, the text content, through semantic recognition. Then, the prompt words related to the target image are extracted from the text content, and the content of the prompt words can be audited to determine whether the prompt words meet the pre-set text security specification conditions. Only when the prompt words meet the text security specification conditions can the content category of the target image be further determined, and the prompt words can be text-expanded.
[0143] Specifically, a sensitive word library can be constructed in advance. The sensitive word library contains multiple keywords related to content such as terrorism, pornography, politics, and hot public opinion. Then, the content of the prompt words is matched with the sensitive word library for keywords. Once a sensitive word is matched in the prompt words, it is considered that the prompt words do not meet the pre-set text security specification conditions, and the image generation instruction needs to be directly filtered out; if no sensitive word is matched in the prompt words, it can be considered that the prompt words meet the pre-set text security specification conditions.
[0144] In this way, it can be ensured that only image generation instructions with high security are executed, thereby ensuring the security of the target image.
[0145] As an alternative embodiment, the image generation model includes a basic image generation model and a specific object generation model, the content category includes a subject category and a style category, and the obtaining of the content enrichment information corresponding to the content category includes:
[0146] Obtain the basic image generation model corresponding to the subject category of the target image, and the specific object generation model corresponding to the style category of the target image;
[0147] Query the subject keyword that has a mapping relationship with the basic image generation model in the pre-set subject mapping table, and determine the subject keyword as the content enrichment information corresponding to the subject category;
[0148] Query the style keyword that has a mapping relationship with the specific object generation model in the pre-set style mapping table, and determine the style keyword as the content enrichment information corresponding to the style category.
[0149] In this embodiment, the image generation model can include a basic image generation model and a specific object generation model. Each basic image generation model is used to implement the text-to-image task of one or more categories of subjects, and each specific object generation model is used to implement the text-to-image task of one or more categories of image styles. Therefore, the corresponding basic image generation model can be determined based on the subject category, and the specific object generation model corresponding to the style category.
[0150] For an image generation model, a main keyword mapping table associated with the basic image generation model can be predefined. In the main keyword mapping table, each basic image generation model has main keywords with which it has a mapping relationship. After determining the basic image generation model, the main keywords that have a mapping relationship with the determined basic image generation model can be retrieved from the main keyword mapping table, and these main keywords can be added to the prompt as content enrichment information.
[0151] Exemplarily, when the main subject of the target image is a kitten, then the main subject category is an animal, and the determined basic image generation model is the basic image generation model corresponding to animals. Thus, in the main keyword mapping table associated with the basic image generation model, the main keywords corresponding to "kitten" can be queried, such as "fluffy, cute, lively", etc. These words are used as content enrichment information. In this way, the content of the prompt can be enriched, and further enrich the content of the target image, making the details of the target image richer and more accurately expressing the user's needs.
[0152] Exemplarily, since each basic image generation model corresponds to at least one main subject type, and each specific object generation model corresponds to one style type, the main keywords that have a mapping relationship with the basic image generation model are used to describe the main subject type corresponding to the basic image generation model, and the style keywords that have a mapping relationship with the specific object generation model are used to describe the image style type corresponding to the specific object generation model.
[0153] Similarly, a style keyword mapping table associated with the specific object generation model can be predefined. In the style keyword mapping table, each specific object generation model has style keywords with which it has a mapping relationship. After determining the specific object generation model, the style keywords that have a mapping relationship with the determined specific object generation model can be retrieved from the style keyword mapping table, and these style keywords can be added to the prompt as content enrichment information.
[0154] In the embodiments of the present application, by adding the keywords corresponding to the model to the prompt, the text content of the prompt is made richer, and it is convenient for the model to better understand the content category corresponding to the text content of the input model, thereby improving the accuracy of the generated image.
[0155] As an optional embodiment, before determining the content category of the target image based on the recognition result, the method further includes:
[0156] When it is detected that there is a special word in the prompt corresponding to the image generation instruction, query the general word corresponding to the special word;
[0157] Use the general word to replace the corresponding special word in the prompt.
[0158] In this embodiment, special words include regional slang, common sayings, ancient Chinese poems, proverbs, Internet terms, professional terms, scientific nouns, and so on. These special words may be commonly used language expressions in a certain region or community, or may be obscure industry terms, and usually have a certain regionality and limitation. Image generation models usually cannot understand these special words.
[0159] A special word mapping table can be set in advance. In the special word mapping table, each special word has a corresponding general word with a mapping relationship, and the general word is an easy-to-understand and popular explanation of the special word.
[0160] Therefore, whenever a special word is detected in the prompt, the general word with a mapping relationship to the special word can be queried in the special word mapping table, and then the corresponding special word in the prompt is replaced with the general word.
[0161] Exemplarily, the special words can be pre-classified into multiple dimensions in the special word mapping table, such as regional special words, industry special words, and community special words. For example, "slang, common sayings, ancient Chinese poems, and proverbs" can be determined as regional special words, and "professional terms and scientific nouns" can be determined as industry special words. In this way, the dimension to which the special word belongs can be first queried in the special word table, and then the general word with a mapping relationship to the special word can be further queried.
[0162] Exemplarily, when the prompt is "A cute tabby cat with a round face and chubby cheeks in an anime style is sunbathing on a lawn full of flowers", where "round face and chubby cheeks" is a Chinese slang and common saying, that is, a special word. Therefore, the general word "round face, chubby" with a first mapping relationship to "round face and chubby cheeks" can be queried, and thus the prompt can be converted to "A cute tabby cat with a round face and chubby in an anime style is sunbathing on a lawn full of flowers".
[0163] The embodiment of the present application converts the obscure special words in the text content to be input into the model, facilitating the model to better understand the text content input into the model, thereby improving the accuracy of the generated images.
[0164] As an alternative embodiment, the using the image generation model corresponding to the content category to convert the final text into a target image includes:
[0165] Obtaining the basic image generation model and the specific object generation model corresponding to the content category;
[0166] Mount the specific object generation model onto the base image generation model to obtain a fused image generation model, where the image generation model includes the base image generation model, the specific object generation model, and the fused image generation model;
[0167] According to the task type corresponding to the image generation instruction, use the fused image generation model to convert the final text into a target image corresponding to the target parameters, where the target parameters refer to the target size and / or target resolution corresponding to the task type.
[0168] In this embodiment, the specific object generation model cannot be used alone. The specific object generation model needs to be mounted on the base image generation model to assist the base image generation model in image generation.
[0169] After mounting the specific object generation model onto the base image generation model to obtain a fused image generation model, a random noise image can be obtained, and the final text can be converted into a text embedding vector. The text embedding vector is embedded into the fused image generation model loaded with a sampler to guide the image generation model to perform multiple iterations of noise reduction on the noise image to obtain a first intermediate image.
[0170] Then, based on the target format and target size in the pre-specified target parameters, the first intermediate image can be further edited. Specifically, since the image format of the first intermediate image output by the image generation model is often not the target format, and the image size of the first intermediate image is also different from the target size. Therefore, it is necessary to edit the first intermediate image to convert the image format of the first intermediate image into the target format and adjust the image size to the target size, so as to obtain a second intermediate image.
[0171] Exemplarily, the target image can be a wallpaper that is full-screen displayed on the central control display screen of a vehicle, or a background image of a vehicle control program in the vehicle. In order to adapt to the size and aspect ratio of the display screen or the vehicle control program display frame, the target size of the target image needs to be pre-specified.
[0172] In addition, since the operation of the image generation model requires a large power consumption, and the computing power that the vehicle controller can allocate to the image generation model is limited, the resolution and details of the second intermediate image generated by the image generation model are often poor. Based on this, the second intermediate image can be semantically segmented, the main area of the second intermediate image can be determined based on the result of the semantic segmentation, and the image details can be repaired in the main area to make the resolution of the image meet the target resolution in the target parameters, so as to obtain the final target image.
[0173] Exemplarily, semantic segmentation can label the pixels in an image as belonging to specific categories to identify and distinguish different objects, structures, or regions in the second intermediate image. Exemplarily, the object in the second intermediate image that is centered and has the largest coverage area can be determined as the main body of the second intermediate image, and the region where the main body is located is the main body region of the second intermediate image. Then, only the main body region is repaired in detail.
[0174] Exemplarily, the second intermediate image can be semantically segmented into multiple regions by the SAM (Segment Anything Model) segmentation large model, and the region that is centered and has the largest coverage area among the multiple regions is determined as the main body region.
[0175] As an alternative embodiment, the converting the final text into a first intermediate image by using the image generation model includes:
[0176] Obtain a noise image;
[0177] Encode the final text to obtain a text embedding vector;
[0178] Use the image generation model embedding the text embedding vector to perform noise reduction processing on the noise image to obtain a third latent feature image;
[0179] Decode the third latent feature image to obtain the first intermediate image.
[0180] In this embodiment, during the process of converting text into an image, a randomly generated noise image can be obtained, and the text content of the prompt is mapped into a high-dimensional vector space to obtain a text embedding vector. After obtaining the text embedding vector, the text embedding vector and the noise image can be jointly input into a fusion model, and the text embedding vector is used to guide the image generation model to perform noise reduction processing on the noise image to obtain a third latent feature image. Then, the third latent feature image can be decoded by using an autoencoder image generation model to obtain a target image that can be recognized by the human eye.
[0181] In this embodiment, the conversion from the prompt to the target image can be accurately completed.
[0182] As an alternative embodiment, the converting the final text into a target image by using the image generation model corresponding to the content category includes:
[0183] Perform a security review on the final text to determine whether the final text meets the preset text security specification conditions;
[0184] If the final text meets the text security specification conditions, use the image generation model corresponding to the content category to convert the final text into the target image;
[0185] If there are sensitive words in the final text that do not meet the text security specification conditions, delete the sensitive words in the final text, and use the image generation model corresponding to the content category to convert the final text after deleting the sensitive words into the target image.
[0186] In this embodiment, after converting the image generation instruction into the final text, the content of the final text can be audited to determine whether the final text meets the preset text security specification conditions. Only when the final text meets the text security specification conditions can the final text be converted into the target image. If the final text does not meet the text security specification conditions, it is necessary to delete the sensitive words in the final text and convert the final text after deleting the sensitive words into the security specification conditions.
[0187] Specifically, a sensitive word library can be constructed in advance. The sensitive word library contains multiple keywords related to content such as terrorism, pornography, politics, and hot public opinion. Then, match the content of the final text with the sensitive word library. Once a sensitive word is matched in the final text, it is considered that the final text does not meet the preset text security specification conditions, and the sensitive words in the final text need to be deleted; if no sensitive word is matched in the final text, it can be considered that the final text meets the preset text security specification conditions.
[0188] In this way, the privacy and security of the generated image can be guaranteed.
[0189] As an optional embodiment, the image size includes the image length and the image width, the target size includes the target length and the target width, and the image editing of the first intermediate image to obtain the second intermediate image includes:
[0190] In the case where the image length of the first intermediate image is less than the target length or the image width of the first intermediate image is less than the target width, crop the first intermediate image to obtain an image to be border-expanded, where the image length of the image to be border-expanded is less than or equal to the target length and the image width of the border-expanded image is less than or equal to the target width;
[0191] Use blank pixels to perform edge expansion on the image to be border-expanded to obtain an expanded image, where the image length of the expanded image is equal to the target length and the image width of the expanded image is equal to the target width;
[0192] Redraw the extended area of the extended image to obtain the second intermediate image, where the extended area is the area filled with blank pixels.
[0193] In this embodiment, if the image length of the first intermediate image is less than the target length, or the image width of the first intermediate image is less than the target width, it means that it is impossible to obtain an image with the target size only by cropping the first intermediate image.
[0194] If the image length of the first intermediate image is less than the target length and the image width of the first intermediate image is less than the target width, the first intermediate image can be directly used as the image to be border-extended; if the image length of the first intermediate image is greater than the target length, or the image width of the first intermediate image is greater than the target width, the first intermediate image needs to be cropped to obtain an image to be border-extended with an image length less than or equal to the target length and an image width less than or equal to the target width.
[0195] Specifically, the main body in the first intermediate image can be recognized through image segmentation or object detection, and according to the results of object detection or image segmentation, a cropping area containing the main body can be determined, and the main body is made to be located as much as possible in the middle of the cropping area, and the length of the cropping area is less than the target length and the width is less than the target width. Then, the cropping area can be cropped out in the first intermediate image, and the cropping area is determined as the image to be border-extended.
[0196] Among them, the target length can be the display area length corresponding to the inner screen of the vehicle-mounted screen, and the target width can be the display area width corresponding to the inner screen of the vehicle-mounted screen.
[0197] Subsequently, blank pixels without content can be applied to splice at the edge part of the image to be border-extended to achieve the edge extension of the image to be border-extended and obtain an extended image that meets the target size. Among them, the extended area outside the image to be border-extended in the extended image is composed of blank pixels without content. To improve the image quality, the blank pixels in the extended area can be redrawn according to the user's requirements to obtain a second intermediate image that matches the target size.
[0198] In this embodiment, through the cropping and extension of the image, the image output by the image generation model can meet the target size required by the target image.
[0199] As an alternative embodiment, the redrawing the extended area of the extended image to obtain the second intermediate image includes:
[0200] Obtain the description information of the first intermediate image;
[0201] Use the grayscale mask image to mask the extended image;
[0202] Add noise to the expanded area according to the pixel values of each pixel point in the expanded area after mask processing;
[0203] Perform noise reduction processing on the expanded area after mask processing based on the description information of the first intermediate image to obtain a third intermediate image;
[0204] Smooth the transition area between the expanded area and the initial area in the third intermediate image to obtain the second intermediate image, where the initial area is the area outside the transition area in the expanded image.
[0205] In this embodiment, the description information of the first intermediate image can represent the image content of the first intermediate image. Specifically, computer vision extraction techniques can be used to extract the features of the first intermediate image, and then these features can be converted into natural language descriptions, that is, the description information of the first intermediate image is obtained. It is also possible to directly use the final text as the description information of the first intermediate image.
[0206] Then, after determining the expanded area of the expanded image, a grayscale mask image can be masked on the expanded image. The part of the grayscale mask image that masks the initial area is a black area with a pixel value of 0, and the pixel values of the part of the grayscale mask image that masks the expanded area are not 0.
[0207] Then, by obtaining a pre-set random number seed, convert the random number seed into an initialized Gaussian noise, and then add the Gaussian noise to the expanded image. Since the pixel value of the initial area is 0, it is impossible to add noise to the initial area. The noise can be added to each pixel in the expanded area based on the pixel values of each pixel point in the expanded area. The larger the pixel value, the more noise is added, indicating a greater redrawing amplitude.
[0208] After adding noise to the expanded area in the first intermediate image, the expanded area with added noise can be subjected to noise reduction processing to complete the redrawing of the expanded area. Specifically, the expanded image with added noise and the description information of the first intermediate image can be input into the image-to-image generation model. The description information can guide the image-to-image generation model to predict the noise on the expanded image, and use the loaded sampler to denoise the predicted noise, thereby completing the redrawing of the expanded image, obtaining a third intermediate image, and smoothing the transition area between the expanded area and the initial area in the third intermediate image to obtain the second intermediate image.
[0209] Exemplarily, the image-to-image generation model can be the base model in the latent diffusion model (Stable Diffusion model).
[0210] In this embodiment, by adding noise and denoising to the extended area composed of blank pixels, the redrawing of the extended area can be completed with only a small amount of computing power, so that the image generated by the model can match the required target size.
[0211] As an alternative embodiment, the smoothing process on the transition area between the extended area and the initial area in the third intermediate image to obtain the second intermediate image includes:
[0212] Adding noise to the transition area in the third intermediate image;
[0213] Performing noise reduction processing on the transition area in the third intermediate image based on the description information of the first intermediate image to obtain the second intermediate image.
[0214] In this embodiment, after redrawing the extended area of the first intermediate image to obtain the third intermediate image, the visual connection or transition between the redrawn area and the non-redrawn area in the third intermediate image may not be natural or smooth enough, and obvious edges, discontinuous transitions, or differences in color, brightness, etc. may occur.
[0215] By adding noise and denoising to the transition area in the third intermediate image, the details of the transition area in the third intermediate image can be repaired, making the connection of the whole image smooth and natural, and obtaining the second intermediate image.
[0216] Exemplarily, the transition area in the third intermediate image can be repaired by adding noise and denoising to the whole third intermediate image.
[0217] In this embodiment, by adding noise and denoising to the partially redrawn image, the detail repair of the image can be completed, making the connection of the whole image smooth and natural, and improving the quality of the image.
[0218] As an alternative embodiment, the image editing of the first intermediate image to obtain the second intermediate image includes:
[0219] When the image length of the first intermediate image is greater than the target length and the image width of the first intermediate image is greater than the target width, cropping the first intermediate image according to the target size to obtain the second intermediate image.
[0220] In this embodiment, if the image length of the first intermediate image is greater than the target length and the image width of the first intermediate image is greater than the target width, then the first intermediate image output by the image generation model can be directly cropped so that the length of the cropped image is the target length and the width of the cropped image is the target width. Thus, the second intermediate image can be directly obtained by cropping.
[0221] Exemplarily, the main body in the first intermediate image can be recognized through image segmentation or object detection, and according to the results of object detection or image segmentation, a cropping area containing the main body can be determined, and the main body is made to be located as much as possible in the middle of the cropping area, and the size of the cropping area meets the target size. Then, the cropping area can be cropped out from the first intermediate image, and the cropping area is determined as the second intermediate image.
[0222] In the embodiments of the present application, the relatively large-sized image output by the image generation model can be cropped so that the cropped image meets the target size specified by the target image.
[0223] As an optional embodiment, the repairing the image details of the second intermediate image to obtain the target image includes:
[0224] Obtaining the description information of the second intermediate image;
[0225] Adding noise to the main body area of the second intermediate image;
[0226] Performing noise reduction processing on the main body area with added noise in the second intermediate image based on the description information of the second intermediate image to obtain the target image.
[0227] In this embodiment, the description information of the second intermediate image can characterize the image content of the second intermediate image. Specifically, computer vision extraction techniques can be used to extract the features of the second intermediate image, and then these features are converted into natural language descriptions, that is, the description information of the second intermediate image is obtained. The final text can also be directly used as the description information of the second intermediate image.
[0228] After determining the main body area of the second intermediate image, by obtaining a preset random number seed, the random number seed can be converted into an initialized Gaussian noise, and then the Gaussian noise is added to each pixel in the main body area.
[0229] After adding noise to the main body area in the second intermediate image, the main body area with added noise can be subjected to noise reduction processing to complete the enrichment of image details. Specifically, the second intermediate image with added noise can be input into the image-to-image model, and the image-to-image model can predict the noise on the second intermediate image and use the loaded sampler to denoise the predicted noise, so as to obtain the target image with detailed repair.
[0230] Exemplarily, the image-to-image model can be the base model in the latent diffusion model (Stable Diffusion model).
[0231] In this embodiment, after determining the main region of the second intermediate image, noise can be added to and removed from only the main region, enriching the details of the main subject in the image with relatively low computing power and power consumption and improving the image quality at a relatively fast processing speed.
[0232] As an alternative embodiment, adding noise to the main region of the second intermediate image includes:
[0233] Masking the non-main region of the second intermediate image with a black mask image having pixel values of 0, where the non-main region is the region outside the main region in the second intermediate image;
[0234] Adding noise to the second intermediate image after the masking process;
[0235] Performing noise reduction on the main region with added noise in the second intermediate image based on the description information of the second intermediate image to obtain the target image, including:
[0236] Encoding the second intermediate image after the masking process to obtain a first latent feature image;
[0237] Performing noise reduction on the first latent feature image based on the description information of the second intermediate image to obtain a second latent feature image;
[0238] Decoding the second latent feature image to obtain the target image.
[0239] In this embodiment, after determining the main region in the second intermediate image, the region outside the main region in the second intermediate image can be determined as the non-main region. Since in the second intermediate image, the main region is the region where the main subject in the image is located, and the main subject is the object that needs to be prominently shown in the image, the main region has relatively high requirements for the picture quality and image details, while the non-main region has relatively low requirements for the picture quality and image details.
[0240] To save the computing power of the model, only the main region in the second intermediate image needs to be detail-repaired, and the non-main region does not need to be detail-repaired. Therefore, the non-main region can be masked with a black mask image having pixel values of 0, that is, the values of all pixel points in the non-main region are set to 0.
[0241] After masking the non-main region, when adding noise to the second intermediate image, the masked non-main region will not be affected by the noise. That is to say, when adding noise to the second intermediate image after the masking process, the noise will only be added to the main part of the second intermediate image.
[0242] In the above manner, it is possible to quickly and accurately add noise only to the main region in the second intermediate image.
[0243] After adding noise to the second intermediate image after mask processing, an image with blurred noise can be obtained, and the blurred area is the main region of the image. Subsequently, the second intermediate image with added noise can be encoded by an autoencoder, that is, the second intermediate image with added noise is mapped into the latent space to generate the first latent feature image.
[0244] Subsequently, the first latent feature image can be input into the base model of the latent diffusion model (Stable Diffusion model). Since the non-main regions have been masked, after loading a suitable sampler in the base model, only the noise of the main region can be predicted, and the main region in the first latent feature image can be denoised based on the predicted result to obtain the second feature image.
[0245] After obtaining the second feature image, the second feature image can be decoded using a variational autoencoder to obtain the visible target image.
[0246] In this embodiment, it is possible to mask the regions outside the main region that need to be detailedly repaired and super-resolved, so that the non-main regions do not participate in the noise addition and denoising processes, and only a small amount of computing power is required to complete the detail enrichment of the specified region.
[0247] As an alternative embodiment, before repairing the image details of the second intermediate image to obtain the target image, the method further includes:
[0248] Using interpolation technology, the second intermediate image is expanded in the length and width directions respectively;
[0249] Image enhancement is performed on the second intermediate image after image expansion through a deblurring algorithm.
[0250] In this embodiment, since the operation of the image generation model consumes a large amount of computing power, the resolution of the intermediate image output by the image generation model is usually low. The intermediate image can be super-resolved by a trained super-resolution model to improve the overall resolution of the intermediate image and obtain a clearer image.
[0251] Exemplarily, the super-resolution model can be an Enhanced Super-Resolution Generative Adversarial Network (ESRGAN), and ESRGAN is used to perform 2X super-resolution on the intermediate image. Specifically, ESRGAN can use deep learning techniques, a generative adversarial network, to improve the spatial resolution of the intermediate image. And 2X super-resolution can generate an image with a resolution twice that of the intermediate image in both the horizontal and vertical directions by processing the second intermediate image, thereby improving the visual quality and details of the image.
[0252] As an alternative embodiment, the super-resolution model can be a Residual in Residual Dense Block (RRDB) module without Batch Normalization (BN). Exemplarily, the RRDB module can include three interconnected Dense Blocks, and each Dense Block consists of five convolutional layers. These convolutional layers may have different filters and feature map depths for learning different levels of image representations. In this way, the super-resolution model has better generalization.
[0253] As an alternative embodiment, after repairing the image details of the second intermediate image to obtain the target image, the method further includes:
[0254] Storing the target image when the target image meets the preset image security specification conditions.
[0255] In this embodiment, after generating the final target image, it is also necessary to review the target image to determine whether the target image meets the pre-set image security specification conditions. Only when the target image meets the image security specification conditions can the target image be stored according to the specified path, otherwise the target image is filtered.
[0256] Specifically, a hash value of the image can be generated in advance using a hash function to establish a hash database of sensitive content. The hash database contains multiple sensitive hash data related to content such as terrorism, politics, pornography, nudity, hot public opinion, religion, specific IPs or persons, etc. Then, the hash database is used to match the target image. Once there is sensitive hash data in the hash database that matches the target image, it is considered that the target image does not meet the preset image security specification conditions; if there is no sensitive hash data in the hash database that matches the target image, it can be considered that the target image meets the preset image security specification conditions.
[0257] In this way, the privacy and security of the generated image can be guaranteed.
[0258] Based on the image generation method provided in the above embodiments, correspondingly, the present application also provides a specific implementation manner of an image generation device. Please refer to the following embodiments.
[0259] First, refer to Figure 2 The image generation device 200 provided in the embodiments of the present application includes the following modules:
[0260] An acquisition module 201, configured to acquire an image generation instruction of a user, where the image generation instruction is used to indicate generating a target image;
[0261] An identification module 202, configured to identify the image generation instruction, determine a content category of the target image based on the identification result, and acquire content enrichment information corresponding to the content category;
[0262] An expansion module 203, configured to use the content enrichment information to perform text expansion on a prompt word corresponding to the image generation instruction to obtain a final text;
[0263] A conversion module 204, configured to use an image generation model corresponding to the content category to convert the final text into a target image.
[0264] The device can identify the image generation instruction of the user, then determine the content category of the target image indicated by the image generation instruction based on the identification result, and use the content enrichment information corresponding to the content category to perform text expansion on the prompt word corresponding to the image generation instruction to obtain a final text, and then use the image generation model corresponding to the content category to convert the final text into a target image. In this way, the complete requirements of the user for the image content can be more clearly expressed through the expanded final text, and the complete requirements of the user for the image content can be more accurately understood through the image generation model corresponding to the content category, so as to obtain a target image that better meets the user's requirements and improve the quality of the target image generated in the vehicle.
[0265] As an implementation manner of the present application, the above identification module 202 may further include:
[0266] A first identification unit, configured to identify the image generation instruction to obtain text content corresponding to the image generation instruction;
[0267] An understanding unit, configured to perform semantic understanding on the text content to obtain a user intention; and extract a prompt word for generating a target image from the user intention;
[0268] A first determination unit, configured to determine the content category of the target image based on feature information in the prompt word.
[0269] As an implementation manner of the present application, the content category includes a main category, and the recognition module 202 may further include:
[0270] A second determination unit, configured to determine a content phrase in the prompt word based on the user intention, where the content phrase is used to indicate a content element generated in the target image;
[0271] A first query unit, configured to, when there are multiple content phrases in the prompt word, query the content weights of the content elements corresponding to each content phrase in a preset content weight table to obtain multiple content weights;
[0272] A third determination unit, configured to determine the content element with the highest content weight among the multiple content weights as the main content of the target image, and determine the category of the main content as the content category of the target image.
[0273] As an implementation manner of the present application, the recognition module 202 may further include:
[0274] A second recognition unit, configured to recognize the image generation instruction, and determine the prompt word corresponding to the image generation instruction based on the recognition result;
[0275] A first review unit, configured to perform a security review on the prompt word to determine whether the prompt word meets the preset text security specification conditions;
[0276] A fourth determination unit, configured to, when the prompt word meets the text security specification conditions, determine the content category of the target image based on the feature information of the prompt word.
[0277] As an implementation manner of the present application, the image generation model includes a basic image generation model and a specific object generation model, and the acquisition module 201 may further include:
[0278] A first acquisition unit, configured to acquire the basic image generation model corresponding to the main category of the target image, and the specific object generation model corresponding to the style category of the target image;
[0279] A second query unit, configured to query, in a preset main body mapping table, a main body keyword having a mapping relationship with the basic image generation model, and determine the main body keyword as the content enrichment information corresponding to the main category;
[0280] A third query unit, configured to query, in a preset style mapping table, a style keyword having a mapping relationship with the specific object generation model, and determine the style keyword as the content enrichment information corresponding to the style category.
[0281] As an implementation manner of the present application, the above-mentioned image generation device 200 may further include:
[0282] A query module, configured to query a general vocabulary corresponding to the special vocabulary when it is detected that the special vocabulary exists in the final text;
[0283] A replacement module, configured to replace the corresponding special vocabulary in the final text with the general vocabulary.
[0284] As an implementation manner of the present application, the above-mentioned conversion module 204 may further include:
[0285] A second acquisition unit, configured to acquire a basic image generation model and a specific object generation model corresponding to the content category;
[0286] A mounting unit, configured to mount the specific object generation model on the basic image generation model to obtain a fused image generation model, where the image generation model includes the basic image generation model, the specific object generation model, and the fused image generation model;
[0287] A first conversion unit, configured to convert the final text into a target image corresponding to the target parameter according to the task type corresponding to the image generation instruction, where the target parameter refers to the target size and / or target resolution corresponding to the task type.
[0288] As an implementation manner of the present application, the above-mentioned first conversion unit may further include:
[0289] A first conversion subunit, configured to use the final text to guide the fused image generation model to convert a randomly generated noise image into a first intermediate image;
[0290] An editing subunit, configured to perform image editing on the first intermediate image to obtain a second intermediate image, where the image size of the second intermediate image matches the target size corresponding to the task type;
[0291] A repair subunit, configured to repair the image details of the second intermediate image based on the target resolution corresponding to the task type to obtain the target image.
[0292] As an implementation manner of the present application, the above-mentioned conversion module 204 may further include:
[0293] A second audit unit, configured to perform a security audit on the final text to determine whether the final text meets the preset text security specification conditions;
[0294] A second conversion unit, configured to convert the final text into the target image by using the image generation model corresponding to the content category if the final text meets the text security specification conditions;
[0295] A third conversion unit, configured to delete the sensitive words in the final text if there are sensitive words in the final text that do not meet the text security specification conditions, and convert the final text after deleting the sensitive words into the target image by using the image generation model corresponding to the content category.
[0296] As an implementation manner of the present application, the image size includes an image length and an image width, the target size includes a target length and a target width, and the above-mentioned editing subunit may further include:
[0297] A first cropping subunit, configured to crop the first intermediate image to obtain an image to be border-expanded if the image length of the first intermediate image is less than the target length or the image width of the first intermediate image is less than the target width, where the image length of the image to be border-expanded is less than or equal to the target length, and the image width of the border-expanded image is less than or equal to the target width;
[0298] An image expansion subunit, configured to perform edge expansion on the image to be border-expanded by using blank pixels to obtain an expanded image, where the image length of the expanded image is equal to the target length, and the image width of the expanded image is equal to the target width;
[0299] A redrawing subunit, configured to redraw the expanded area of the expanded image to obtain the second intermediate image, where the expanded area is the area filled with the blank pixels.
[0300] As an implementation manner of the present application, the above-mentioned redrawing subunit may further be configured to:
[0301] Obtain the description information of the first intermediate image;
[0302] Perform masking processing on the expanded image by using a grayscale mask image;
[0303] Add noise to the expanded area according to the pixel values of each pixel point in the expanded area after the masking processing;
[0304] Perform noise reduction processing on the expanded area after the masking processing based on the description information of the first intermediate image to obtain a third intermediate image;
[0305] Perform smoothing processing on the transition area between the expanded area and the initial area in the third intermediate image to obtain the second intermediate image, where the initial area is the area outside the transition area in the expanded image.
[0306] As an implementation manner of the present application, the above repair subunit may further include:
[0307] An acquisition subunit, configured to acquire the description information of the second intermediate image;
[0308] A noise addition subunit, configured to add noise to the main body area of the second intermediate image;
[0309] A noise reduction subunit, configured to perform noise reduction processing on the main body area with added noise in the second intermediate image based on the description information of the second intermediate image to obtain the target image.
[0310] As an implementation manner of the present application, the above image generation device 200 may further include:
[0311] An interpolation module, configured to perform image expansion on the second intermediate image along the length direction and the width direction respectively;
[0312] A deblurring module, configured to perform image enhancement on the second intermediate image after image expansion through a deblurring algorithm.
[0313] The image generation device provided by the embodiment of the present invention can implement each step in the above method embodiment. To avoid repetition, it will not be elaborated here.
[0314] The embodiment of the present application further provides an image generation system, and the system includes:
[0315] A server, configured to acquire an image generation instruction of a user, where the image generation instruction is used to indicate generating a target image;
[0316] An application end, configured to identify the image generation instruction, determine the content category of the target image based on the recognition result, and acquire the content enrichment information corresponding to the content category;
[0317] The application end is further configured to use the content enrichment information to perform text expansion on the prompt word corresponding to the image generation instruction to obtain a final text;
[0318] The application end is further configured to convert the final text into a target image corresponding to target parameters according to the task type corresponding to the image generation instruction, where the target parameters refer to the target size and / or target resolution corresponding to the task type.
[0319] In some embodiments, the server is further configured to:
[0320] Receive a voice instruction of a user and send the voice instruction to the application end;
[0321] Receive the final text obtained by converting the voice command from the application end;
[0322] Display the final text on the display screen of the vehicle;
[0323] Send the final text to the application end and receive the target image generated by the application end based on the final text;
[0324] Display the target image on the display screen of the vehicle.
[0325] Among them, the server in the image generation system is located at the vehicle end, and the application end can be set in the cloud or at the vehicle end. The application end and the server are communicatively connected through an intermediate layer. Algorithms related to large models can be deployed in the application end, and the server can support the interface display and interaction of the in-vehicle screen in the vehicle cabin.
[0326] Specifically, the server can include a voice program and a drawing program. The voice program receives the user's voice command, which is transferred through the intermediate layer and sent to the ASR algorithm of the application end for processing to obtain a prompt word, and the processed prompt word is sent to the application end for on-screen display.
[0327] After that, the voice program sends the final text to the LLM algorithm of the application end. The LLM algorithm processes the final text, converts the prompt word into the final text, and sends the final text to the server.
[0328] The voice program on the server side forwards the final text to the drawing program, which invokes the drawing program. After the drawing program obtains the final text, it calls the text-to-image related algorithm of the application end for drawing to obtain the target image. After content review, the target image is returned to the drawing program for display.
[0329] The image generation system provided by the embodiments of the present invention can implement each step in the above method embodiments. To avoid repetition, it will not be elaborated here.
[0330] Figure 3 Shows the hardware structure schematic diagram of the image generation device provided by the embodiments of the present application.
[0331] The image generation device may include a processor 301 and a memory 302 storing computer program instructions.
[0332] Specifically, the above-mentioned processor 301 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0333] The memory 302 may include a mass storage for data or instructions. By way of example and not limitation, the memory 302 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. In a suitable case, the memory 302 may include removable or non-removable (or fixed) media. In a suitable case, the memory 302 may be inside or outside the integrated gateway disaster recovery device. In a specific embodiment, the memory 302 is a non-volatile solid-state memory.
[0334] The memory may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk storage media device, an optical storage media device, a flash memory device, an electrical, optical, or other physical / tangible memory storage device. Thus, generally, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to an aspect of the present disclosure.
[0335] The processor 301 reads and executes the computer program instructions stored in the memory 302 to implement any one of the image generation methods in the above embodiments.
[0336] In one example, the image generation device may further include a communication interface 303 and a bus 310. Among them, as Figure 3 shown, the processor 301, the memory 302, and the communication interface 303 are connected through the bus 310 to complete the communication with each other.
[0337] The communication interface 303 is mainly used to implement the communication between the modules, devices, units, and / or devices in the embodiments of the present application.
[0338] The bus 310 includes hardware, software, or both, and couples the components of the image generation device to each other. By way of example and not limitation, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front-Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses or a combination of two or more of these. Where appropriate, the bus 310 may include one or more buses. Although the embodiments of the present application describe and illustrate specific buses, the present application contemplates any suitable bus or interconnect.
[0339] The image generation device may be based on the above embodiments, so as to implement the image generation method and device in combination with the above.
[0340] In addition, in combination with the image generation method in the above embodiments, the embodiments of the present application may provide a computer storage medium to implement. Computer program instructions are stored on the computer storage medium; when the computer program instructions are executed by a processor, any one of the image generation methods in the above embodiments is implemented, and the same technical effects can be achieved. For the sake of brevity, details are not repeated here. Among them, the above computer-readable storage medium may include a non-transitory computer-readable storage medium, such as a Read-Only Memory (ROM), a Random Access Memory (RAM), a magnetic disk, or an optical disc, etc., which are not limited herein.
[0341] In addition, the embodiments of the present application also provide a vehicle, including computer program instructions, which can implement the steps and corresponding contents of the foregoing method embodiments when executed by a processor.
[0342] It should be clear that the present application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present application.
[0343] The functional blocks shown in the above structural block diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, and so on. When implemented in software, the elements of the present application are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted via a data signal carried in a carrier wave on a transmission medium or a communication link. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical discs, hard disks, fiber optic media, radio frequency (RF) links, and so on. The code segment can be downloaded via a computer network such as the Internet, an intranet, and so on.
[0344] It should also be noted that the exemplary embodiments mentioned in the present application describe some methods or systems based on a series of steps or devices. However, the present application is not limited to the order of the above steps, that is, the steps can be executed in the order mentioned in the embodiments, can be different from the order in the embodiments, or several steps can be executed simultaneously.
[0345] As described above with reference to the flowcharts and / or block diagrams of the methods, apparatuses, and vehicles according to embodiments of the present disclosure. It should be understood that each block in the flowchart and / or block diagram, and the combination of blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to generate a machine such that the instructions executed by the processor of the computer or other programmable data processing device enable the implementation of the functions / actions specified in one or more blocks of the flowchart and / or block diagram. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field programmable logic circuit. It can also be understood that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can also be implemented by dedicated hardware that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0346] The above is only the specific implementation manner of the present application. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, modules, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be repeated here. It should be understood that the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present application.
Claims
1. An image generation method, characterized in that, The method includes: Obtaining an image generation instruction of a user, where the image generation instruction is used to indicate generating a target image; Identifying the image generation instruction, determining a content category of the target image based on the identification result, and obtaining content enrichment information corresponding to the content category; Using the content enrichment information to perform text expansion on a prompt word corresponding to the image generation instruction to obtain a final text; Using an image generation model corresponding to the content category to convert the final text into a target image.
2. The image generation method according to claim 1, characterized in that, The identifying the image generation instruction and determining the content category of the target image based on the identification result includes: Identifying the image generation instruction to obtain text content corresponding to the image generation instruction; Performing semantic understanding on the text content to obtain a user intention; and extracting a prompt word for generating the target image from the user intention; Determining the content category of the target image based on feature information in the prompt word.
3. The image generation method according to claim 1, characterized in that, The identifying the image generation instruction and determining the content category of the target image based on the identification result includes: Identifying the image generation instruction and determining content phrases in a prompt word corresponding to the image generation instruction based on the identification result, where the content phrases are used to indicate content elements to be generated in the target image; In a case where there are multiple content phrases in the prompt word, querying content weights of content elements corresponding to each content phrase in a preset content weight table to obtain multiple content weights; Determining a content element with the highest content weight among the multiple content weights as the main content of the target image, and determining a category of the main content as the content category of the target image.
4. The image generation method according to claim 1, characterized in that, The identifying the image generation instruction and determining the content category of the target image based on the identification result includes: Identifying the image generation instruction and determining a prompt word corresponding to the image generation instruction based on the identification result; Performing a security review on the prompt word to determine whether the prompt word meets preset text security specification conditions; In a case where the prompt word meets the text security specification conditions, determining the content category of the target image based on feature information of the prompt word.
5. The image generation method according to claim 1, characterized in that, The content category includes a subject category and a style category, and the obtaining content enrichment information corresponding to the content category includes: Obtaining a basic image generation model corresponding to the subject category of the target image and a specific object generation model corresponding to the style category of the target image; Querying a subject keyword having a mapping relationship with the basic image generation model in a preset subject mapping table, and determining the subject keyword as the content enrichment information corresponding to the subject category; Querying a style keyword having a mapping relationship with the specific object generation model in a preset style mapping table, and determining the style keyword as the content enrichment information corresponding to the style category.
6. The image generation method according to claim 1, characterized in that, Before determining the content category of the target image based on the identification result, the method further includes: When it is detected that there are special words in the prompt corresponding to the image generation instruction, query the general words that have a corresponding relationship with the special words; Use the general words to replace the corresponding special words in the prompt.
7. The image generation method according to claim 1, characterized in that, The converting the final text into a target image by using the image generation model corresponding to the content category includes: Conduct a security review on the final text to determine whether the final text meets the preset text security specification conditions; If the final text meets the text security specification conditions, use the image generation model corresponding to the content category to convert the final text into the target image; If there are sensitive words in the final text that do not meet the text security specification conditions, delete the sensitive words in the final text, and use the image generation model corresponding to the content category to convert the final text after deleting the sensitive words into the target image.
8. The image generation method according to claim 1, characterized in that, The converting the final text into a target image by using the image generation model corresponding to the content category includes: Obtain the basic image generation model and the specific object generation model corresponding to the content category; Mount the specific object generation model onto the basic image generation model to obtain a fused image generation model, where the image generation model includes the basic image generation model, the specific object generation model, and the fused image generation model; According to the task type corresponding to the image generation instruction, use the fused image generation model to convert the final text into a target image corresponding to the target parameters, where the target parameters refer to the target size and / or target resolution corresponding to the task type.
9. The image generation method according to claim 8, wherein, The converting the final text into a target image corresponding to the target parameters by using the fused image generation model includes: Use the final text to guide the fused image generation model to convert a randomly generated noise image into a first intermediate image; Perform image editing on the first intermediate image to obtain a second intermediate image, where the image size of the second intermediate image matches the target size corresponding to the task type; Based on the target resolution corresponding to the task type, repair the image details of the second intermediate image to obtain the target image.
10. The image generation method according to claim 9, wherein, The image size includes the image length and the image width, the target size includes the target length and the target width, and the performing image editing on the first intermediate image to obtain a second intermediate image includes: When the image length of the first intermediate image is less than the target length, or the image width of the first intermediate image is less than the target width, crop the first intermediate image to obtain an image to be border-expanded, where the image length of the image to be border-expanded is less than or equal to the target length, and the image width of the border-expanded image is less than or equal to the target width; Use blank pixels to perform edge expansion on the image to be border-expanded to obtain an expanded image, where the image length of the expanded image is equal to the target length, and the image width of the expanded image is equal to the target width; Redraw the enlarged area of the enlarged image to obtain the second intermediate image, where the enlarged area is the area filled with blank pixels.
11. The image generation method according to claim 10, wherein, The redrawing of the enlarged area of the enlarged image to obtain the second intermediate image includes: Obtain the description information of the first intermediate image; Mask the enlarged image using a grayscale mask image; Add noise to the enlarged area according to the pixel values of each pixel point in the enlarged area after the masking process; Perform noise reduction on the enlarged area after the masking process based on the description information of the first intermediate image to obtain a third intermediate image; Smooth the transition area between the enlarged area and the initial area in the third intermediate image to obtain the second intermediate image, where the initial area is the area outside the transition area in the enlarged image.
12. The image generation method according to claim 9, wherein, The repairing of the image details of the second intermediate image to obtain the target image includes: Obtain the description information of the second intermediate image; Add noise to the main area of the second intermediate image; Perform noise reduction on the main area with added noise in the second intermediate image based on the description information of the second intermediate image to obtain the target image.
13. The image generation method according to claim 9, wherein, Before the repairing of the image details of the second intermediate image to obtain the target image, the method further includes: Adopt an interpolation technique to expand the second intermediate image along the length and width directions respectively; Perform image enhancement on the second intermediate image after image expansion through a deblurring algorithm.
14. An image generation device, wherein, The device includes: An acquisition module for acquiring an image generation instruction of a user, where the image generation instruction is used to indicate the generation of a target image; An identification module for identifying the image generation instruction, determining the content category of the target image based on the identification result, and acquiring the content rich information corresponding to the content category; An expansion module for expanding the prompt word corresponding to the image generation instruction using the content rich information to obtain a final text; A conversion module for converting the final text into a target image using the image generation model corresponding to the content category.
15. An image generation system, wherein, The system includes: A server for acquiring an image generation instruction of a user, where the image generation instruction is used to indicate the generation of a target image; An application terminal for identifying the image generation instruction, determining the content category of the target image based on the identification result, and acquiring the content rich information corresponding to the content category; The application terminal is further used for expanding the prompt word corresponding to the image generation instruction using the content rich information to obtain a final text; The application terminal is further used for converting the final text into a target image using the image generation model corresponding to the content category.
16. The image generation system according to claim 15, wherein, The server is further used for: Receiving a voice instruction of a user and sending the voice instruction to the application terminal; Receiving the final text converted from the voice instruction by the application terminal; Displaying the final text on the display screen of the vehicle; Send the final text to the application side and receive the target image generated by the application side based on the final text; Display the target image on the display screen of the vehicle.
17. An image generation device, wherein, The image generation device includes: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, the image generation method according to any one of claims 1-13 is implemented.
18. A computer storage medium, wherein, Computer program instructions are stored on the computer storage medium, and when the computer program instructions are executed by a processor, the image generation method according to any one of claims 1-13 is implemented.
19. A vehicle, wherein, The vehicle includes at least one of the following: The image generation device according to claim 14; the image generation system according to claim 15 and claim 16; the image generation device according to claim 17; the computer storage medium according to claim 18.