Image generation method and device

By introducing user input data to adjust model parameters in the image generation method, and optimizing the image generation model using knowledge base or real-time training set data, the problem of insufficient image generation accuracy is solved, and higher matching and applicability is achieved.

CN120296188APending Publication Date: 2025-07-11HUAWEI TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202410028193.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-08
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The accuracy of the existing image generation method is insufficient, and the degree of correspondence between the generated image and the user input data is not high.

Method used

By introducing user input data to adjust the parameters of the image generation model, and using the weight coefficients in the knowledge base or real-time training set data to optimize the model parameters, improve the matching degree between the model and the user input data.

Benefits of technology

The accuracy of the image generation method is improved, so that the degree of correspondence between the generated image and the user input data is increased, and the applicability is wider.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296188A_ABST
    Figure CN120296188A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an image generation method and device, relates to the technical field of image processing, and is used for improving the accuracy of the image generation method. The method comprises the steps of obtaining first data; determining first information according to the first data; adjusting parameters of the second model according to the first weight coefficient; and inputting the second data into the second model to obtain a target image. Wherein the first data is used for generating a target image, the first information comprises a first weight coefficient, the first weight coefficient is a weight coefficient of a first model, the first model is used for generating an adjustment parameter of a second model, the second model is used for generating the target image, and the second data comprises the first data. The image generation model parameters are adjusted through the user input data, the matching degree of the image generation model and the user input data can be improved, the corresponding degree of the generated image output by the image generation model and the user input data is increased, and therefore the accuracy of the image generation method is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of image processing technology, and in particular, to an image generation method and apparatus. Background Art

[0002] In the field of artificial intelligence, computer vision technology is applied in many fields such as image generation, semantic segmentation, and object detection. As an important part of computer vision technology, image generation technology has gradually gained popularity in the fields of science and technology and art.

[0003] An image generation method can convert user input data (such as text, images, etc.) into a generated image corresponding to the user input data. The generated image of the image generation method can be a real photo, a painting, a three-dimensional (3D) rendering, or a completely imaginary image.

[0004] The accuracy of the image generation method is used to represent the correspondence degree between the generated image and the user input data. Summary of the Invention

[0005] The embodiments of the present application provide an image generation method and apparatus for improving the accuracy of the image generation method. To achieve the above object, the embodiments of the present application adopt the following technical solutions:

[0006] In a first aspect, the embodiments of the present application provide an image generation method, the method includes: obtaining first data; determining first information according to the first data; adjusting parameters of a second model according to a first weight coefficient; inputting second data into the second model to obtain a target image. Wherein, the first data is used to generate the target image, the first information includes the first weight coefficient, the first weight coefficient is the weight coefficient of a first model, the first model is used to generate adjustment parameters of the second model, the second model is used to generate the target image, and the second data includes the first data.

[0007] It can be seen that the method provided by the embodiments of the present application, on the basis of the image generation technology, additionally introduces a step of adjusting the parameters of the image generation model (i.e., the second model) through user input data (i.e., the first data). By adjusting the parameters of the image generation model through user input data, the matching degree between the image generation model and the user input data can be improved, so that the correspondence degree between the generated image output by the image generation model and the user input data increases, thereby improving the accuracy of the image generation method.

[0008] In a possible implementation, the first feature vector can be determined based on the above-mentioned first data. When there is a target second feature vector in the first database, the first feature vector is input into the first database to obtain the above-mentioned first information. The first database includes Q second feature vectors and P first information corresponding to the Q second feature vectors. The target second feature vector is a second feature vector whose similarity to the first feature vector is greater than the similarity threshold. Both Q and P are positive integers. When the target second feature vector does not exist in the first database, the first feature vector is input into the second database to obtain training set data, and the training set data is input into the second model to obtain the above-mentioned first information. The second database includes W third feature vectors and E training set data corresponding to the W third feature vectors. Both W and E are positive integers.

[0009] It can be seen that for the method provided in the embodiments of the present application, on the one hand, when there is a weight coefficient matching the user input data in the knowledge base, the parameters of the image generation model can be directly adjusted through the existing weight coefficient in the knowledge base to improve the matching degree between the image generation model and the user input data. On the other hand, when there is no weight coefficient matching the user input data in the knowledge base, the existing data in the knowledge base can be used as a training set to generate a weight coefficient matching the user input data in real time, so as to improve the matching degree between the image generation model and the user input data, as well as the versatility of the image generation model, so that the image generation model can be applied to more image generation scenarios.

[0010] In a possible implementation, the similarity includes cosine similarity, Euclidean distance, Manhattan distance, Chebyshev distance, Hamming distance or other similarities.

[0011] In a possible implementation, the second database may include a real-time database, and the real-time database includes real-time training set data, which is training set data obtained through real-time search.

[0012] It can be seen that the method provided in the embodiments of the present application can obtain training set data matching the user data through real-time search, so as to improve the matching degree between the image generation model and the user input data, as well as the real-time performance of the image generation model, so that the image generation model can generate images with a higher matching degree to real-time hotspots.

[0013] In a possible implementation, when the target image generated by the above-mentioned first data is a portrait or an animal picture, the above-mentioned first weight information and the above-mentioned first data are input into the third model to obtain target intermediate data, and the target intermediate data is the intermediate data of the portrait or the intermediate data of the animal picture; the parameters of the second model are adjusted according to the target intermediate data.

[0014] It can be seen that for the method provided by the embodiments of the present application, when generating a human figure or an animal figure, the parameters of the image generation model can be adjusted according to the intermediate data (such as a human contour figure, an animal contour figure, or a pose figure) generated by the third model, so that the image generation model can more accurately output a human figure or an animal figure that matches the user input data.

[0015] In a possible implementation manner, the above second data can be input into the above second model to obtain M generated images, where M is a positive integer; N generated images among the above M generated images are determined as the above target images, where N is a positive integer.

[0016] For example, objective metrics can be calculated for the M generated images, and the N generated images with the highest objective metric scores among the M generated images are determined as the above target images.

[0017] In a possible implementation manner, the first information may further include an identifier, and the identifier is used to characterize the style of the above target image.

[0018] It can be understood that by adjusting the parameters of the image generation model through the identifier used to characterize the image style corresponding to the user data, the matching degree between the image generation model and the user input data can be further improved, so that the correspondence between the generated image output by the image generation model and the user input data is further increased, thereby further improving the accuracy of the image generation method.

[0019] In a possible implementation manner, the second data may further include an identifier, and the identifier is used to characterize the style of the above target image.

[0020] It can be understood that additionally inputting the identifier used to characterize the image style corresponding to the user data into the image generation model can further increase the correspondence between the generated image output by the image generation model and the user input data, thereby further improving the accuracy of the image generation method.

[0021] In a possible implementation manner, the first data may include at least one of text, picture, audio, or video.

[0022] It can be understood that on the one hand, compared with using only text as the input data for the image generation model to generate images, additionally introducing pictures, audio, and video as the input data for the image generation model to generate images can improve the applicability of the image generation model and make the image generation model applicable to more image generation scenarios. On the other hand, by adjusting the parameters of the image generation model through the text, pictures, audio, and video input by the user, the matching degree between the image generation model and the user input data can be improved, so that the correspondence between the generated image output by the image generation model and the user input data is increased, thereby improving the accuracy of the image generation method.

[0023] In a second aspect, an embodiment of the present application provides an image generation device, which includes a transceiver unit and a processing unit. The above transceiver unit is used to obtain first data, and the above first data is used to generate a target image. The above processing unit is used to determine first information according to the first data. The above first information includes a first weight coefficient, and the above first weight coefficient is the weight coefficient of a first model. The above first model is used to generate adjustment parameters of a second model, and the above second model is used to generate the above target image. The above processing unit is further used to adjust the parameters of the above second model according to the above first weight coefficient. The above transceiver unit is further used to input second data into the above second model to obtain the above target image, and the above second data includes the above first data.

[0024] In a possible implementation manner, the above processing unit is specifically used for: determining a first feature vector according to the above first data; when there is a target second feature vector in a first database, inputting the above first feature vector into the first database to obtain the above first information. The above first database includes Q second feature vectors and P pieces of first information corresponding to the Q second feature vectors. The above target second feature vector is a second feature vector whose similarity to the above first feature vector is greater than a similarity threshold. Both Q and P are positive integers; when there is no above target second feature vector in the first database, inputting the above first feature vector into a second database to obtain training set data and inputting the above training set data into the above second model to obtain the above first information. The above second database includes W third feature vectors and E pieces of training set data corresponding to the W third feature vectors. Both W and E are positive integers.

[0025] In a possible implementation manner, the above second database includes a real-time database, and the above real-time database includes real-time training set data, and the above real-time training set data is training set data obtained through real-time search.

[0026] In a possible implementation manner, the above processing unit is specifically used for: when the target image generated by the above first data is a portrait or an animal picture, inputting the above first weight information and the above first data into a third model to obtain target intermediate data, and the above target intermediate data is the intermediate data of the above portrait or the intermediate data of the above animal picture; adjusting the parameters of the above second model according to the above target intermediate data.

[0027] In a possible implementation manner, the above transceiver unit is specifically used for: inputting the above second data into the above second model to obtain M generated images, where M is a positive integer; determining N generated images among the above M generated images as the above target image, where N is a positive integer.

[0028] In a possible implementation, the above first information further includes an identifier, and the identifier is used to characterize the style of the above target image.

[0029] In a possible implementation, the above second data further includes an identifier, and the identifier is used to characterize the style of the above target image.

[0030] In a possible implementation, the above first data includes at least one of text, picture, audio, or video.

[0031] In a third aspect, an embodiment of the present application further provides an image generation device, which includes: at least one processor, and when the at least one processor executes program code or instructions, the method described in the above first aspect or any possible implementation manner thereof is implemented.

[0032] Optionally, the image generation device may further include at least one memory, and the at least one memory is used to store the program code or instructions.

[0033] In a fourth aspect, an embodiment of the present application further provides a chip, including: an input interface, an output interface, and at least one processor. Optionally, the chip further includes a memory. The at least one processor is used to execute the code in the memory, and when the at least one processor executes the code, the chip implements the method described in the above first aspect or any possible implementation manner thereof.

[0034] Optionally, the above chip may further be an integrated circuit.

[0035] In a fifth aspect, an embodiment of the present application further provides a computer-readable storage medium for storing a computer program, and the computer program includes the method for implementing the above first aspect or any possible implementation manner thereof.

[0036] In a sixth aspect, an embodiment of the present application further provides a computer program product containing instructions, and when it runs on a computer, it causes the computer to implement the method described in the above first aspect or any possible implementation manner thereof. Description of the Drawings

[0037] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0038] Figure 1 It is a schematic structural diagram of an image processing system provided by an embodiment of the present application;

[0039] Figure 2 A structural schematic diagram of a database provided by an embodiment of the present application;

[0040] Figure 3 A structural schematic diagram of a first database provided by an embodiment of the present application;

[0041] Figure 4 A structural schematic diagram of a second database provided by an embodiment of the present application;

[0042] Figure 5 A flowchart of an image generation method provided by an embodiment of the present application;

[0043] Figure 6 A schematic diagram of a data filtering process provided by an embodiment of the present application;

[0044] Figure 7 A schematic diagram of a data retrieval process provided by an embodiment of the present application;

[0045] Figure 8 A schematic diagram of a model training process provided by an embodiment of the present application;

[0046] Figure 9 A schematic diagram of an image generation process provided by an embodiment of the present application;

[0047] Figure 10 A structural schematic diagram of an image generation device provided by an embodiment of the present application;

[0048] Figure 11 A structural schematic diagram of a chip provided by an embodiment of the present application;

[0049] Figure 12 A structural schematic diagram of an electronic device provided by an embodiment of the present application;

[0050] Figure 13 A structural schematic diagram of another image generation device provided by an embodiment of the present application. Detailed implementation manners

[0051] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the embodiments of the present application.

[0052] As used herein, the term "and / or" is merely a description of the relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone.

[0053] In the description of the embodiments of the present application, terms such as "first" and "second" in the specification and drawings are used to distinguish different objects or different treatments of the same object, rather than to describe a specific order of the objects.

[0054] In addition, the terms "including" and "having" and any variations thereof mentioned in the description of the embodiments of the present application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes other steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.

[0055] It should be noted that in the description of the embodiments of the present application, words such as "exemplarily" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design solution described as "exemplarily" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplarily" or "for example" is intended to present relevant concepts in a specific manner.

[0056] Before introducing the technical solution of the present application, some terms related to the present application are explained. The following related explanations can be arbitrarily combined with the technical solution of the embodiments of the present application as optional solutions, and they all fall within the protection scope of the embodiments of the present application. The embodiments of the present application include at least some of the following contents.

[0057] Latent Diffusion Model (StableDiffusion, SD): An open-source text-to-image generation model based on deep learning, mainly used to generate detailed images according to text descriptions.

[0058] Large Language Model (LLM): An extremely large deep learning model pre-trained on a large amount of data, which can perform open-domain tasks such as text summarization and language translation, and has certain logical thinking and reasoning abilities.

[0059] Low Rank Adaptation (LORA): A technique that approximates the high-dimensional structure of a large model with a low-dimensional structure to reduce its complexity, enabling efficient fine-tuning of the model. This technique can adapt the large model to specific tasks or domains faster and more effectively.

[0060] Contrastive Language-Image Pre-training (CLIP): A pre-trained neural network model for matching images and texts. This task is widely used in the multi-modal field for text-image retrieval. The model is pre-trained with a large amount of paired Internet data. Its structure consists of two branches, a text encoder and an image encoder, which are used to learn the matching similarity between images and texts. Thus, it can retrieve images according to texts or retrieve texts according to images. In addition, the training strategy of image-text contrastive learning can also be applied to retrieval tasks in other modalities (such as audio, video, etc.).

[0061] DreamBooth (DB): A method for generating personalized text-to-image. It can fine-tune the diffusion model through a custom theme and can achieve good results with only a small amount of training data.

[0062] Bootstrapping Language-Image Pre-training 2 (BLIP2): A vision-language model for multi-modal understanding of images and texts. The model consists of a visual encoder, a text encoder, a query transformer (Q-Former), and an LLM, which is used to align the two modalities of vision and language. This method uses pre-trained visual encoders and LLMs to enable tasks such as image caption generation, visual question answering, and fine-grained image retrieval. At the same time, replacing the LLM with a model with a larger number of parameters and with instruction-following and chain-of-thought capabilities can achieve the retrieval of local information of images and texts.

[0063] Prompt: The descriptive text input to the Stable Diffusion model.

[0064] Text-to-Image: Generating pictures that conform to the semantics according to the text description input by the user, such as the core elements and semantic relationships contained in the text.

[0065] Image Synthesis: Fusing the current image according to the style of a set of images provided by the user.

[0066] Image Inpainting and Super-Resolution: Adjusting the details and resolution of the image according to the original image provided by the user.

[0067] Image-to-Image: Editing a partial area of a picture according to the image provided by the user and controlling image generation using external conditions.

[0068] In the field of artificial intelligence, computer vision technology is applied in many fields such as image generation, semantic segmentation, and object detection. As an important part of computer vision technology, image generation technology is also gradually popularized in the fields of science and technology and art.

[0069] An image generation method can convert user input data (such as text, images, etc.) into a generated image corresponding to the user input data. The generated image of the image generation method can be a real photo, a painting, a 3D rendering, or a completely imaginary image.

[0070] The accuracy of the image generation method is used to represent the corresponding degree between the generated image and the user input data.

[0071] For this reason, the embodiments of the present application provide an image generation method for improving the accuracy of the image generation method.

[0072] The technical solutions provided by the embodiments of the present application can be applied to an image processing system. The application scenarios of the technical solutions provided by the embodiments of the present application can include various types, such as text-to-image generation, image synthesis (image style transfer), image restoration and super-resolution, image-to-image generation, etc.

[0073] Figure 1 Shows a possible and non-limiting schematic diagram of the above image processing system. As Figure 1 shown, the image processing system 100 includes an electronic device 110 and a database 120.

[0074] The electronic device 110 can be an electronic device such as a mobile phone, a desktop computer, a tablet computer, a laptop computer, a vehicle-mounted terminal, a server, an intelligent robot, an intelligent TV, a multimedia playback device, etc., or other electronic devices, and the embodiments of the present application do not limit this.

[0075] The database 120, the database 120 can be a knowledge base. The knowledge base is a special database for knowledge management. The knowledge base can store information content such as text, images, videos, and audios, and retrieve data corresponding to the request from the data stored in the knowledge base according to the request.

[0076] It can be understood that Figure 1 the structure of the image processing system 100 shown does not constitute a specific limitation on the image processing system 100. In other embodiments of the present application, the image processing system 100 may include more or fewer components than shown, or combine certain components, or split certain components, or different component arrangements. The illustrated components can be implemented by hardware, software, or a combination of software and hardware.

[0077] Figure 2 Shows a possible and non-limiting schematic diagram of the above database. As Figure 2 shown, the database 120 may include a first database 121 and a second database 122.

[0078] The first database 121 is used to store the weight coefficients of the model.

[0079] The second database 122 is used to store the model training data set (such as data like pictures, texts, videos or audios, etc.).

[0080] It can be understood that Figure 2 The structure of the illustrated database 120 does not constitute a specific limitation on the database 120. In some other embodiments of the present application, the database 120 may include more or fewer components than those shown, or combine certain components, or split certain components, or have different component arrangements. The illustrated components can be implemented by hardware, software, or a combination of software and hardware.

[0081] Figure 3 A schematic diagram of a possible and non - restrictive above - mentioned first database is shown. As Figure 3 The illustrated first database 121 may include a LORA knowledge base and a DB knowledge base.

[0082] The LORA knowledge base is used to provide LORA weight information matching the style or specific concept elements.

[0083] The DB knowledge base is used to provide identifiers and LORA weight information matching the concept.

[0084] It can be understood that Figure 3 The structure of the illustrated first database 121 does not constitute a specific limitation on the first database 121. In some other embodiments of the present application, the first database 121 may include more or fewer components than those shown, or combine certain components, or split certain components, or have different component arrangements. The illustrated components can be implemented by hardware, software, or a combination of software and hardware.

[0085] Figure 4 A schematic diagram of a possible and non - restrictive above - mentioned second database is shown. As Figure 4 The illustrated second database 122 may include a general knowledge base, a domain knowledge base, and a real - time knowledge base.

[0086] The general knowledge base is used to provide data such as texts, drawings, videos, and audios in the open domain.

[0087] The domain knowledge base is used to provide data such as texts, drawings, videos, and audios in the vertical domain (such as portraits, animals, etc.).

[0088] The real - time knowledge base is used to provide data such as texts, drawings, videos, and audios of real - time hot concepts, and data such as texts, drawings, videos, and audios that can be queried online.

[0089] It can be understood that Figure 4The structure of the second database 122 shown does not specifically limit the second database 122. In some other embodiments of the present application, the second database 122 may include more or fewer components than those shown, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented by hardware, software, or a combination of software and hardware.

[0090] It can be understood that in the present application, an electronic device and a database are used as examples of the execution subjects for interactive illustration, but the present application does not limit the execution subjects for interactive illustration. For example, the electronic device in the method of the present application may also be a chip, a chip system, or a processor applied to the electronic device, and may also be a logical node, a logical module, or software that can implement all or part of the functions of the electronic device; the knowledge base in the method of the present application may also be a chip, a chip system, or a processor applied to the knowledge base, and may also be a logical node, a logical module, or software that can implement all or part of the functions of the knowledge base.

[0091] Figure 5 The following shows an image generation method provided by an embodiment of the present application, as Figure 5 shown, the method includes:

[0092] S501. Obtain first data.

[0093] Among them, the above-mentioned first data is used to generate a target image.

[0094] In a possible implementation manner, the above-mentioned first data may include at least one of text, pictures, audio, or video, that is, the first data may be composed of text, pictures, audio, or video.

[0095] Among them, the text may include category words. The category words are used to indicate the type of the target image. For example, the category word - vehicle is used to indicate that the target image is a vehicle image.

[0096] Among them, the text may include positive prompt words and / or negative prompt words.

[0097] Exemplarily, when the user wants to generate an image of a zebra, the first data obtained by the electronic device may be the text "generate an image of a zebra".

[0098] Again exemplarily, when the user wants to generate an image of a train, the first data obtained by the electronic device may be a picture of a train.

[0099] Of course, the first data obtained by the electronic device may also be composed of multiple types of data. For example, it may be data combined by multiple types of data such as picture + text, picture + audio, audio + text, etc.

[0100] In a possible implementation, the first data can be filtered to filter out potential sensitive information in the first data.

[0101] Exemplarily, as Figure 6 shown, the content in the first data can be divided into text, pictures, audio, and video.

[0102] As Figure 6 shown, for the text in the first data, a word list (such as a stop words list) can be used for filtering. In the case where there are sensitive words in the text of the first data, the first data or the text content of the first data can be processed. For example, the first data or the text content of the first data can be invalidated, that is, no image is generated based on the first data or no image is generated based on the text content of the first data.

[0103] As Figure 6 shown, for the pictures in the first data, a convolutional neural network (such as a You Only Look Once (YOLO) network) can be used with a content detection model (such as a Not Safe For Work (NSFW) detection model) for filtering. In the case where the detection value of the picture is greater than the detection threshold (i.e., it is determined that there is sensitive content in the picture), the first data or the picture content of the first data can be processed. For example, the first data or the picture content of the first data can be invalidated, that is, no image is generated based on the first data or no image is generated based on the picture content of the first data.

[0104] As Figure 6 shown, for the videos in the first data, a convolutional neural network can be used with a content detection model for filtering. In the case where the detection value of the video is greater than the detection threshold (i.e., it is determined that there is sensitive content in the video), the first data or the video content of the first data can be processed. For example, the first data or the video content of the first data can be invalidated, that is, no video is generated based on the first data or no image is generated based on the video content of the first data.

[0105] As Figure 6 shown, for the audio in the first data, an acoustic model can be used for filtering. In the case where the detection value of the audio is greater than the detection threshold (i.e., it is determined that there is sensitive content in the audio), the first data or the audio content of the first data can be processed. For example, the first data or the audio content of the first data can be invalidated, that is, no audio is generated based on the first data or no image is generated based on the audio content of the first data.

[0106] In a possible implementation, the first data can be normalized to make the first data conform to the input of the knowledge base.

[0107] AsFigure 6 As shown in Figure 6 , the text in the first data can be text-normalized to obtain normalized text.

[0108] For example, special characters in the text content of the first data can be filtered to obtain normalized text.

[0109] Such as Figure 6 As shown in Figure 6 , the images in the first data can be image-normalized to obtain normalized images.

[0110] For example, the images in the first data can be uniformly converted into normalized images in the red green blue (RGB) format, resize format, or sample format.

[0111] Such as Figure 6 As shown in Figure 6 , the videos in the first data can be image-normalized to obtain normalized videos.

[0112] For example, the video content in the first data can be uniformly converted into normalized videos in the RGB format, resize format, or sample format.

[0113] Such as Figure 6 As shown in Figure 6 , the audio in the first data can be audio-normalized.

[0114] For example, the audio content in the first data can be uniformly converted into normalized audio in the sample format.

[0115] S502. Determine the first information according to the first data.

[0116] Wherein, the above first information includes a first weight coefficient, and the above first weight coefficient is the weight coefficient of the first model, and the first model is used to generate adjustment parameters of the second model, and the second model is used to generate the above target image.

[0117] Exemplarily, the first model can be a Lora model, and the second model can be an SD model.

[0118] Exemplarily, the first information may further include an identifier, and the above identifier is used to characterize the style of the above target image.

[0119] For example, if the identifier is [token][cartoon], then the identifier characterizes that the style of the target image is a cartoon style.

[0120] Exemplarily, a task query instruction can be generated according to the first data through an Agent index, and the first information can be obtained by querying a database according to the task query instruction.

[0121] In a possible implementation, the first feature vector can be determined based on the above-mentioned first data. When the target second feature vector exists in the first database, the above-mentioned first feature vector is input into the first database to obtain the above-mentioned first information; when the target second feature vector does not exist in the first database, the above-mentioned first feature vector is input into the second database to obtain training set data, and the above-mentioned training set data is input into the above-mentioned second model to obtain the above-mentioned first information.

[0122] As shown in Table 1, the above-mentioned first database may include Q second feature vectors and P first information corresponding to the Q second feature vectors. The target second feature vector is a second feature vector whose similarity to the above-mentioned first feature vector is greater than the similarity threshold, and both Q and P are positive integers. Each second feature vector may correspond to one or more first information, and each first information may correspond to one or more second feature vectors.

[0123] Table 1

[0124] Second eigenvector First information Second eigenvector 1 First information 1 Second eigenvector 2 First information 2 Second eigenvector 3 First information 3 Second eigenvector 4 First information 4 …… …… …… …… …… …… Second eigenvector Q First information P

[0125] It can be understood that when the target second feature vector exists, the first information obtained by inputting the above-mentioned first feature vector into the first database includes the first information corresponding to the target second feature vector.

[0126] For example, the similarities between the second feature vector 2 and the second feature vector 3 and the first feature vector are greater than the similarity, that is, the second feature vector 2 and the second feature vector 3 are the target second feature vectors. The first information obtained by inputting the first feature vector into the first database includes the first information 2 corresponding to the second feature vector 2 and the first information 3 corresponding to the second feature vector 3.

[0127] As shown in Table 2, the above-mentioned second database includes W third feature vectors and E training set data corresponding to the W third feature vectors, and both W and E are positive integers. Each third feature vector may correspond to one or more training data sets, and each training data set may correspond to one or more third feature vectors.

[0128] Table 2

[0129] Third eigenvector Training data set Third eigenvector 1 Training data set 1 Third eigenvector 2 Training data set 2 Third eigenvector 3 Training data set 3 Third eigenvector 4 Training data set 4 …… …… …… …… …… …… Third eigenvector W Training data set E

[0130] In a possible implementation, the similarity includes cosine similarity, Euclidean distance, Manhattan distance, Chebyshev distance, Hamming distance or other similarities.

[0131] It can be understood that in the absence of the target second feature vector, the training set obtained by inputting the above first feature vector into the first database includes the training data set corresponding to the target third feature vector. The above target third feature vector is a third feature vector whose similarity to the above first feature vector is greater than the similarity threshold.

[0132] For example, the similarity between the third feature vector 1 and the first feature vector is greater than the similarity, that is, the third feature vector 1 is the target third feature vector. The training data set obtained by inputting the first feature vector into the second database includes the training data set 1 corresponding to the third feature vector 1.

[0133] Exemplarily, as Figure 7 shown, in the case where the target second feature vector exists in the first database, inputting the first feature vector obtained from the first data into the first database can obtain the first information.

[0134] Also exemplarily, as Figure 7 shown, in the case where the above target second feature vector does not exist in the first database, inputting the first feature vector obtained from the first data into the second database can obtain the training set data. Inputting the obtained training set data into the second model can obtain the first information.

[0135] Optionally, the above second database includes a real-time database (real-time knowledge base), and the real-time database includes real-time training set data, and the real-time training set data is training set data obtained through real-time search.

[0136] For example, after the first feature vector is input into the real-time database, the real-time database can perform real-time search according to the first feature vector to obtain the training set data, and output the training set data obtained by real-time search according to the first feature vector.

[0137] In a possible implementation manner, the first database can be trained according to the first information obtained from the training set data. That is, the first feature vector and the first information obtained from the training set data can be sent to the first database. In this way, the first database can directly output the corresponding first information the next time it receives the first feature vector, without having to obtain the training data set from the second database again.

[0138] In a possible implementation manner, the first information can also be [displayed]. For example, the first information can be displayed through the display screen of the electronic device.

[0139] In a possible implementation manner, the first feature vector may include a first coarse-grained feature vector and / or a first fine-grained feature vector.

[0140] For the text content in the first data, the first fine-grained feature vector represents the global text, and the first fine-grained feature vector represents the key text segments.

[0141] For the video, image, and audio in the first data, the first fine-grained feature vector represents the sampled global content, and the first fine-grained feature vector represents the sampled high-frequency and regional main body information.

[0142] In a possible implementation, the first data can be input into a vector model to obtain a first feature vector.

[0143] In a possible implementation, the above vector model can include a CLIP model, a BLIP2 model, or an LLM model.

[0144] Exemplarily, the first data can be input into the CLIP model to obtain a first coarse-grained feature vector.

[0145] Another exemplarily, the first data can be input into the BLIP2 model to obtain a first fine-grained feature vector.

[0146] S503. Adjust the parameters of the second model according to the first weight coefficient.

[0147] In a possible implementation, when the target image generated by the above first data is a portrait or an animal image, the above first weight information and the above first data can be input into a third model to obtain target intermediate data; adjust the parameters of the second model according to the above target intermediate data.

[0148] Wherein, the above target intermediate data is the intermediate data of the above portrait or the intermediate data of the above animal image. For example, the target intermediate data can be a portrait contour map, an animal contour map, or a pose map.

[0149] Exemplarily, the above third model can be a control network model.

[0150] Exemplarily, as Figure 8 shown, when the target image generated by the above first data is a portrait or an animal image, the above first weight information and the above first data are input into a third model to obtain target intermediate data, and then the parameters of the second model are adjusted according to the above target intermediate data to obtain a trained second model.

[0151] Another exemplarily, as Figure 8 shown, when the target image generated by the above first data is not a portrait and not an animal image, the parameters of the second model can be directly adjusted according to the first weight coefficient to obtain a trained second model.

[0152] It should be noted that the first model can be understood as a fine-tuning model. The fine-tuning model can be connected to other models, and the parameters of other models can be frozen. Only by fine-tuning the weight coefficients of the first model can other models be trained. Therefore, the parameters of the second model can be adjusted through the first weight coefficient.

[0153] For example, take the first model as the LoRA model and the second model as the SD model. The LoRA model can be a small model trained under the SD large model for fine-tuning the large model. LoRA can adjust the characters, styles, etc.

[0154] The fine-tuning model can be regarded as a plugin for a large model (such as the SD model). Without modifying the SD model, it can use a small amount of data to train a painting style / intellectual property (IP) / character to meet customized requirements. The training resources required are smaller than those for training the SD model. To reduce training costs, the LoRA model only trains low-rank matrices. When used, the parameters of the LoRA model are injected into the SD model to change the generation style of the SD model or add new characters / IPs to the SD model. The whole process is a linear relationship. It can be considered that after the original SD model is superimposed with the LoRA model, a model with a brand-new effect is obtained.

[0155] S504. Input the second data into the second model to obtain a target image.

[0156] Among them, the above-mentioned second data includes the above-mentioned first data.

[0157] In a possible implementation, the above-mentioned second data may further include an identifier. The identifier is used to characterize the style of the above-mentioned target image. For example, if the identifier [token][Gothic] represents that the style of the target image is Gothic style.

[0158] In a possible implementation, the above-mentioned second data can be input into the above-mentioned second model to obtain M generated images, where M is a positive integer; N generated images among the above-mentioned M generated images are determined as the above-mentioned target image, and N is a positive integer.

[0159] Exemplarily, as Figure 9 shown, the above-mentioned second data can be input into the trained second model to obtain M generated images, and then the M generated images are input into an image evaluation model to obtain evaluation scores of the M images. Then, the N generated images with the highest input evaluation scores are used as the target images.

[0160] For example, the above second data can be encoded by an Encoder to obtain encoded data. The encoded data is input into the trained second model to obtain M generated images, and then the M generated images are input into an image evaluation model to obtain evaluation scores of the M images. Subsequently, the 4 generated images with the highest evaluation scores are input as target images.

[0161] Exemplarily, the above image evaluation model may include a semantic consistency model (such as a CLIP model) and / or an aesthetics model (such as a Human Preference Score (HPS) model).

[0162] In a possible implementation manner, the target images can also be displayed. For example, the target images can be displayed through the display screen of an electronic device.

[0163] It can be seen that the method provided in the embodiments of the present application, on the basis of the image generation technology, additionally introduces the step of adjusting the parameters of the image generation model (i.e., the second model) through user input data (i.e., the first data). By adjusting the parameters of the image generation model through user input data, the matching degree between the image generation model and the user input data can be improved, and the corresponding degree between the generated images output by the image generation model and the user input data can be increased, thereby improving the accuracy of the image generation method.

[0164] Next, in combination with Figure 10 An image generation device for executing the above image generation method will be introduced.

[0165] It can be understood that in order to implement the above functions, the image generation device includes corresponding hardware and / or software modules for executing each function. Combining the algorithm steps of each example described in the embodiments disclosed in this article, the embodiments of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application in combination with the embodiments, but such implementation should not be considered to exceed the scope of the embodiments of the present application.

[0166] The embodiments of the present application can divide the function modules of the image generation device according to the above method examples. For example, each function module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware. It should be noted that the division of modules in this embodiment is illustrative, only a logical function division, and there may be other division methods in actual implementation.

[0167] In the case of dividing each function module corresponding to each function, Figure 10FIG. 0 shows a possible schematic composition of the image generation device involved in the above embodiments. The device can be an electronic device, or a module applied to an electronic device (such as a processor, a chip, or a chip system, etc.), or a logical node, a logical module, or software that can implement all or part of the functions of an electronic device. As Figure 10 shown, the image generation device 1000 may include: a transceiver unit 1001 and a processing unit 1002.

[0168] The transceiver unit 1001 is configured to obtain first data, and the first data is used to generate a target image.

[0169] The processing unit 1002 is configured to determine first information according to the first data. The first information includes a first weight coefficient, and the first weight coefficient is a weight coefficient of a first model. The first model is used to generate adjustment parameters of a second model, and the second model is used to generate the target image.

[0170] The processing unit 1002 is further configured to adjust parameters of the second model according to the first weight coefficient.

[0171] The transceiver unit 1001 is further configured to input second data into the second model to obtain the target image, and the second data includes the first data.

[0172] In a possible implementation manner, the processing unit 1001 is specifically configured to: determine a first feature vector according to the first data; when a target second feature vector exists in a first database, input the first feature vector into the first database to obtain the first information. The first database includes Q second feature vectors and P first information corresponding to the Q second feature vectors. The target second feature vector is a second feature vector whose similarity to the first feature vector is greater than a similarity threshold, and both Q and P are positive integers; when the target second feature vector does not exist in the first database, input the first feature vector into a second database to obtain training set data and input the training set data into the second model to obtain the first information. The second database includes W third feature vectors and E training set data corresponding to the W third feature vectors, and both W and E are positive integers.

[0173] In a possible implementation manner, the second database includes a real-time database, and the real-time database includes real-time training set data, and the real-time training set data is training set data obtained through real-time search.

[0174] In a possible implementation manner, the above-mentioned processing unit 1002 is specifically configured to: when the target image generated by the above-mentioned first data is a portrait or an animal picture, input the above-mentioned first weight information and the above-mentioned first data into a third model to obtain target intermediate data, where the target intermediate data is the intermediate data of the portrait or the intermediate data of the animal picture; adjust the parameters of the above-mentioned second model according to the target intermediate data.

[0175] In a possible implementation manner, the above-mentioned transceiver unit 1001 is specifically configured to: input the above-mentioned second data into the above-mentioned second model to obtain M generated images, where M is a positive integer; determine N generated images among the above-mentioned M generated images as the above-mentioned target images, where N is a positive integer.

[0176] In a possible implementation manner, the above-mentioned first information further includes an identifier, and the identifier is used to characterize the style of the above-mentioned target image.

[0177] In a possible implementation manner, the above-mentioned second data further includes an identifier, and the identifier is used to characterize the style of the above-mentioned target image.

[0178] In a possible implementation manner, the above-mentioned first data includes at least one of text, picture, audio or video.

[0179] The embodiment of the present application also provides a chip. Figure 11 The structural schematic diagram of a chip 1100 is shown. The chip 1100 includes one or more processors 1101 and an interface circuit 1102. Optionally, the above-mentioned chip 1100 may further include a bus 1103.

[0180] The processor 1101 may be an integrated circuit chip with signal processing capabilities. During the implementation process, each step of the above-mentioned image generation method may be completed by the integrated logic circuit in the hardware of the processor 1101 or by instructions in software form.

[0181] Optionally, the above-mentioned processor 1101 may be a general-purpose processor, a digital signal processing (DSP) processor, an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods and steps disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0182] The interface circuit 1102 can be used for sending or receiving data, instructions, or information. The processor 1101 can utilize the data, instructions, or other information received by the interface circuit 1102 for processing, and can send the processed information through the interface circuit 1102.

[0183] Optionally, the chip further includes a memory, which can include a read-only memory and a random access memory, and provides operation instructions and data to the processor. A part of the memory can also include a non-volatile random access memory (NVRAM).

[0184] Optionally, the memory stores executable software modules or data structures. The processor can execute corresponding operations by calling the operation instructions stored in the memory (the operation instructions can be stored in the operating system).

[0185] Optionally, the chip can be used in the image generation device involved in the embodiments of the present application. Optionally, the interface circuit 1102 can be used to output the execution result of the processor 1101. For the image generation method provided by one or more embodiments of the present application, reference can be made to the foregoing various embodiments, which will not be elaborated here.

[0186] It should be noted that the functions corresponding to the processor 1101 and the interface circuit 1102 can be implemented through hardware design, can also be implemented through software design, or can be implemented through a combination of software and hardware, and are not limited here.

[0187] Figure 12 This is a schematic structural diagram of an electronic device provided by an embodiment of the present application. The electronic device 1200 can be a processor, a chip, or a functional module in a processor. As Figure 12 shown, the electronic device 1200 includes a processor 1201, a transceiver 1202, and a communication line 1203.

[0188] Among them, the processor 1201 is used to execute any step in the image generation method provided by the embodiment of the present application, and during the process of executing any step in the image generation method provided by the embodiment of the present application, it can optionally call the transceiver 1202 and the communication line 1203 to complete the corresponding operations.

[0189] Further, the electronic device 1200 can also include a memory 1204. Among them, the processor 1201, the memory 1204, and the transceiver 1202 can be connected through the communication line 1203.

[0190] Among them, the processor 1201 is a processor, a general - purpose processor, a network processor (NP), a digital signal processor (DSP), a microprocessor, a micro - controller, a programmable logic device (PLD), or any combination thereof. The processor 1201 can also be other devices with processing functions, such as circuits, devices, or software modules, without limitation.

[0191] The transceiver 1202 is used to communicate with other devices or other communication networks. The other communication networks can be Ethernet, radio access network (RAN), wireless local area networks (WLAN), etc. The transceiver 1202 can be a module, a circuit, a transceiver, or any device capable of realizing communication.

[0192] The transceiver 1202 is mainly used for sending and receiving commands, information, etc. It can include a transmitter and a receiver, which are respectively used for sending and receiving commands, information, etc. Operations other than sending and receiving commands, information, etc. are implemented by the processor.

[0193] The communication line 1203 is used to transmit information between the components included in the electronic device 1200.

[0194] In one design, the processor can be regarded as a logic circuit, and the transceiver can be regarded as an interface circuit.

[0195] The memory 1204 is used to store instructions. Among them, the instructions can be computer programs.

[0196] Among them, the memory 1204 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DRRAM). The memory 1204 can also be a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, etc. It should be noted that the memory of the systems and methods described herein is intended to include but not be limited to these and any other suitable types of memory.

[0197] It should be noted that the memory 1204 can exist independently of the processor 1201 or can be integrated with the processor 1201. The memory 1204 can be used to store instructions, program codes, or some data, etc. The memory 1204 can be located inside the electronic device 1200 or outside the electronic device 1200, without limitation. The processor 1201 is used to execute the instructions stored in the memory 1204 to implement the method provided in the above embodiments of the present application.

[0198] In one example, the processor 1201 can include one or more processors, such as Figure 12 processor 0 and processor 1 in

[0199] As an alternative implementation, the electronic device 1200 includes multiple processors. For example, in addition to Figure 12 the processor 1201 therein, it may further include a processor 1209.

[0200] As an alternative implementation, the electronic device 1200 further includes an output device 1205 and an input device 1206. Exemplarily, the input device 1206 is a device such as a keyboard, a mouse, a microphone, or a joystick, and the output device 1205 is a device such as a display screen or a speaker.

[0201] It should be noted that the electronic device 1200 may be a chip system or a device with a Figure 12 similar structure therein. Among them, the chip system may be composed of chips or may include chips and other discrete devices. Actions, terms, etc. involved among the embodiments of the present application may refer to each other without limitation. The message names or parameter names in the messages exchanged between the devices in the embodiments of the present application are only examples, and other names may also be adopted in specific implementations without limitation. In addition, Figure 12 the component structure shown therein does not constitute a limitation on the electronic device 1200. Except Figure 12 for the components shown, the electronic device 1200 may include more or fewer components than Figure 12 shown, or combine certain components, or have different component arrangements.

[0202] The processors and transceivers described in the present application may be implemented on an integrated circuit (IC), an analog IC, a radio frequency integrated circuit, a mixed-signal IC, an application specific integrated circuit (ASIC), a printed circuit board (PCB), an electronic device, etc. The processors and transceivers may also be manufactured using various IC process technologies, such as complementary metal oxide semiconductor (CMOS), N-type metal oxide semiconductor (NMOS), P-type metal oxide semiconductor (PMOS), bipolar junction transistor (BJT), BiCMOS, silicon germanium (SiGe), gallium arsenide (GaAs), etc.

[0203] Figure 13The figure is a schematic structural diagram of an image generation device provided by an embodiment of the present application. The image generation device can be applied to the scenarios shown in the above method embodiments. For ease of explanation, Figure 13 only the main components of the image generation device are shown, including a processor 1301, a memory 1302, a control circuit 1303, and an input / output device 1304. The processor 1301 is mainly used for processing communication protocols and communication data, executing software programs, and processing data of software programs. The memory 1302 is mainly used for storing software programs and data. The control circuit 1303 is mainly used for power supply and transmission of various electrical signals. The input / output device 1304 is mainly used for receiving data input by users and outputting data to users.

[0204] When the image generation device is the processor 1301, the control circuit 1303 can be a main board, the memory 1302 includes media with storage functions such as a hard disk, RAM, and ROM, the processor 1301 can include a baseband processor 1301 and a central processor. The baseband processor is mainly used for processing communication protocols and communication data, and the central processor is mainly used for controlling the entire image generation device, executing software programs, and processing data of software programs. The input / output device 1304 includes a display screen, a keyboard, a mouse, etc.; the control circuit 1303 can further include or be connected to a transceiver circuit or a transceiver, for example: a network cable interface, etc., for sending or receiving data or signals, for example, for data transmission and communication with other devices. Further, an antenna can also be included for wireless signal transceiver for data / signal transmission with other devices.

[0205] An embodiment of the present application further provides an image generation device, which includes: at least one processor, when the at least one processor executes program code or instructions, the above-mentioned related method steps are implemented to implement the image generation method in the above embodiment.

[0206] Optionally, the device can further include at least one memory, and the at least one memory is used for storing the program code or instructions.

[0207] An embodiment of the present application further provides a computer storage medium, in which computer instructions are stored. When the computer instructions run on the image generation device, the image generation device is enabled to execute the above-mentioned related method steps to implement the image generation method in the above embodiment.

[0208] An embodiment of the present application further provides a computer program product. When the computer program product runs on a computer, the computer is enabled to execute the above-mentioned related steps to implement the image generation method in the above embodiment.

[0209] An embodiment of the present application further provides an image generation device, which may specifically be a chip, an integrated circuit, a component or a module. Specifically, the device may include a processor connected to a memory for storing instructions, or the device includes at least one processor for obtaining instructions from an external memory. When the device runs, the processor may execute the instructions to cause the chip to execute the image generation method in each of the above method embodiments.

[0210] It should be understood that in various embodiments of the present application, the sequence numbers of the above processes do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0211] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0212] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0213] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the above division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling, direct coupling or communication connection may be through some interfaces, and the indirect coupling or communication connection of devices or units may be in an electrical, mechanical or other form.

[0214] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units. They may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0215] In addition, in each embodiment of the present application, each functional unit may be integrated into one processing unit, may exist physically alone for each unit, or two or more units may be integrated into one unit.

[0216] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the above methods in each embodiment of the present application. The foregoing storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0217] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. An image generation method, characterized in that, Including: Obtain first data, where the first data is used to generate a target image; Determine first information according to the first data, the first information includes a first weight coefficient, the first weight coefficient is the weight coefficient of a first model, the first model is used to generate adjustment parameters of a second model, and the second model is used to generate the target image; Adjust the parameters of the second model according to the first weight coefficient; Input second data into the second model to obtain the target image, where the second data includes the first data.

2. The method according to claim 1, wherein The determining first information according to the first data includes: Determine a first feature vector according to the first data; When there is a target second feature vector in a first database, input the first feature vector into the first database to obtain the first information, the first database includes Q second feature vectors and P first information corresponding to the Q second feature vectors, the target second feature vector is a second feature vector whose similarity to the first feature vector is greater than a similarity threshold, and both Q and P are positive integers; When there is no such target second feature vector in the first database, input the first feature vector into a second database to obtain training set data and input the training set data into the second model to obtain the first information, the second database includes W third feature vectors and E training set data corresponding to the W third feature vectors, and both W and E are positive integers.

3. The method according to claim 2, wherein The second database includes a real-time database, and the real-time database includes real-time training set data, and the real-time training set data is training set data obtained through real-time search.

4. The method according to any one of claims 1 to 3, characterized in that, The adjusting the parameters of the second model according to the first weight coefficient includes: When the target image generated by the first data is a portrait or an animal picture, input the first weight information and the first data into a third model to obtain target intermediate data, and the target intermediate data is the intermediate data of the portrait or the intermediate data of the animal picture; Adjust the parameters of the second model according to the target intermediate data.

5. The method according to any one of claims 1 to 4, characterized in that The inputting second data into the second model to obtain the target image includes: Input the second data into the second model to obtain M generated images, where M is a positive integer; Determine N generated images among the M generated images as the target image, where N is a positive integer.

6. The method according to any one of claims 1 to 5, characterized in that, The first information further includes an identifier, and the identifier is used to characterize the style of the target image.

7. The method according to any one of claims 1 to 6, characterized in that The second data further includes an identifier, and the identifier is used to characterize the style of the target image.

8. The method according to any one of claims 1 to 7, characterized in that The first data includes at least one of text, picture, audio, or video.

9. An image generation device, characterized in that, Including: A transceiver unit and a processing unit; The transceiver unit is used to obtain first data, where the first data is used to generate a target image; The processing unit is used to determine first information according to the first data, the first information includes a first weight coefficient, the first weight coefficient is the weight coefficient of a first model, the first model is used to generate adjustment parameters of a second model, and the second model is used to generate the target image; The processing unit is further configured to adjust the parameters of the second model according to the first weight coefficient; The transceiver unit is further configured to input the second data into the second model to obtain the target image, where the second data includes the first data.

10. The device according to claim 9, characterized in that, Specifically, the processing unit is configured to: Determine a first feature vector according to the first data; When there is a target second feature vector in the first database, input the first feature vector into the first database to obtain the first information, where the first database includes Q second feature vectors and P first information corresponding to the Q second feature vectors, and the target second feature vector is a second feature vector whose similarity to the first feature vector is greater than a similarity threshold, and both Q and P are positive integers; When there is no target second feature vector in the first database, input the first feature vector into the second database to obtain training set data and input the training set data into the second model to obtain the first information, where the second database includes W third feature vectors and E training set data corresponding to the W third feature vectors, and both W and E are positive integers.

11. The device according to claim 10, characterized in that, The second database includes a real-time database, and the real-time database includes real-time training set data, where the real-time training set data is training set data obtained through real-time search.

12. The device according to any one of claims 9 to 11, characterized in that, Specifically, the processing unit is configured to: When the target image generated by the first data is a human figure or an animal figure, input the first weight information and the first data into a third model to obtain target intermediate data, where the target intermediate data is the intermediate data of the human figure or the intermediate data of the animal figure; Adjust the parameters of the second model according to the target intermediate data.

13. The device according to any one of claims 9 to 12, characterized in that, Specifically, the transceiver unit is configured to: Input the second data into the second model to obtain M generated images, where M is a positive integer; Determine N generated images among the M generated images as the target image, where N is a positive integer.

14. The device according to any one of claims 9 to 13, characterized in that, The first information further includes an identifier, and the identifier is used to characterize the style of the target image.

15. The device according to any one of claims 9 to 14, characterized in that, The second data further includes an identifier, and the identifier is used to characterize the style of the target image.

16. The device according to any one of claims 9 to 15, characterized in that The first data includes at least one of text, picture, audio, or video.

17. An image generation device, comprising at least one processor and a memory, characterized in that, The at least one processor executes a program or instruction stored in the memory, so that the image generation device implements the method described in any one of claims 1 to 8 above.

18. A computer-readable storage medium for storing a computer program, characterized in that, When the computer program runs on a computer or a processor, the computer or the processor implements the method described in any one of claims 1 to 8 above.

19. A computer program product, characterized in that the computer program product contains instructions, When the instruction runs on a computer or a processor, the computer or the processor implements the method described in any one of claims 1 to 8 above.

20. A chip, comprising at least one processor and a memory, characterized in that, The at least one processor executes a program or instruction stored in the memory to implement the method described in any one of claims 1 to 8 above.

Citation Information

Cited By

  • Image generation method and apparatus

    WO2025148372A1