Image generation method and apparatus
By obtaining user input data to adjust the parameters of the image generation model, and using knowledge base and real-time database to optimize the image generation method, the problem of insufficient image generation accuracy is solved, and higher matching and applicability is achieved.
Patent Information
- Application Number
- PCT/CN2024/118259
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-08
- Filing Date
- 2024-09-11
- Publication Date
- 2025-07-17
AI Technical Summary
The accuracy of the existing image generation method is insufficient, and the degree of correspondence between the generated image and the user input data is not high.
By obtaining user input data, the first information is determined, including the first weight coefficient, the parameters of the second model are adjusted, and the second data is input into the second model to generate a target image, and the parameters of the image generation model are optimized using the knowledge base and real-time database to improve the matching degree.
The accuracy of the image generation method is improved, so that the degree of correspondence between the generated image and the user input data is increased, and it is suitable for more image generation scenarios.
Smart Images

Figure CN2024118259_17072025_PF_FP_ABST
Abstract
Description
Image generation method and device
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on January 8, 2024, with application number 202410028193.5 and application name “Image Generation Method and Device”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The embodiments of the present application relate to the field of image processing technology, and in particular to an image generation method and apparatus. Background Art
[0003] In the field of artificial intelligence, computer vision technology is applied in many fields such as image generation, semantic segmentation, and target detection. As an important part of computer vision technology, image generation technology is gradually becoming popular in the fields of science and art.
[0004] Image generation methods can convert user input data (such as text, images, etc.) into generated images corresponding to the user input data. The generated images of image generation methods can be real photos, paintings, three-dimensional (3D) renderings, or completely imaginary images.
[0005] The accuracy of an image generation method is used to indicate how well the generated image corresponds to the user input data.
[0006] Summary of the Invention
[0007] The present invention provides an image generation method and apparatus for improving the accuracy of the image generation method. To achieve the above-mentioned purpose, the present invention adopts the following technical solutions:
[0008] In a first aspect, embodiments of the present application provide an image generation method, comprising: acquiring first data; determining first information based on the first data; adjusting parameters of a second model based on a first weight coefficient; and inputting second data into the second model to obtain a target image. The first data is used to generate the target image, the first information includes a first weight coefficient, which is a weight coefficient of the first model, the first model is used to generate adjustment parameters of the second model, the second model is used to generate the target image, and the second data includes the first data.
[0009] It can be seen that the method provided in the embodiment of the present application, based on the image generation technology, additionally introduces the step of adjusting the parameters of the image generation model (i.e., the second model) through user input data (i.e., the first data). By adjusting the parameters of the image generation model through user input data, the matching degree between the image generation model and the user input data can be improved, and the degree of correspondence between the generated image output by the image generation model and the user input data can be increased, thereby improving the accuracy of the image generation method.
[0010] In one possible implementation, the first eigenvector can be determined based on the first data. If the target second eigenvector exists in the first database, the first eigenvector is input into the first database to obtain the first information. The first database includes Q second eigenvectors and P first information corresponding to the Q second eigenvectors. The target second eigenvector is a second eigenvector whose similarity to the first eigenvector is greater than a similarity threshold, and Q and P are both positive integers. If the target second eigenvector does not exist in the first database, the first eigenvector is input into the second database to obtain training set data and the training set data is input into the second model to obtain the first information. The second database includes W third eigenvectors and E training set data corresponding to the W third eigenvectors, and W and E are both positive integers.
[0011] It can be seen that the method provided in the embodiment of the present application, on the one hand, can directly adjust the parameters of the image generation model using the existing weight coefficients in the knowledge base to improve the matching degree between the image generation model and the user input data when the knowledge base has weight coefficients that match the user input data. On the other hand, when the knowledge base does not have weight coefficients that match the user input data, the existing data in the knowledge base can be used as a training set to generate weight coefficients that match the user input data in real time to improve the matching degree between the image generation model and the user input data, as well as the versatility of the image generation model, so that the image generation model can be applied to more image generation scenarios.
[0012] In a possible implementation, the similarity includes sine-cosine similarity, Euclidean distance, Manhattan distance, Chebyshev distance, Hamming distance or other similarities.
[0013] In a possible implementation, the second database may include a real-time database, the real-time database includes real-time training set data, and the real-time training set data is training set data obtained through real-time search.
[0014] It can be seen that the method provided in the embodiment of the present application can obtain training set data that matches user data through real-time search, so as to improve the matching degree between the image generation model and the user input data, as well as the real-time performance of the image generation model, so that the image generation model can generate images with a higher degree of matching with real-time hotspots.
[0015] In one possible implementation, when the target image generated by the first data is a picture of a person or an animal, the first weight information and the first data can be input into a third model to obtain target intermediate data, where the target intermediate data is the intermediate data of the person or the intermediate data of the animal; and the parameters of the second model are adjusted according to the target intermediate data.
[0016] It can be seen that the method provided in the embodiment of the present application can adjust the parameters of the image generation model according to the intermediate data generated by the third model (such as a human outline image, an animal outline image or a posture image) when generating a human image or an animal image, so that the image generation model can more accurately output a human image or an animal image that matches the user input data.
[0017] In a possible implementation, the second data may be input into the second model to obtain M generated images, where M is a positive integer; and N generated images among the M generated images are determined as the target images, where N is a positive integer.
[0018] For example, objective indicators may be calculated for M generated images, and N generated images with the highest objective indicator scores among the M generated images may be determined as the target images.
[0019] In a possible implementation, the first information may further include an identifier, where the identifier is used to characterize the style of the target image.
[0020] It can be understood that by adjusting the parameters of the image generation model through the identifier used to characterize the image style corresponding to the user data, the matching degree between the image generation model and the user input data can be further improved, so that the degree of correspondence between the generated image output by the image generation model and the user input data can be further increased, thereby further improving the accuracy of the image generation method.
[0021] In a possible implementation, the second data may further include an identifier, where the identifier is used to characterize the style of the target image.
[0022] It is understandable that additionally inputting an identifier representing the image style corresponding to the user data into the image generation model can further increase the degree of correspondence between the generated image output by the image generation model and the user input data, thereby further improving the accuracy of the image generation method.
[0023] In a possible implementation manner, the first data may include at least one of text, picture, audio, or video.
[0024] It can be understood that, on the one hand, compared to using only text as input data for the image generation model to generate images, the additional introduction of pictures, audio, and video as input data for the image generation model to generate images can improve the applicability of the image generation model and make the image generation model applicable to more image generation scenarios. On the other hand, by adjusting the parameters of the image generation model through the text, pictures, audio, and video input by the user, the matching degree between the image generation model and the user input data can be improved, and the degree of correspondence between the generated image output by the image generation model and the user input data can be increased, thereby improving the accuracy of the image generation method.
[0025] In a second aspect, an embodiment of the present application provides an image generation device, which includes: a transceiver unit and a processing unit. The transceiver unit is used to obtain first data, and the first data is used to generate a target image. The processing unit is used to determine first information based on the first data, and the first information includes a first weight coefficient, and the first weight coefficient is the weight coefficient of the first model. The first model is used to generate adjustment parameters of the second model, and the second model is used to generate the target image. The processing unit is also used to adjust the parameters of the second model according to the first weight coefficient. The transceiver unit is also used to input second data into the second model to obtain the target image, and the second data includes the first data.
[0026] In one possible implementation, the processing unit is specifically used to: determine a first feature vector based on the first data; if the target second feature vector exists in the first database, input the first feature vector into the first database to obtain the first information, the first database includes Q second feature vectors and P first information corresponding to the Q second feature vectors, the target second feature vector is a second feature vector whose similarity with the first feature vector is greater than a similarity threshold, and Q and P are both positive integers; if the target second feature vector does not exist in the first database, input the first feature vector into the second database to obtain training set data and input the training set data into the second model to obtain the first information, the second database includes W third feature vectors and E training set data corresponding to the W third feature vectors, and W and E are both positive integers.
[0027] In a possible implementation, the second database includes a real-time database, the real-time database includes real-time training set data, and the real-time training set data is training set data obtained through real-time search.
[0028] In one possible implementation, the processing unit is specifically used to: when the target image generated by the first data is a picture of a person or an animal, input the first weight information and the first data into a third model to obtain target intermediate data, where the target intermediate data is the intermediate data of the person or the intermediate data of the animal; and adjust the parameters of the second model according to the target intermediate data.
[0029] In one possible implementation, the transceiver unit is specifically used to: input the second data into the second model to obtain M generated images, where M is a positive integer; and determine N of the M generated images as the target images, where N is a positive integer.
[0030] In a possible implementation, the first information further includes an identifier, and the identifier is used to represent the style of the target image.
[0031] In a possible implementation, the second data further includes an identifier, and the identifier is used to characterize the style of the target image.
[0032] In a possible implementation, the first data includes at least one of text, picture, audio, or video.
[0033] In a third aspect, an embodiment of the present application further provides an image generating device, which includes: at least one processor, which implements the above method in the above first aspect or any possible implementation thereof when the above at least one processor executes program code or instructions.
[0034] Optionally, the image generating device may further include at least one memory, and the at least one memory is used to store the program code or instruction.
[0035] In a fourth aspect, embodiments of the present application further provide a chip comprising: an input interface, an output interface, and at least one processor. Optionally, the chip further comprises a memory. The at least one processor is configured to execute code in the memory. When the at least one processor executes the code, the chip implements the method described in the first aspect or any possible implementation thereof.
[0036] Optionally, the chip may also be an integrated circuit.
[0037] In a fifth aspect, an embodiment of the present application further provides a computer-readable storage medium for storing a computer program, wherein the computer program includes methods for implementing the above-mentioned first aspect or any possible implementation thereof.
[0038] In a sixth aspect, an embodiment of the present application further provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to implement the above-mentioned method in the above-mentioned first aspect or any possible implementation thereof. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0040] FIG1 is a schematic diagram of the structure of an image processing system provided in an embodiment of the present application;
[0041] FIG2 is a schematic diagram of the structure of a database provided in an embodiment of the present application;
[0042] FIG3 is a schematic diagram of the structure of a first database provided in an embodiment of the present application;
[0043] FIG4 is a schematic diagram of the structure of a second database provided in an embodiment of the present application;
[0044] FIG5 is a schematic diagram of a flow chart of an image generation method provided in an embodiment of the present application;
[0045] FIG6 is a schematic diagram of a data filtering process provided in an embodiment of the present application;
[0046] FIG7 is a schematic diagram of a data retrieval process provided in an embodiment of the present application;
[0047] FIG8 is a schematic diagram of a model training process provided in an embodiment of the present application;
[0048] FIG9 is a schematic diagram of an image generation process provided in an embodiment of the present application;
[0049] FIG10 is a schematic structural diagram of an image generating device provided in an embodiment of the present application;
[0050] FIG11 is a schematic diagram of the structure of a chip provided in an embodiment of the present application;
[0051] FIG12 is a schematic structural diagram of an electronic device provided in an embodiment of the present application;
[0052] FIG13 is a schematic structural diagram of another image generating device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0053] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments of this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the embodiments of this application.
[0054] The term "and / or" in this article is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone.
[0055] The terms "first" and "second" and so on in the description and drawings of the embodiments of this application are used to distinguish different objects, or to distinguish different processing of the same object, rather than to describe a specific order of objects.
[0056] Furthermore, the terms "including," "having," and any variations thereof, mentioned in the description of the embodiments of the present application are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or units is not limited to the listed steps or units but may optionally include other steps or units not listed, or may optionally include other steps or units inherent to the process, method, product, or apparatus.
[0057] It should be noted that in the description of the embodiments of this application, words such as "exemplarily" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "exemplarily" or "for example" in the embodiments of this application should not be interpreted as having priority or advantage over other embodiments or designs. Rather, the use of words such as "exemplarily" or "for example" is intended to present the relevant concepts in a concrete manner.
[0058] Before introducing the technical solutions of this application, some of the terms involved in this application are explained. The following related explanations can be combined with the technical solutions of the embodiments of this application as optional solutions, and they all fall within the scope of protection of the embodiments of this application. The embodiments of this application include at least part of the following contents.
[0059] Stable diffusion model (SD): An open-source generative model for text-to-image translation based on deep learning, mainly used to generate detailed images based on text descriptions.
[0060] Large language model (LLM): A very large deep learning model pre-trained based on a large amount of data. It can perform open-domain tasks such as text summarization and language translation, and has certain logical thinking and reasoning capabilities.
[0061] Low rank adaptation (LORA): A technique that reduces the complexity of large models by approximating their high-dimensional structure with a low-dimensional structure, enabling efficient fine-tuning of the model. This technique can adapt large models to specific tasks or domains more quickly and effectively.
[0062] Contrastive Language-Image Pre-training (CLIP): A pre-trained neural network model for matching images and text. This task is widely used in multimodal fields and is used for text-image retrieval. The model is pre-trained using a large amount of paired internet data. The structure uses two branches: a text encoder and an image encoder. It learns the matching similarity between images and text, and can then retrieve images based on text, or text based on images. In addition, the training strategy of image-text contrast learning can also be applied to retrieval tasks in other modalities (such as audio and video).
[0063] Dream Booth (DB): A customized text-to-image generation method that can fine-tune the diffusion model through custom themes and achieve good results with only a small amount of training data.
[0064] Bootstrapping Language-Image Pre-training 2 (BLIP2): A visual language model for multimodal understanding of images and text. The model consists of a visual encoder, a text encoder, a query transformer (Q-Former), and an LLM to align the visual and language modalities. This method leverages the pretrained visual encoder and LLM to enable image captioning, visual question answering, and fine-grained image retrieval. Furthermore, by replacing the LLM with a model with larger parameters and the ability to follow instructions and chain thoughts, it enables retrieval of local information from both images and text.
[0065] Prompt: Enter the descriptive text for the stable diffusion model.
[0066] Text-generated image: It generates semantically consistent images based on the text description entered by the user, such as the core elements and semantic relationships contained in the text.
[0067] Image synthesis: It is to fuse the current image according to the style of a set of images provided by the user.
[0068] Image restoration and super-resolution: adjust the details and resolution of the image based on the original image provided by the user.
[0069] Image generation: based on the image provided by the user, part of the image area is edited and the image generation is controlled by external conditions.
[0070] In the field of artificial intelligence, computer vision technology is applied in many fields such as image generation, semantic segmentation, and target detection. As an important part of computer vision technology, image generation technology is gradually becoming popular in the fields of science and art.
[0071] Image generation methods can convert user input data (such as text, images, etc.) into generated images corresponding to the user input data. The generated images of the image generation method can be real photos, paintings, 3D renderings, or completely imaginary images.
[0072] The accuracy of an image generation method is used to indicate how well the generated image corresponds to the user input data.
[0073] To this end, an embodiment of the present application provides an image generation method to improve the accuracy of the image generation method.
[0074] The technical solutions provided in the embodiments of this application can be applied to image processing systems. The application scenarios of the technical solutions provided in the embodiments of this application can include various scenarios, such as text generation into images, image synthesis (image style conversion), image restoration and super-resolution, image generation into images, etc.
[0075] FIG1 shows a schematic diagram of a possible, non-limiting image processing system as described above. As shown in FIG1 , the image processing system 100 includes an electronic device 110 and a database 120 .
[0076] The electronic device 110 may be an electronic device such as a mobile phone, a desktop computer, a tablet computer, a laptop computer, a vehicle-mounted terminal, a server, an intelligent robot, a smart TV, a multimedia playback device, or other electronic devices, and the embodiments of the present application are not limited thereto.
[0077] The database 120 may be a knowledge base. A knowledge base is a special database used for knowledge management. The knowledge base can store information such as text, images, videos, and audio, and retrieve corresponding data from the data stored in the knowledge base upon request.
[0078] It should be understood that the structure of the image processing system 100 shown in FIG1 does not constitute a specific limitation on the image processing system 100. In other embodiments of the present application, the image processing system 100 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0079] FIG2 shows a possible, non-limiting schematic diagram of the above database. As shown in FIG2 , the database 120 may include a first database 121 and a second database 122 .
[0080] The first database 121 is used to store the weight coefficients of the model.
[0081] The second database 122 is used to store model training data sets (such as pictures, texts, videos or audio data).
[0082] It should be understood that the structure of database 120 shown in FIG2 does not constitute a specific limitation on database 120. In other embodiments of the present application, database 120 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0083] Fig. 3 shows a possible, non-limiting schematic diagram of the first database. As shown in Fig. 3, the first database 121 may include a LORA knowledge base and a DB knowledge base.
[0084] The LORA knowledge base is used to provide LORA weight information that matches style or specific concept elements.
[0085] The DB knowledge base is used to provide identifiers and LORA weight information matching concepts.
[0086] It should be understood that the structure of the first database 121 shown in FIG3 does not constitute a specific limitation on the first database 121. In other embodiments of the present application, the first database 121 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0087] Fig. 4 shows a possible, non-limiting schematic diagram of the second database. As shown in Fig. 4, the second database 122 may include a general knowledge base, a domain knowledge base, and a real-time knowledge base.
[0088] A general knowledge base that provides open-domain data such as text, drawings, videos, and audio.
[0089] Domain knowledge base, used to provide text, drawings, videos, audio and other data in vertical fields (such as portraits, animals, etc.).
[0090] A real-time knowledge base is used to provide text, drawings, videos, audio and other data of real-time hot concepts, and text, drawings, videos, audio and other data that can be queried online.
[0091] It should be understood that the structure of the second database 122 shown in FIG4 does not constitute a specific limitation on the second database 122. In other embodiments of the present application, the second database 122 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0092] It is understood that this application uses electronic devices and databases as examples of the execution entities of the interactive diagrams, but this application does not limit the execution entities of the interactive diagrams. For example, the electronic device in the method of this application can also be a chip, chip system or processor applied to the electronic device, or a logical node, logical module or software that can realize all or part of the electronic device; the knowledge base in the method of this application can also be a chip, chip system or processor applied to the knowledge base, or a logical node, logical module or software that can realize all or part of the knowledge base functions.
[0093] FIG5 shows an image generation method provided by an embodiment of the present application. As shown in FIG5 , the method includes:
[0094] S501: Acquire first data.
[0095] The first data is used to generate a target image.
[0096] In a possible implementation, the first data may include at least one of text, picture, audio, or video, that is, the first data may be composed of text, picture, audio, or video.
[0097] The text may include a classifier. A classifier is used to indicate the type of the target image. For example, the classifier "vehicle" is used to indicate that the target image is a vehicle image.
[0098] The text may include positive prompt words and / or negative prompt words.
[0099] For example, the user wants to generate an image of a zebra, and the first data acquired by the electronic device may be the text "generate an image of a zebra".
[0100] As another example, the user wants to generate an image of a train, and the first data acquired by the electronic device may be a picture of a train.
[0101] Of course, the first data acquired by the electronic device may also be composed of multiple types of data, such as picture+text, picture+audio, audio+text, and other types of data.
[0102] In a possible implementation, the first data may be filtered to filter out potential sensitive information in the first data.
[0103] Exemplarily, as shown in FIG6 , the content in the first data may be divided into text, picture, audio, and video.
[0104] As shown in FIG6 , a vocabulary (such as a stop word vocabulary) can be used to filter the text in the first data. If sensitive words are present in the text of the first data, the first data or the text content of the first data can be processed. For example, the first data or the text content of the first data can be invalidated, that is, an image is no longer generated based on the first data or the text content in the first data.
[0105] As shown in FIG6 , a convolutional neural network (such as a You Only Look Once (YOLO) network) may be used to filter images in the first data using a content detection model (such as a Not Safe for Work (NSFW) detection model). If the detection value of the image is greater than a detection threshold (i.e., it is determined that sensitive content exists in the image), the first data or the image content of the first data may be processed. For example, the first data or the image content of the first data may be invalidated, i.e., an image is no longer generated based on the first data or based on the image content in the first data.
[0106] As shown in FIG6 , a convolutional neural network can be used to filter the video in the first data using a content detection model. When the detection value of the video is greater than a detection threshold (i.e., it is determined that sensitive content exists in the video), the first data or the video content of the first data can be processed. For example, the first data or the video content of the first data can be invalidated, that is, a video is no longer generated based on the first data, or an image is no longer generated based on the video content in the first data.
[0107] As shown in FIG6 , the acoustic model may be used to filter the audio in the first data. If the audio detection value is greater than a detection threshold (i.e., it is determined that sensitive content exists in the audio), the first data or the audio content of the first data may be processed. For example, the first data or the audio content of the first data may be invalidated, i.e., audio is no longer generated based on the first data, or an image is no longer generated based on the audio content in the first data.
[0108] In a possible implementation, the first data may be standardized so as to make the first data conform to the input of the knowledge base.
[0109] As shown in FIG6 , text normalization processing may be performed on the text in the first data to obtain standardized text.
[0110] For example, special characters in the text content of the first data may be filtered to obtain standardized text.
[0111] As shown in FIG6 , image normalization processing is performed on the pictures in the first data to obtain normalized pictures.
[0112] For example, the pictures in the first data may be uniformly converted into standardized pictures in a red, green, blue (RGB) format, a resize format, or a sample format.
[0113] As shown in FIG6 , image normalization processing is performed on the video in the first data to obtain a standardized video.
[0114] For example, the video content in the first data may be uniformly converted into standardized videos in RGB format, resize format, or sample format.
[0115] As shown in FIG6 , the audio in the first data may be subjected to audio normalization processing.
[0116] For example, the audio content in the first data may be uniformly converted into standardized audio in a sample format.
[0117] S502: Determine first information according to the first data.
[0118] The first information includes a first weight coefficient, which is a weight coefficient of a first model. The first model is used to generate adjustment parameters of a second model, and the second model is used to generate the target image.
[0119] Exemplarily, the first model may be a Lora model, and the second model may be an SD model.
[0120] Exemplarily, the first information may further include an identifier, where the identifier is used to characterize the style of the target image.
[0121] For example, if the identifier is [token][cartoon], then the identifier indicates that the style of the target image is cartoon style.
[0122] For example, a task query instruction may be generated according to the first data through an agent index, and the database may be queried according to the task query instruction to obtain the first information.
[0123] In one possible implementation, a first feature vector can be determined based on the first data. If the target second feature vector exists in the first database, the first feature vector is input into the first database to obtain the first information. If the target second feature vector does not exist in the first database, the first feature vector is input into the second database to obtain training set data, and the training set data is input into the second model to obtain the first information.
[0124] As shown in Table 1, the first database may include Q second eigenvectors and P first information corresponding to the Q second eigenvectors. The target second eigenvector is a second eigenvector whose similarity to the first eigenvector is greater than a similarity threshold, where Q and P are both positive integers. Each second eigenvector may correspond to one or more first information, and each first information may correspond to one or more second eigenvectors.
[0125] Table 1
[0126] It can be understood that, in the case where the target second eigenvector exists, the first information obtained by inputting the above-mentioned first eigenvector into the first database includes the first information corresponding to the target second eigenvector.
[0127] For example, the similarity between the second eigenvector 2 and the second eigenvector 3 and the first eigenvector is greater than the similarity, that is, the second eigenvector 2 and the second eigenvector 3 are the target second eigenvectors, and the first information obtained by inputting the first eigenvector into the first database includes the first information 2 corresponding to the second eigenvector 2 and the first information 3 corresponding to the second eigenvector 3.
[0128] As shown in Table 2, the second database includes W third eigenvectors and E training set data corresponding to the W third eigenvectors, where W and E are both positive integers. Each third eigenvector may correspond to one or more training data sets, and each training data set may correspond to one or more third eigenvectors.
[0129] Table 2
[0130] In a possible implementation, the similarity includes sine-cosine similarity, Euclidean distance, Manhattan distance, Chebyshev distance, Hamming distance or other similarities.
[0131] It is understood that, in the absence of a target second eigenvector, the training set obtained by inputting the first eigenvector into the first database includes a training data set corresponding to the target third eigenvector. The target third eigenvector is a third eigenvector having a similarity with the first eigenvector greater than a similarity threshold.
[0132] For example, the similarity between the third eigenvector 1 and the first eigenvector is greater than the similarity, that is, the third eigenvector 1 is the target third eigenvector, and the training data set obtained by inputting the first eigenvector into the second database includes the training data set 1 corresponding to the third eigenvector 1.
[0133] Exemplarily, as shown in FIG7 , when the target second feature vector exists in the first database, the first feature vector obtained according to the first data is input into the first database to obtain the first information.
[0134] As another example, as shown in FIG7 , when the target second feature vector does not exist in the first database, the first feature vector obtained based on the first data is input into the second database to obtain training set data. The obtained training set data is input into the second model to obtain the first information.
[0135] Optionally, the second database includes a real-time database (real-time knowledge base), and the real-time database includes real-time training set data, which is training set data obtained through real-time search.
[0136] For example, after the first feature vector is input into the real-time database, the real-time database can search in real time according to the first feature vector to obtain training set data, and output the training set data obtained by the real-time search according to the first feature vector.
[0137] In one possible implementation, the first database can be trained based on the first information obtained from the training set data. Specifically, the first feature vector and the first information obtained from the training set data can be sent to the first database. This allows the first database to directly output the corresponding first information upon receiving the first feature vector, without having to obtain the training data set from the second database.
[0138] In a possible implementation, the first information may also be displayed. For example, the first information may be displayed on a display screen of the electronic device.
[0139] In a possible implementation, the first feature vector may include a first coarse-grained feature vector and / or a first fine-grained feature vector.
[0140] For the text content in the first data, the first fine-grained feature vector represents the global text, and the first fine-grained feature vector represents the key fragment of the text.
[0141] For the video, image, and audio in the first data, the first fine-grained feature vector represents the global content after sampling, and the first fine-grained feature vector represents the high-frequency and regional subject information after sampling.
[0142] In a possible implementation, the first data may be input into a vector model to obtain a first feature vector.
[0143] In a possible implementation, the above-mentioned vector model may include a CLIP model, a BLIP2 model, or an LLM model.
[0144] Exemplarily, the first data may be input into the CLIP model to obtain a first coarse-grained feature vector.
[0145] As another example, the first data may be input into the BLIP2 model to obtain a first fine-grained feature vector.
[0146] S503: Adjust the parameters of the second model according to the first weight coefficient.
[0147] In one possible implementation, when the target image generated by the first data is a picture of a person or an animal, the first weight information and the first data can be input into a third model to obtain target intermediate data; and the parameters of the second model can be adjusted according to the target intermediate data.
[0148] The target intermediate data may be the intermediate data of the human figure or the intermediate data of the animal figure. For example, the target intermediate data may be a human outline figure, an animal outline figure, or a posture figure.
[0149] Exemplarily, the third model may be a control network (controlnet) model.
[0150] Exemplarily, as shown in Figure 8, when the target image generated by the above-mentioned first data is a picture of a person or an animal, the above-mentioned first weight information and the above-mentioned first data are input into the third model to obtain target intermediate data, and then the parameters of the above-mentioned second model are adjusted according to the above-mentioned target intermediate data to obtain the trained second model.
[0151] As another example, as shown in FIG8 , when the target image generated by the first data is neither a person image nor an animal image, the parameters of the second model can be adjusted directly according to the first weight coefficient to obtain the trained second model.
[0152] It should be noted that the first model can be understood as a fine-tuning model. The fine-tuning model can be connected to other models and the parameters of other models can be frozen. By only fine-tuning the weight coefficient of the first model, other models can be trained. Therefore, the parameters of the second model can be adjusted through the first weight coefficient.
[0153] For example, if the first model is a LoRA model and the second model is an SD model, the LoRA model can be a small model trained under the SD model and used to fine-tune the large model. LoRA can adjust the character, style, and so on.
[0154] A fine-tuning model can be considered a plug-in for a larger model (such as the SD model). Without modifying the SD model, it uses a small amount of data to train a specific style / intellectual property (IP) / character to meet customized requirements. This requires fewer training resources than training an SD model. To reduce training costs, the LoRA model only trains low-rank matrices. When used, the LoRA model's parameters are injected into the SD model, thereby changing the SD model's generation style or adding new characters / IPs. The entire process is a linear relationship, and can be thought of as a completely new model obtained by superimposing the LoRA model on the original SD model.
[0155] S504: Input the second data into the second model to obtain a target image.
[0156] The second data includes the first data.
[0157] In one possible implementation, the second data may further include an identifier that is used to characterize the style of the target image. For example, if the identifier [token][Gothic] characterizes the style of the target image as Gothic.
[0158] In a possible implementation, the second data may be input into the second model to obtain M generated images, where M is a positive integer; and N generated images among the M generated images are determined as the target images, where N is a positive integer.
[0159] For example, as shown in FIG9 , the second data can be input into the trained second model to obtain M generated images, and then the M generated images can be input into the image evaluation model to obtain evaluation scores for the M images. The N generated images with the highest evaluation scores are then input as target images.
[0160] For example, the second data can be encoded using an encoder to generate encoded data. This encoded data can then be fed into the trained second model to generate M generated images. These M generated images can then be fed into an image evaluation model to obtain evaluation scores for the M images. The four generated images with the highest evaluation scores are then fed into the model as target images.
[0161] Exemplarily, the above-mentioned image evaluation model may include a semantic consistency model (such as a CLIP model) and / or an aesthetic model (such as a human preference assessment (HPS) model).
[0162] In a possible implementation, a target image may also be displayed. For example, the target image may be displayed on a display screen of an electronic device.
[0163] It can be seen that the method provided in the embodiment of the present application, based on the image generation technology, additionally introduces the step of adjusting the parameters of the image generation model (i.e., the second model) through user input data (i.e., the first data). By adjusting the parameters of the image generation model through user input data, the matching degree between the image generation model and the user input data can be improved, and the degree of correspondence between the generated image output by the image generation model and the user input data can be increased, thereby improving the accuracy of the image generation method.
[0164] The following will introduce an image generating device for executing the above-mentioned image generating method with reference to FIG10 .
[0165] It is understandable that, in order to realize the above functions, the image generating device includes hardware and / or software modules corresponding to the execution of each function. In combination with the algorithm steps of each example described in the embodiments disclosed herein, the embodiments of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application in combination with the embodiments, but such implementation should not be considered to exceed the scope of the embodiments of the present application.
[0166] In the embodiment of the present application, the image generation device can be divided into functional modules according to the above method example. For example, each functional module can be divided according to each function, or two or more functions can be integrated into one processing module. The above integrated modules can be implemented in the form of hardware. It should be noted that the division of modules in this embodiment is schematic and is only a logical function division. In actual implementation, other division methods may be used.
[0167] FIG10 illustrates a possible schematic diagram of the image generation device involved in the above embodiments, where functional modules are divided according to their respective functions. The device may be an electronic device, a module (such as a processor, chip, or chip system) used in an electronic device, or a logical node, logic module, or software that implements all or part of the functions of an electronic device. As shown in FIG10 , the image generation device 1000 may include a transceiver unit 1001 and a processing unit 1002.
[0168] The transceiver unit 1001 is configured to obtain first data, where the first data is used to generate a target image.
[0169] The processing unit 1002 is used to determine the first information based on the first data. The above-mentioned first information includes a first weight coefficient. The above-mentioned first weight coefficient is the weight coefficient of the first model. The above-mentioned first model is used to generate the adjustment parameters of the second model. The above-mentioned second model is used to generate the above-mentioned target image.
[0170] The processing unit 1002 is further configured to adjust the parameters of the second model according to the first weight coefficient.
[0171] The transceiver unit 1001 is further configured to input second data into the second model to obtain the target image, where the second data includes the first data.
[0172] In one possible implementation, the processing unit 1001 is specifically used to: determine a first feature vector based on the first data; if the target second feature vector exists in the first database, input the first feature vector into the first database to obtain the first information, the first database includes Q second feature vectors and P first information corresponding to the Q second feature vectors, the target second feature vector is a second feature vector whose similarity with the first feature vector is greater than a similarity threshold, and Q and P are both positive integers; if the target second feature vector does not exist in the first database, input the first feature vector into the second database to obtain training set data and input the training set data into the second model to obtain the first information, the second database includes W third feature vectors and E training set data corresponding to the W third feature vectors, and W and E are both positive integers.
[0173] In a possible implementation, the second database includes a real-time database, the real-time database includes real-time training set data, and the real-time training set data is training set data obtained through real-time search.
[0174] In one possible implementation, the processing unit 1002 is specifically used to: when the target image generated by the first data is a person image or an animal image, input the first weight information and the first data into a third model to obtain target intermediate data, where the target intermediate data is the intermediate data of the person image or the intermediate data of the animal image; and adjust the parameters of the second model according to the target intermediate data.
[0175] In a possible implementation, the transceiver unit 1001 is specifically configured to: input the second data into the second model to obtain M generated images, where M is a positive integer; and determine N of the M generated images as the target images, where N is a positive integer.
[0176] In a possible implementation, the first information further includes an identifier, and the identifier is used to represent the style of the target image.
[0177] In a possible implementation, the second data further includes an identifier, and the identifier is used to characterize the style of the target image.
[0178] In a possible implementation, the first data includes at least one of text, picture, audio, or video.
[0179] The present application also provides a chip. FIG11 shows a schematic diagram of the structure of a chip 1100. The chip 1100 includes one or more processors 1101 and an interface circuit 1102. Optionally, the chip 1100 may also include a bus 1103.
[0180] The processor 1101 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above image generation method can be completed by the hardware integrated logic circuit in the processor 1101 or the software instruction.
[0181] Optionally, the processor 1101 may be a general-purpose processor, a digital signal processing (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The methods and steps disclosed in the embodiments of the present application may be implemented or executed. The general-purpose processor may be a microprocessor or any conventional processor.
[0182] The interface circuit 1102 can be used to send or receive data, instructions or information. The processor 1101 can use the data, instructions or other information received by the interface circuit 1102 to process it, and can send the processing completion information through the interface circuit 1102.
[0183] Optionally, the chip also includes a memory, which may include a read-only memory and a random access memory, and provides operating instructions and data to the processor. Part of the memory may also include a non-volatile random access memory (NVRAM).
[0184] Optionally, the memory stores an executable software module or a data structure, and the processor can perform corresponding operations by calling an operation instruction stored in the memory (the operation instruction may be stored in an operating system).
[0185] Optionally, the chip can be used in an image generation device according to an embodiment of the present application. Optionally, the interface circuit 1102 can be used to output the execution result of the processor 1101. For the image generation method provided in one or more embodiments of the present application, reference can be made to the aforementioned embodiments and will not be repeated here.
[0186] It should be noted that the corresponding functions of the processor 1101 and the interface circuit 1102 can be implemented through hardware design, software design, or a combination of hardware and software, and there is no limitation here.
[0187] FIG12 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device 1200 may be a processor or a chip or functional module in a processor. As shown in FIG12 , the electronic device 1200 includes a processor 1201 , a transceiver 1202 , and a communication line 1203 .
[0188] Among them, the processor 1201 is used to execute any step in the image generation method provided in the embodiment of the present application, and in the process of executing any step in the image generation method provided in the embodiment of the present application, the transceiver 1202 and the communication line 1203 can be optionally called to complete the corresponding operation.
[0189] Furthermore, the electronic device 1200 may further include a memory 1204 . The processor 1201 , the memory 1204 and the transceiver 1202 may be connected via a communication line 1203 .
[0190] The processor 1201 is a processor, a general-purpose processor, a network processor (NP), a digital signal processor (DSP), a microprocessor, a microcontroller, a programmable logic device (PLD), or any combination thereof. The processor 1201 may also be other devices with processing functions, such as circuits, devices, or software modules, without limitation.
[0191] Transceiver 1202 is used to communicate with other devices or other communication networks, such as Ethernet, radio access networks (RAN), wireless local area networks (WLAN), etc. Transceiver 1202 can be a module, circuit, transceiver, or any device capable of implementing communication.
[0192] The transceiver 1202 is mainly used for sending and receiving commands and information, and may include a transmitter and a receiver for sending and receiving commands and information, respectively. Operations other than sending and receiving commands and information are implemented by the processor.
[0193] The communication line 1203 is used to transmit information between the various components included in the electronic device 1200.
[0194] In one design, the processor can be considered as the logic circuit and the transceiver as the interface circuit.
[0195] The memory 1204 is used to store instructions, where the instructions may be computer programs.
[0196] The memory 1204 may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM may be used, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct RAM bus RAM (DRRAM). Memory 1204 may also be a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), magnetic disk storage media or other magnetic storage devices, etc. It should be noted that the memory of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0197] It should be noted that the memory 1204 can exist independently of the processor 1201 or can be integrated with the processor 1201. The memory 1204 can be used to store instructions, program code, or some data. The memory 1204 can be located within the electronic device 1200 or outside the electronic device 1200, without limitation. The processor 1201 is configured to execute the instructions stored in the memory 1204 to implement the methods provided in the above embodiments of the present application.
[0198] In one example, processor 1201 may include one or more processors, such as processor 0 and processor 1 in FIG. 12 .
[0199] As an optional implementation, the electronic device 1200 includes multiple processors. For example, in addition to the processor 1201 in FIG. 12 , it may also include a processor 1209 .
[0200] As an optional implementation, the electronic device 1200 further includes an output device 1205 and an input device 1206. For example, the input device 1206 is a keyboard, a mouse, a microphone, a joystick, and the like, and the output device 1205 is a display screen, a speaker, and the like.
[0201] It should be pointed out that the electronic device 1200 can be a chip system or a device with a similar structure as shown in Figure 12. Among them, the chip system can be composed of chips, or it can include chips and other discrete devices. The actions, terms, etc. involved in the various embodiments of this application can refer to each other without limitation. The message names or parameter names in the messages exchanged between the various devices in the embodiments of this application are only an example. Other names can also be used in the specific implementation without limitation. In addition, the component structure shown in Figure 12 does not constitute a limitation on the electronic device 1200. In addition to the components shown in Figure 12, the electronic device 1200 may include more or fewer components than those shown in Figure 12, or combine certain components, or arrange the components differently.
[0202] The processor and transceiver described in this application can be implemented on an integrated circuit (IC), an analog IC, a radio frequency integrated circuit, a mixed-signal IC, an application specific integrated circuit (ASIC), a printed circuit board (PCB), an electronic device, etc. The processor and transceiver can also be manufactured using various IC process technologies, such as complementary metal oxide semiconductor (CMOS), N-type metal oxide semiconductor (NMOS), P-type metal oxide semiconductor (positive channel metal oxide semiconductor, PMOS), bipolar junction transistor (BJT), bipolar CMOS (BiCMOS), silicon germanium (SiGe), gallium arsenide (GaAs), etc.
[0203] Figure 13 is a schematic diagram of the structure of an image generation device provided in an embodiment of the present application. The image generation device can be applied to the scenario shown in the above method embodiment. For ease of explanation, Figure 13 only shows the main components of the image generation device, including a processor 1301, a memory 1302, a control circuit 1303, and an input / output device 1304. The processor 1301 is mainly used to process communication protocols and communication data, execute software programs, and process data of software programs. The memory 1302 is mainly used to store software programs and data. The control circuit 1303 is mainly used for power supply and transmission of various electrical signals. The input / output device 1304 is mainly used to receive data input by the user and output data to the user.
[0204] When the image generating device is a processor 1301, the control circuit 1303 may be a motherboard, the memory 1302 may include a hard disk, RAM, ROM, or other storage media, and the processor 1301 may include a baseband processor 1301 and a central processing unit. The baseband processor is primarily used to process communication protocols and communication data, while the central processing unit is primarily used to control the entire image generating device, execute software programs, and process software program data. The input and output devices 1304 include a display screen, a keyboard, and a mouse. The control circuit 1303 may further include or be connected to a transceiver circuit or transceiver, such as a network cable interface, for sending or receiving data or signals, such as for data transmission and communication with other devices. Furthermore, an antenna may be included for transmitting and receiving wireless signals for data / signal transmission with other devices.
[0205] An embodiment of the present application also provides an image generation device, which includes: at least one processor, when the at least one processor executes program code or instructions, it implements the above-mentioned related method steps to implement the image generation method in the above-mentioned embodiment.
[0206] Optionally, the apparatus may further include at least one memory configured to store the program code or instruction.
[0207] An embodiment of the present application also provides a computer storage medium, which stores computer instructions. When the computer instructions are executed on an image generating device, the image generating device executes the above-mentioned related method steps to implement the image generating method in the above-mentioned embodiment.
[0208] An embodiment of the present application further provides a computer program product. When the computer program product is run on a computer, the computer is caused to execute the above-mentioned related steps to implement the image generation method in the above-mentioned embodiment.
[0209] The present application also provides an image generation device, which may be a chip, integrated circuit, component, or module. Specifically, the device may include a processor and a memory for storing instructions, or the device may include at least one processor configured to retrieve instructions from an external memory. When the device is in operation, the processor may execute the instructions, causing the chip to perform the image generation method described in each of the above method embodiments.
[0210] It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0211] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0212] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0213] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the above units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0214] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, and may be located in one place or distributed across multiple network units. Some or all of these units may be selected based on actual needs to achieve the objectives of this embodiment.
[0215] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0216] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the above methods of each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0217] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. An image generation method, characterized in that, Comprising: Obtain first data, where the first data is used to generate a target image; Determine first information according to the first data, where the first information includes a first weight coefficient, and the first weight coefficient is the weight coefficient of a first model, and the first model is used to generate adjustment parameters of a second model, and the second model is used to generate the target image; Adjust the parameters of the second model according to the first weight coefficient; Input second data into the second model to obtain the target image, where the second data includes the first data.
2. The method according to claim 1, characterized in that The determining first information according to the first data includes: Determine a first feature vector according to the first data; When there is a target second feature vector in a first database, input the first feature vector into the first database to obtain the first information, where the first database includes Q second feature vectors and P first information corresponding to the Q second feature vectors, and the target second feature vector is a second feature vector whose similarity to the first feature vector is greater than a similarity threshold, and both Q and P are positive integers; When there is no such target second feature vector in the first database, input the first feature vector into a second database to obtain training set data and input the training set data into the second model to obtain the first information, where the second database includes W third feature vectors and E training set data corresponding to the W third feature vectors, and both W and E are positive integers.
3. The method according to claim 2, characterized in that, The second database includes a real-time database, and the real-time database includes real-time training set data, and the real-time training set data is training set data obtained through real-time search.
4. The method according to any one of claims 1 to 3, characterized in that, The adjusting the parameters of the second model according to the first weight coefficient includes: When the target image generated by the first data is a portrait or an animal picture, input the first weight information and the first data into a third model to obtain target intermediate data, where the target intermediate data is the intermediate data of the portrait or the intermediate data of the animal picture; Adjust the parameters of the second model according to the target intermediate data.
5. The method according to any one of claims 1 to 4, characterized in that The inputting second data into the second model to obtain the target image includes: Input the second data into the second model to obtain M generated images, where M is a positive integer; Determine N generated images among the M generated images as the target image, where N is a positive integer.
6. The method according to any one of claims 1 to 5, characterized in that, The first information further includes an identifier, and the identifier is used to characterize the style of the target image.
7. The method according to any one of claims 1 to 6, characterized in that The second data further includes an identifier, and the identifier is used to characterize the style of the target image.
8. The method according to any one of claims 1 to 7, characterized in that, The first data includes at least one of text, picture, audio or video.
9. An image generation device, characterized in that, Comprising: A transceiver unit and a processing unit; The transceiver unit is used to obtain first data, where the first data is used to generate a target image; The processing unit is used to determine first information according to the first data, where the first information includes a first weight coefficient, and the first weight coefficient is the weight coefficient of a first model, and the first model is used to generate adjustment parameters of a second model, and the second model is used to generate the target image; The processing unit is further configured to adjust the parameters of the second model according to the first weight coefficient; The transceiver unit is further configured to input the second data into the second model to obtain the target image, where the second data includes the first data.
10. The device according to claim 9, wherein Specifically, the processing unit is configured to: Determine a first feature vector according to the first data; When a target second feature vector exists in the first database, input the first feature vector into the first database to obtain the first information, where the first database includes Q second feature vectors and P first information corresponding to the Q second feature vectors, and the target second feature vector is a second feature vector whose similarity to the first feature vector is greater than a similarity threshold, and both Q and P are positive integers; When the target second feature vector does not exist in the first database, input the first feature vector into the second database to obtain Training set data and input the training set data into the second model to obtain the first information, where the second database includes W third feature vectors and E training set data corresponding to the W third feature vectors, and both W and E are positive integers.
11. The device according to claim 10, wherein The second database includes a real-time database, and the real-time database includes real-time training set data, and the real-time training set data is training set data obtained through real-time search.
12. The device according to any one of claims 9 to 11, characterized in that, Specifically, the processing unit is configured to: When the target image generated by the first data is a portrait or an animal picture, input the first weight information and the first data into a third model to obtain target intermediate data, where the target intermediate data is the intermediate data of the portrait or the intermediate data of the animal picture; Adjust the parameters of the second model according to the target intermediate data.
13. The device according to any one of claims 9 to 12, characterized in that Specifically, the transceiver unit is configured to: Input the second data into the second model to obtain M generated images, where M is a positive integer; Determine N generated images among the M generated images as the target image, where N is a positive integer.
14. The device according to any one of claims 9 to 13, characterized in that, The first information further includes an identifier, and the identifier is used to characterize the style of the target image.
15. The device according to any one of claims 9 to 14, characterized in that The second data further includes an identifier, and the identifier is used to characterize the style of the target image.
16. The device according to any one of claims 9 to 15, characterized in that, The first data includes at least one of text, picture, audio, or video.
17. An image generation device, comprising at least one processor and a memory, characterized in that, The at least one processor executes the program or instructions stored in the memory, so that the image generation device implements the method described in any one of claims 1 to 8 above.
18. A computer-readable storage medium for storing a computer program, characterized in that, When the computer program runs on a computer or a processor, the computer or the processor implements the method described in any one of claims 1 to 8 above.
19. A computer program product, comprising instructions therein, characterized in that, When the instruction runs on a computer or a processor, the computer or the processor implements the method described in any one of claims 1 to 8 above.
20. A chip, comprising at least one processor and a memory, characterized in that, The at least one processor executes the program or instructions stored in the memory to implement the method described in any one of claims 1 to 8 above.
Citation Information
Patent Citations
Image generation method and device
CN120296188A
Image style migration method and device, equipment and storage medium
CN114266943A
Picture generation method and device, storage medium and computing equipment
CN117036546A
Model training method, photo generation method and related equipment
CN117078509A
Pentograph model training method and device, equipment and storage medium
CN117173504A