Image generation method, device and equipment and computer readable storage medium

By acquiring the information feature vectors of the original image and prompt information, and using the feature acquisition model to generate the target image, the problem of cumbersome image generation steps in the prior art is solved, and an efficient image generation process is realized.

CN120047579APending Publication Date: 2025-05-27BOE TECHNOLOGY GROUP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510213074.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The prior art requires users to manually enter text during image generation, resulting in cumbersome steps and low efficiency.

Method used

By obtaining the information feature vectors of the original image and prompt information, using the feature acquisition model to generate the target image, and automatically adding the target text to the original image.

Benefits of technology

No manual editing is required, which significantly saves image generation time and improves image generation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047579A_ABST
    Figure CN120047579A_ABST
Patent Text Reader

Abstract

The invention discloses an image generation method, device and equipment and a computer readable storage medium, and belongs to the technical field of computers. The method comprises the steps that an original image and prompt information are obtained, the original image is an image to which a target text is to be added, and the prompt information is description information of the target text; obtaining an information feature vector used for representing the prompt information through a feature obtaining model; and according to the original image and the information feature vector, generating a target image, the target image being an image in which the target text is added in the original image. According to the method, the time required for image generation is saved, and the image generation efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Claims

1. An image generation method, characterized in that: The method comprises: Acquire an original image and prompt information, wherein the original image is an image to which a target text is to be added, and the prompt information is description information of the target text; Acquire an information feature vector for characterizing the prompt information through a feature acquisition model; A target image is generated according to the original image and the information feature vector, wherein the target image is an image in which the target text is added to the original image.

2. The method according to claim 1, characterized in that Before acquiring the information feature vector for characterizing the prompt information through the feature acquisition model, the method further includes: Acquire multiple reference images and annotation information of each reference image, wherein the multiple reference images are images of multiple reference texts, the annotation information of the reference images is description information of the reference texts included in the reference images, and the description information of the reference texts includes the reference texts and the font style of the reference texts; A first model is trained according to the multiple reference images and the annotation information of each reference image to obtain the feature acquisition model.

3. The method according to claim 2, characterized in that The step of training a first model according to the plurality of reference images and the annotation information of each reference image to obtain the feature acquisition model comprises: Inputting the plurality of reference images and the annotation information of each reference image into the first model to obtain an image feature vector of each reference image and an information feature vector of the annotation information of each reference image; Determining a first loss value of the first model according to the image feature vectors of the respective reference images and the information feature vectors of the annotation information of the respective reference images; Based on the first loss value, the first model is adjusted until the first loss value is lower than a first loss threshold, and the corresponding first model is determined to be the feature acquisition model.

4. The method according to claim 3, characterized in that: The determining, according to the image feature vectors of the respective reference images and the information feature vectors of the annotation information of the respective reference images, a first loss value of the first model comprises: Generate a similarity matrix according to the image feature vectors of the reference images and the information feature vectors of the annotation information of the reference images, wherein the similarity matrix includes a plurality of elements, and the element in the i-th row and the j-th column is used to represent the similarity between the annotation information of the i-th reference image and the j-th reference image, and both i and j are integers greater than 0 and not greater than the total number of the plurality of reference images; Normalizing each row of the similarity matrix to obtain a first similarity matrix; Normalizing each column of the similarity matrix to obtain a second similarity matrix; A first loss value of the first model is determined according to the first similarity matrix and the second similarity matrix.

5. The method according to claim 4, characterized in that The determining a first loss value of the first model according to the first similarity matrix and the second similarity matrix includes: Determine a first reference loss value according to the first similarity matrix, where the first reference loss value is used to characterize the loss from the image feature vector to the information feature vector; Determine a second reference loss value according to the second similarity matrix, where the second reference loss value is used to characterize the loss from the information feature vector to the image feature vector; Determine the average of the first reference loss value and the second reference loss value as the first loss value of the first model.

6. The method according to claim 3, characterized in that The first model includes an image encoder and a text encoder, the image encoder is used to obtain an image feature vector of the image, and the text encoder is used to obtain an information feature vector of the annotation information; The step of inputting the plurality of reference images and the annotation information of each reference image into the first model to obtain an image feature vector of each reference image and an information feature vector of the annotation information of each reference image includes: For a first reference image among the multiple reference images, input the first reference image into the image encoder to obtain an intermediate feature vector of the first reference image; Inputting the annotation information of the first reference image into the text encoder to obtain an intermediate feature vector of the annotation information of the first reference image; An image feature vector of the first reference image and an information feature vector of the annotation information of the first reference image are determined according to an intermediate feature vector of the first reference image and an intermediate feature vector of the annotation information of the first reference image.

7. The method according to claim 6, characterized in that The determining, according to the intermediate feature vector of the first reference image and the intermediate feature vector of the annotation information of the first reference image, the image feature vector of the first reference image and the information feature vector of the annotation information of the first reference image comprises: Determining an intermediate feature vector of the first reference image as an image feature vector of the first reference image; An intermediate feature vector of the annotation information of the first reference image is determined as an information feature vector of the annotation information of the first reference image.

8. The method according to claim 6, characterized in that The determining, according to the intermediate feature vector of the first reference image and the intermediate feature vector of the annotation information of the first reference image, the image feature vector of the first reference image and the information feature vector of the annotation information of the first reference image comprises: Determine a first query vector, a first key vector, and a first value vector based on the intermediate feature vector of the first reference image, and determine a second query vector, a second key vector, and a second value vector based on the intermediate feature vector of the annotation information of the first reference image; determining, based on the first query vector, the second key vector, and the second value vector, an information feature vector of the annotation information of the first reference image, so that the information feature vector of the annotation information of the first reference image incorporates features of the first reference image; Based on the second query vector, the first key vector and the first value vector, an image feature vector of the first reference image is determined, so that the image feature vector of the first reference image incorporates features of the annotation information of the first reference image.

9. The method according to any one of claims 2 to 8, characterized in that: The acquiring of multiple reference images comprises: For a first reference image among the multiple reference images, acquiring a text image and at least one style image corresponding to the first reference image, wherein the text image indicates content of a reference text in the first reference image, and the at least one style image indicates at least one font style; The text image and the at least one style image are input into a font generation model to obtain the first reference image, wherein the font style of the reference text included in the first reference image is a fusion style of the font styles indicated by the at least one style image.

10. The method according to any one of claims 1 to 8, characterized in that: The description information of the target text includes the target text and a target font style. The description information of the target text is description information in a natural language, and the font style of the target text in the target image is the target font style.

11. The method according to claim 10, characterized in that The description information of the target text also includes a target color, and the color of the target text in the target image is the target color.

12. The method according to any one of claims 1 to 8, characterized in that: The step of generating a target image according to the original image and the information feature comprises: The target image is generated according to the original image and the information feature vector through an image generation model.

13. The method according to claim 12, characterized in that The method further comprises: Acquire a mask image of the original image, wherein the mask image indicates a location in the original image where the target text is to be added; The step of generating the target image by using an image generation model according to the original image and the information feature vector comprises: The target image is generated by the image generation model according to the original image, the information feature vector and the mask image of the original image. The target image is an image with the target text added to the position indicated by the mask image in the original image.

14. An image generating device, characterized in that: The device comprises: An acquisition module, used to acquire an original image and prompt information, wherein the original image is an image to which a target text is to be added, and the prompt information is description information of the target text; The acquisition module is further used to acquire an information feature vector for representing the prompt information through a feature acquisition model; A generating module is used to generate a target image according to the original image and the information feature vector, wherein the target image is an image in which the target text is added to the original image.

15. A computer device, characterized in that: The computer device includes a processor and a memory, wherein at least one program code is stored in the memory, and the at least one program code is loaded and executed by the processor so that the computer device implements the image generation method according to any one of claims 1 to 13.

16. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores at least one program code, and the at least one program code is loaded and executed by a processor so that a computer implements the image generation method according to any one of claims 1 to 13.

17. A computer program product, characterized in that The computer program product stores at least one computer instruction, and the at least one computer instruction is loaded and executed by a processor so that a computer implements the image generation method according to any one of claims 1 to 13.