Image processing method and device, equipment and storage medium

By extracting and adjusting the gaze direction vector generated by the text image model, the problem of gaze direction randomness is solved, precise control of gaze direction is achieved, creation time cost is reduced and the determinism of generated images is improved.

CN120953404APending Publication Date: 2025-11-14SHANGHAI IQIYI NEW MEDIA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510967151.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-14
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

In existing technologies, the randomness of the gaze direction in images generated by large-scale image processing models is difficult to control, causing the actor's gaze direction to deviate from expectations, increasing the time cost and uncertainty of creation.

Method used

By acquiring the gaze feature information input by the user, and using pre-trained text-to-image model, feature recognition model and gaze analysis model, the gaze direction vectors of the target prompt word and the reference image are extracted respectively. The gaze angle deviation is calculated and adjusted to ensure that the gaze direction of the generated image is accurately matched with the target prompt word.

Benefits of technology

It reduces creation time and costs, improves the certainty and accuracy of generated images, and ensures that the direction of gaze accurately matches the semantic requirements of the target prompt.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120953404A_ABST
    Figure CN120953404A_ABST
Patent Text Reader

Abstract

The invention relates to an image processing method and device, equipment and a storage medium, and the method comprises the steps: obtaining a target prompt word inputted by a user, and the target prompt word comprises sight line feature information; the target cue word is input into a pre-trained text graph model, so that the pre-trained text graph model outputs a reference image, and the reference image comprises an eye region; analyzing sight line feature information in the target cue word by using a pre-trained feature recognition model to obtain a first sight line direction vector sequence; identifying an eye region in the reference image by using a pre-trained sight line analysis model to obtain a second sight line direction vector sequence; and adjusting the reference image according to the first line-of-sight direction vector sequence and the second line-of-sight direction vector sequence to obtain a target image. The method can ensure that the sight direction of the finally generated image is accurately matched with the semantic requirement of the target cue word, reduces the time cost of creation, and ensures the certainty of the generated image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to an image processing method, apparatus, device, and storage medium. Background Technology

[0002] In film and television production and advertising creativity, AI-powered costume fitting for actors and AI-generated virtual group photos of multiple actors are becoming increasingly popular. Film and television teams are using this technology to quickly preview the effects of different costumes and styles, and to plan promotional posters featuring actors appearing together across different time periods. To make the virtual images more realistic and story-driven, precisely controlling the actors' gaze is crucial, as it directly impacts the sense of interaction between characters and the overall visual impact.

[0003] In existing technologies, in order to control the actor's gaze direction, cue words including gaze information are input into a text-based image model to obtain an image that satisfies the gaze direction conditions.

[0004] However, randomness remains an unavoidable problem in the process of generating images using the Wenshengtu large model. Even with explicit gaze control cues, the direction of the actor's gaze in the final generated image may still deviate from the expected direction, requiring repeated adjustments to the cues or parameters, which increases the time cost and uncertainty of the creative process. Summary of the Invention

[0005] This application provides an image processing method, apparatus, device, and storage medium. After acquiring the output image of the text-based image model, semantic analysis is performed on a reference image and a target prompt word, extracting the gaze direction vectors of both. By calculating the gaze angle deviation between the two, the gaze direction in the reference image is adjusted to ensure that the gaze direction of the final generated image accurately matches the semantic requirements of the target prompt word, reducing the creation time cost and guaranteeing the determinism of the generated image.

[0006] In a first aspect, this application provides an image processing method, the method comprising:

[0007] Obtain target prompt words input by the user, wherein the target prompt words include gaze feature information;

[0008] The target prompt word is input into a pre-trained text image model so that the pre-trained text image model outputs a reference image, which includes the eye region;

[0009] Using a pre-trained feature recognition model, the gaze feature information in the target prompt is analyzed to obtain a first gaze direction vector sequence;

[0010] Using a pre-trained gaze analysis model, the eye region in the reference image is identified to obtain a second gaze direction vector sequence;

[0011] The reference image is adjusted based on the first line-of-sight vector sequence and the second line-of-sight vector sequence to obtain the target image.

[0012] Secondly, this application provides an image processing apparatus, the apparatus comprising:

[0013] The acquisition unit is used to acquire target prompt words input by the user, wherein the target prompt words include gaze feature information;

[0014] An input unit is used to input the target prompt word into a pre-trained text image model so that the pre-trained text image model outputs a reference image, the reference image including an eye region;

[0015] The analysis unit is used to analyze the gaze feature information in the target prompt word using a pre-trained feature recognition model to obtain a first gaze direction vector sequence;

[0016] The recognition unit is used to identify the eye region in the reference image using a pre-trained gaze analysis model to obtain a second gaze direction vector sequence.

[0017] The adjustment unit is used to adjust the reference image according to the first line-of-sight vector sequence and the second line-of-sight vector sequence to obtain the target image.

[0018] Thirdly, this application provides an image processing device, comprising: at least one communication interface; at least one bus connected to the at least one communication interface; at least one processor connected to the at least one bus; and at least one memory connected to the at least one bus, wherein the processor is configured to:

[0019] Obtain target prompt words input by the user, wherein the target prompt words include gaze feature information;

[0020] The target prompt word is input into a pre-trained text image model so that the pre-trained text image model outputs a reference image, which includes the eye region;

[0021] Using a pre-trained feature recognition model, the gaze feature information in the target prompt is analyzed to obtain a first gaze direction vector sequence;

[0022] Using a pre-trained gaze analysis model, the eye region in the reference image is identified to obtain a second gaze direction vector sequence;

[0023] The reference image is adjusted based on the first line-of-sight vector sequence and the second line-of-sight vector sequence to obtain the target image.

[0024] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described image processing method.

[0025] Compared with the prior art, the technical solution provided in this application has the following advantages: In this application embodiment, a target prompt word including gaze feature information is obtained from user input, and the target prompt word is input into a pre-trained text-based image model so that the pre-trained text-based image model outputs a reference image including the eye region. Then, a pre-trained gaze analysis model is used to identify the eye region in the reference image, obtaining a second gaze direction vector sequence. Based on the first and second gaze direction vector sequences, the reference image is adjusted to obtain a target image that meets the requirements. It can be seen that after obtaining the output image of the text-based image model, this application performs semantic analysis on the reference image and the target prompt word, extracting their respective gaze direction vectors. By calculating the gaze angle deviation between the two, the gaze direction in the reference image is adjusted to ensure that the gaze direction of the final generated image accurately matches the semantic requirements of the target prompt word, reducing the creation time cost and ensuring the determinism of the generated image. Attached Figure Description

[0026] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0027] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0028] One or more embodiments are illustrated by way of example with reference numerals in the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.

[0029] Figure 1 A schematic flowchart of an image processing method provided in an embodiment of this application;

[0030] Figure 2 A flowchart illustrating a method for determining a first line-of-sight direction vector sequence provided in an embodiment of this application;

[0031] Figure 3A flowchart illustrating a method for determining a second line-of-sight direction vector sequence provided in an embodiment of this application;

[0032] Figure 4 A flowchart illustrating a method for determining a second line-of-sight direction vector sequence provided in an embodiment of this application;

[0033] Figure 5 A schematic flowchart illustrating a target image determination method provided in an embodiment of this application;

[0034] Figure 6 A schematic flowchart illustrating a target image determination method provided in an embodiment of this application;

[0035] Figure 7 A flowchart illustrating a method for determining an included angle sequence provided in an embodiment of this application;

[0036] Figure 8 This is a schematic flowchart of an image processing apparatus provided in an embodiment of this application;

[0037] Figure 9 This is a schematic diagram of an image processing device provided in an embodiment of this application. Detailed Implementation

[0038] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0039] The following disclosure provides numerous different embodiments or examples for implementing various structures of the invention. To simplify the disclosure, specific examples of components and arrangements are described below. These are merely examples and are not intended to limit the scope of the invention. Furthermore, reference numerals and / or letters may be repeated in different examples. Such repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed.

[0040] In film and television production and advertising creativity, AI-powered costume fitting for actors and AI-generated virtual group photos of multiple actors are becoming increasingly popular. Film and television teams utilize this technology to quickly preview the effects of different costumes and styles, and to plan promotional posters featuring actors appearing together across different time periods. To make the virtual images more realistic and narrative-driven, precise control of the actors' gaze direction is crucial, directly impacting the interaction between characters and the overall visual impact. Current technologies control the actors' gaze direction by inputting cue words containing gaze information into a large-scale image model to obtain images where the gaze direction meets certain conditions. However, randomness remains a persistent problem in the image generation process of the large-scale image model. Even with explicit gaze control cue words, the actors' gaze direction in the final generated image may still deviate from expectations, requiring repeated adjustments to the cue words or parameters, increasing the time cost and uncertainty of the creative process.

[0041] To address the aforementioned issues, this application provides an image processing method. After acquiring the output image of the text-based image model, this method performs semantic analysis on a reference image and a target prompt word, extracting the gaze direction vectors of both. By calculating the gaze angle deviation between the two, the gaze direction in the reference image is adjusted to ensure that the gaze direction of the final generated image accurately matches the semantic requirements of the target prompt word, reducing creation time and guaranteeing the determinism of the generated image. Figure 1 As shown, the specific steps include:

[0042] Step 101: Obtain the target prompt words input by the user.

[0043] Among them, the target prompt words include gaze feature information, which is feature information describing the direction of gaze.

[0044] In this step, when the user enters a target suggestion word, the target suggestion word is obtained, and subsequent operations are performed based on the suggestion word.

[0045] Step 102: Input the target prompt word into the pre-trained text-to-image model so that the pre-trained text-to-image model outputs a reference image.

[0046] The reference image includes the eye region. A pre-trained text-to-image model is used to generate corresponding images based on prompts. This model is a machine-trained model, specifically a diffusion model. The training method of this model is similar to existing training methods, using a large-scale text-to-image dataset (such as LAION-5B) for end-to-end optimization. It improves semantic alignment ability through an iterative training strategy of noise addition and progressive denoising, which will not be elaborated here.

[0047] In this step, since the target prompt includes gaze feature information, the image generated based on the target prompt must include the eye region. Therefore, the target prompt is input into a pre-trained text-to-image model so that the pre-trained text-to-image model processes the target prompt and outputs a reference image including the eye region.

[0048] Step 103: Using a pre-trained feature recognition model, analyze the gaze feature information in the target prompt words to obtain the first gaze direction vector sequence.

[0049] In this step, the target prompt word can be directly input into a pre-trained feature recognition model, allowing the model to recognize the prompt word and output a first gaze direction vector sequence. Alternatively, gaze feature information and spatial location information for each object can be extracted from the target prompt word; the pre-trained feature recognition model can then be used to recognize the gaze feature information for each object, obtaining a first gaze direction vector for each object; and based on the spatial location information of each object, the first gaze direction vectors for each object can be sorted to obtain a first gaze direction vector sequence.

[0050] The feature recognition model, used to convert gaze feature information into gaze direction vectors, is a Large Language Model (LLM). A LLM is a deep learning-based Natural Language Processing (NLP) model with massive parameters (typically billions or even trillions), capable of understanding and generating human language. It is trained on large-scale text data (such as books, web pages, and dialogues) to learn the statistical patterns, semantic relationships, and contextual associations of language, thereby performing various language tasks such as text generation, translation, question answering, summarization, and code writing.

[0051] Step 104: Using a pre-trained gaze analysis model, the eye region in the reference image is identified to obtain a second gaze direction vector sequence.

[0052] In this step, the reference image can be directly input into a pre-trained gaze analysis model to enable the model to recognize the reference image and obtain a second gaze direction vector sequence. Alternatively, the eye region within each face detection box can be determined from the reference image; for each face detection box, the pre-trained gaze analysis model is used to identify the eye region within the face detection box to determine the corresponding second gaze direction vector; based on the spatial position information of each face detection box in the reference image, the second gaze direction vectors corresponding to each eye region are sorted to obtain a second gaze direction vector sequence.

[0053] The gaze analysis model is a pre-trained artificial intelligence model whose core function is to analyze images (or the eye region within an image) to identify the gaze direction of a person in the image and output the result as a "gaze direction vector." The model learns from a large number of image samples containing different gaze angles (especially eye feature data) to master the correlation between gaze direction and eye shape (such as pupil position, eyeball rotation angle, eyelid state, etc.) and facial posture. When a reference image (or an extracted eye region) is input, the model extracts key visual features, combines them with the learned patterns, calculates the corresponding gaze direction, and converts it into a vector form (i.e., the "gaze direction vector") for easy subsequent processing. This vector intuitively reflects the horizontal, vertical, and other directional information of the gaze.

[0054] Step 105: Adjust the reference image according to the first line-of-sight direction vector sequence and the second line-of-sight direction vector sequence to obtain the target image.

[0055] In this step, the first and second line-of-sight vectors corresponding to the same position in the first and second line-of-sight vector sequences are obtained. Then, the cosine or sine value between them is calculated, and the angle between them is calculated based on this value. When the angle is within a preset range, the reference image is not adjusted, and the target image is obtained. If the angle is within the preset range, the line-of-sight direction of the corresponding object is adjusted according to this angle until the angle is within the preset range, thus obtaining the target image.

[0056] In this embodiment, a target prompt word including gaze feature information is obtained from user input and input into a pre-trained text-to-image model, causing the pre-trained text-to-image model to output a reference image including the eye region. Then, a pre-trained gaze analysis model is used to identify the eye region in the reference image, obtaining a second gaze direction vector sequence. Based on the first and second gaze direction vector sequences, the reference image is adjusted to obtain a target image that meets the requirements. Therefore, after obtaining the output image of the text-to-image model, this application performs semantic analysis on the reference image and the target prompt word, extracting the gaze direction vectors of both. By calculating the gaze angle deviation between the two, the gaze direction in the reference image is adjusted to ensure that the gaze direction of the final generated image accurately matches the semantic requirements of the target prompt word, reducing the creation time cost and ensuring the determinism of the generated image.

[0057] In this embodiment, the target prompt may involve the gaze feature information of at least one object. Therefore, a pre-trained feature recognition model is needed to identify the gaze feature information of each object, obtaining a first gaze direction vector corresponding to each object. Simultaneously, the spatial location information of each object is extracted from the target prompt. Finally, based on the spatial location information of each object, the first gaze direction vectors corresponding to each object are sorted to obtain a first gaze direction vector sequence. Therefore, this embodiment provides a method for determining a first gaze direction vector sequence, as follows: Figure 2 As shown, the specific steps include:

[0058] Step 201: Extract the gaze feature information and spatial location information of each object from the target prompt words.

[0059] The target prompt includes at least one object's gaze characteristics and spatial location information. The gaze characteristics are used to define the object's gaze orientation in the image (e.g., "looking up" or "looking at an object on the right"), while the spatial location information is used to specify the object's layout and positioning in the image (e.g., "located in the lower left corner of the image" or "the upper right third area").

[0060] In this step, a pre-trained feature extraction model can be set up. The target prompt words are input into the pre-trained feature extraction model, which then outputs the gaze feature information and spatial location information of each object. The feature extraction model can be a natural language recognition model. Alternatively, the target prompt words can be identified from scratch. Based on an object keyword library, the first object is determined from the target prompt words. Then, based on the gaze keyword library and the location keyword library, the gaze feature information and spatial location information of the first object are extracted from the target prompt words. This process is repeated until all target prompt words have been identified, obtaining the gaze feature information and spatial location information of all objects in the target prompt words.

[0061] Step 202: Using a pre-trained feature recognition model, the gaze feature information of each object is identified to obtain the first gaze direction vector corresponding to each object.

[0062] In this step, for each object, the object's gaze feature information is input into a pre-trained feature recognition model to obtain the gaze direction vector corresponding to the object, and this vector is used as the first gaze direction vector corresponding to the object.

[0063] Step 203: Based on the spatial location information of each object, sort the first line of sight direction vector corresponding to each object to obtain a sequence of first line of sight direction vectors.

[0064] In this step, based on the spatial location information of each object, the first line of sight vectors corresponding to these objects are sorted according to a preset sorting strategy to obtain a sequence of first line of sight vectors.

[0065] For example, the target cue includes three objects: the first object is located on the left, the second in the middle, and the third on the right. The default strategy is to arrange them from left to right. Therefore, the first gaze direction vector of the first object is placed in the first position of the first gaze direction vector sequence, the first gaze direction vector of the second object is placed in the second position, and the first gaze direction vector of the third object is placed in the third position. This results in the first gaze direction vector sequence.

[0066] In this embodiment, when performing face recognition on a reference image, face detection boxes are marked on the faces in the reference image. Each face detection box includes a face, which typically includes the eye region. Therefore, this embodiment can identify the eye region within each face detection box, determine its corresponding second gaze direction vector, and thus obtain the second gaze direction vector corresponding to each face detection box. Then, based on the spatial position information of each face detection box in the reference image, the second gaze direction vectors corresponding to each eye region are sorted to obtain a second gaze direction vector sequence. Therefore, this embodiment provides a method for determining a second gaze direction vector sequence, as follows: Figure 3 As shown, the specific steps include:

[0067] Step 301: Determine the eye region within each face detection box from the reference image.

[0068] In this step, MTCNN (Multi-Task Convolutional Neural Network) is used to identify face detection boxes in the reference image. For each face detection box, key points are located in the region where the face detection box is located, and the coordinates of the two eye regions (left eye corner, right eye corner, and eyelid contour key points) are extracted. Based on these key points, the eye region within the face detection box is determined.

[0069] Step 302: For each face detection box, use a pre-trained gaze analysis model to identify the eye region within the face detection box and determine the second gaze direction vector corresponding to the face detection box.

[0070] In this step, for the eye region within each face detection box, a pre-trained gaze analysis model is input so that the pre-trained gaze analysis model can identify the eye region and obtain the second gaze direction vector corresponding to the eye region, which is the second gaze direction vector of the corresponding face detection box.

[0071] Step 303: Based on the spatial position information of each face detection box in the reference image, sort the second gaze direction vector corresponding to each eye region to obtain a second gaze direction vector sequence.

[0072] The spatial location information is either the center coordinates of the face detection box or the coordinates of the top left corner of the face detection box; it is not limited here.

[0073] In this step, based on the spatial position information of each face detection box in the reference image, the second gaze direction vector corresponding to each eye region is sorted according to a preset sorting strategy to obtain a second gaze direction vector sequence.

[0074] The preset sorting strategy mentioned above is the same as the preset sorting strategy in step 203. For example, if the preset sorting strategy in step 203 is to arrange from left to right, then the preset sorting strategy here is also to arrange from left to right.

[0075] For example, the reference image includes three face detection boxes. The spatial location information of the first face detection box is (x1, y1), the spatial location information of the second face detection box is (x2, y2), and the spatial location information of the third face detection box is (x3, y3), where x1, x2, and x3 are distributed from left to right. The default strategy is to arrange them sequentially from left to right. Therefore, the second gaze direction vector of the first face detection box is placed in the first position of the second gaze direction vector sequence, the second gaze direction vector of the second face detection box is placed in the second position, and the second gaze direction vector of the third face detection box is placed in the third position. In this way, the second gaze direction vector sequence is obtained.

[0076] In this embodiment, since the eye region often contains two eyes, a gaze determination module in a pre-trained gaze analysis model can be used to determine the third gaze direction vector corresponding to each eye. Then, a processing module is used to obtain the second gaze direction vector corresponding to the face detection box using the third gaze direction vectors corresponding to the two eyes. Therefore, this embodiment provides a method for determining the second gaze direction vector, as follows: Figure 4 As shown, the specific steps include:

[0077] Step 401: Determine the third gaze direction vector corresponding to each eye within the face detection box.

[0078] In this step, each eye within the face detection box is input into the gaze determination module, along with the third gaze direction vector corresponding to each eye.

[0079] The gaze determination module is a machine learning module that can analyze the eye region to obtain its corresponding gaze direction vector. For example, the gaze determination module can be the 3DgazeNet algorithm module, or other models. It is not limited here.

[0080] Step 402: Determine the second gaze direction vector corresponding to the face detection box based on the third gaze direction vector corresponding to each eye within the face detection box.

[0081] In this step, the average vector of the third gaze direction vector corresponding to each of the two eyes can be calculated, and the average vector can be normalized to obtain the normalized vector. This vector is then used as the second gaze direction vector corresponding to the face detection box.

[0082] In this embodiment, if the gaze direction of an object in the reference image differs significantly from the gaze direction described in the target prompt, then the gaze direction of the object in the reference image needs to be adjusted. If the gaze direction of an object in the reference image is not significantly different from the gaze direction described in the target prompt, then no adjustment of the gaze direction of the object in the reference image is required. Therefore, this embodiment provides a target image determination method, as follows: Figure 5 As shown, the specific steps include:

[0083] Step 501: Determine the included angle sequence corresponding to the reference image based on the first line-of-sight direction vector sequence and the second line-of-sight direction vector sequence.

[0084] In this step, the first and second line-of-sight direction vectors corresponding to the first sequence position are obtained from the first and second line-of-sight direction vector sequences, and their cosine values ​​are calculated. The included angle is then determined based on this cosine value. An angle sequence is obtained using this method.

[0085] according to Calculate the cosine value of the two vectors, where V1 is the first line of sight vector, V2 is the second line of sight vector, and cosθ is the cosine value between the first line of sight vector and the second line of sight vector.

[0086] Step 502: Detect whether the target angle exists in the angle sequence.

[0087] Specifically, if the target angle is outside the preset range, it indicates that the viewing direction of the corresponding object in the reference image is not significantly different from the viewing direction described in the target prompt. Conversely, if the angle is outside the preset range, it indicates that the viewing direction of the corresponding object in the reference image differs significantly from the viewing direction described in the target prompt. The preset range is generally from -10 degrees to 10 degrees.

[0088] In this step, it is detected whether each angle in the angle sequence is within a preset range. When an angle is not within the preset range, the angle is determined as the target angle.

[0089] Step 503: When there is no target angle in the angle sequence, the reference image is used as the target image.

[0090] In this step, if there is no target angle in the angle sequence, it means that the gaze direction of all objects in the reference image is not much different from the gaze direction described in the target prompt, and the reference image can be used as the final target image.

[0091] Step 504: When a target angle exists in the angle sequence, adjust the reference image according to the target angle to obtain the target image.

[0092] In this step, when a target angle exists in the angle sequence, it means that the gaze direction of the object corresponding to the target angle in the reference image differs too much from the gaze direction described in the target prompt. The reference image needs to be adjusted to obtain the target image.

[0093] In this embodiment, when a target angle exists in the angle sequence, the eye region corresponding to the target angle needs to be adjusted so that the corresponding angle is within a preset range. Therefore, this embodiment provides a target image determination method, as follows: Figure 6 As shown, the specific steps include:

[0094] Step 601: Determine the target eye region to be adjusted from the reference image based on the sequence position of the target angle in the first line of sight vector sequence.

[0095] In this step, based on the target angle's position in the first line of sight vector sequence and the preset sorting strategy, the sorting position of the face detection box of the target eye region in the reference image is determined, and the eye region where the sorting position is located is determined as the target eye region.

[0096] Step 602: Using a pre-trained gaze adjustment model, the target eye region in the reference image is adjusted according to the target angle so that the angle corresponding to the adjusted target eye region is within a preset range, thereby obtaining the target image.

[0097] In this step, the target eye region and the target angle are input into a pre-trained gaze adjustment model. This model adjusts the target eye region to change its corresponding gaze direction, resulting in an adjusted target eye region. Then, the corresponding second gaze direction vector is recalculated, and the angle is recalculated based on this vector. If the angle is outside a preset range, the model is readjusted using the above method until the angle falls within the preset range, at which point the adjustment stops, and the target image is obtained.

[0098] Furthermore, the gaze adjustment model is the LivePortrait algorithm model. Since the LivePortrait algorithm model may not be able to directly recognize the target angle, the target angle can be converted into a value that the LivePortrait algorithm model can understand. Then, the value and the target eye area are input into the LivePortrait algorithm model to obtain the redrawn image, that is, the adjusted target eye area.

[0099] In this embodiment, the first and second line-of-sight direction vectors corresponding to the same sequence position are found in the first and second line-of-sight direction vector sequences, and their cosine values ​​are calculated. Then, the included angle corresponding to the sequence position is determined based on the cosine value. An included angle sequence is obtained based on the included angle corresponding to each sequence position. Therefore, this embodiment provides a method for determining included angle sequences, as follows: Figure 7 As shown, the specific steps include:

[0100] Step 701: Determine the first line of sight direction vector and the second line of sight direction vector corresponding to each sequence position from the first line of sight direction vector sequence and the second line of sight direction vector sequence.

[0101] In this step, for each sequence position, the first line of sight direction vector and the second line of sight direction vector corresponding to that sequence position are extracted from the first line of sight direction vector sequence and the second line of sight direction vector sequence.

[0102] Step 702: Determine the included angle corresponding to each sequence position based on the first line of sight direction vector and the second line of sight direction vector corresponding to each sequence position.

[0103] In this step, for each sequence position, the first and second line-of-sight direction vectors corresponding to that position are determined, and their cosine values ​​are calculated. The corresponding angle is then determined based on these cosine values. The angle corresponding to each sequence position is obtained using the above method.

[0104] Step 703: Sort all the included angles according to the sequence position corresponding to each included angle to obtain the included angle sequence corresponding to the reference image.

[0105] like Figure 8 As shown, this application provides an image processing apparatus, which corresponds to the method embodiment, and specifically includes:

[0106] The acquisition unit 801 is used to acquire the target prompt word input by the user, wherein the target prompt word includes gaze feature information;

[0107] The input unit 802 is used to input the target prompt word into a pre-trained text image model so that the pre-trained text image model outputs a reference image, the reference image including an eye region;

[0108] Analysis unit 803 is used to analyze the gaze feature information in the target prompt word using a pre-trained feature recognition model to obtain a first gaze direction vector sequence;

[0109] The recognition unit 804 is used to identify the eye region in the reference image using a pre-trained gaze analysis model to obtain a second gaze direction vector sequence.

[0110] The adjustment unit 805 is used to adjust the reference image according to the first line-of-sight vector sequence and the second line-of-sight vector sequence to obtain the target image.

[0111] Optionally, the analysis unit 803 is used for:

[0112] Extract the gaze feature information and spatial location information of each object from the target prompt words;

[0113] Using a pre-trained feature recognition model, the gaze feature information of each object is identified to obtain the first gaze direction vector corresponding to each object;

[0114] Based on the spatial location information of each object, the first line of sight direction vector corresponding to each object is sorted to obtain the first line of sight direction vector sequence.

[0115] Optionally, the identification unit 804 is used for:

[0116] The eye region within each face detection box is determined from the reference image;

[0117] For each face detection bounding box, a pre-trained gaze analysis model is used to identify the eye region within the face detection bounding box and determine the second gaze direction vector corresponding to the face detection bounding box.

[0118] Based on the spatial position information of each face detection box in the reference image, the second gaze direction vector corresponding to each eye region is sorted to obtain the second gaze direction vector sequence.

[0119] Optionally, the identification unit 804 is used for:

[0120] Determine the third gaze direction vector corresponding to each eye within the face detection box;

[0121] The second gaze direction vector corresponding to the face detection box is determined based on the third gaze direction vector corresponding to each eye within the face detection box.

[0122] Optionally, the adjustment unit 805 is used for:

[0123] Based on the first line-of-sight vector sequence and the second line-of-sight vector sequence, determine the included angle sequence corresponding to the reference image;

[0124] Detect whether a target angle exists in the angle sequence;

[0125] When the target angle is not present in the angle sequence, the reference image is used as the target image;

[0126] When a target angle exists in the angle sequence, the reference image is adjusted according to the target angle to obtain the target image.

[0127] Optionally, the adjustment unit 805 is used for:

[0128] Based on the sequence position of the target angle in the first line of sight vector sequence, the target eye region to be adjusted is determined from the reference image;

[0129] Using a pre-trained gaze adjustment model, the target eye region in the reference image is adjusted according to the target angle so that the angle corresponding to the adjusted target eye region is within a preset range, thereby obtaining the target image.

[0130] Optionally, the adjustment unit 805 is used for:

[0131] From the first line of sight direction vector sequence and the second line of sight direction vector sequence, determine the first line of sight direction vector and the second line of sight direction vector corresponding to each sequence position;

[0132] Determine the included angle corresponding to each sequence position based on the first line of sight direction vector and the second line of sight direction vector corresponding to each sequence position;

[0133] Based on the sequence position corresponding to each included angle, all included angles are sorted to obtain the included angle sequence corresponding to the reference image.

[0134] like Figure 9As shown in the figure, this application provides an image processing device, including a processor 901, a communication interface 902, a memory 903, and a communication bus 904, wherein the processor 901, the communication interface 902, and the memory 903 communicate with each other through the communication bus 904.

[0135] Memory 903 is used to store computer programs;

[0136] In one embodiment of this application, when the processor 901 executes a program stored in the memory 903, it implements the image processing method provided in any of the foregoing method embodiments, including:

[0137] Obtain target prompt words input by the user, wherein the target prompt words include gaze feature information;

[0138] The target prompt word is input into a pre-trained text image model so that the pre-trained text image model outputs a reference image, which includes the eye region;

[0139] Using a pre-trained feature recognition model, the gaze feature information in the target prompt is analyzed to obtain a first gaze direction vector sequence;

[0140] Using a pre-trained gaze analysis model, the eye region in the reference image is identified to obtain a second gaze direction vector sequence;

[0141] The reference image is adjusted based on the first line-of-sight vector sequence and the second line-of-sight vector sequence to obtain the target image.

[0142] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the image processing method provided in any of the foregoing method embodiments.

[0143] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0144] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented using software plus a general-purpose hardware platform, or of course, using hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0145] It should be understood that the terminology used herein is for the purpose of describing particular exemplary embodiments only and is not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “described” as used herein may also include the plural forms. The terms “comprising,” “including,” “containing,” and “having” are inclusive and therefore indicate the presence of the stated features, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, elements, components, and / or combinations thereof. The method steps, processes, and operations described herein are not construed as requiring them to be performed in a particular order described or illustrated unless the order of performance is explicitly indicated. It should also be understood that additional or alternative steps may be used.

[0146] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. An image processing method, characterized in that, The method includes: Obtain target prompt words input by the user, wherein the target prompt words include gaze feature information; The target prompt word is input into a pre-trained text image model so that the pre-trained text image model outputs a reference image, which includes the eye region; Using a pre-trained feature recognition model, the gaze feature information in the target prompt is analyzed to obtain a first gaze direction vector sequence; Using a pre-trained gaze analysis model, the eye region in the reference image is identified to obtain a second gaze direction vector sequence; The reference image is adjusted based on the first line-of-sight vector sequence and the second line-of-sight vector sequence to obtain the target image.

2. The method according to claim 1, characterized in that, The step of using a pre-trained feature recognition model to analyze the gaze feature information in the target prompt word to obtain the first gaze direction vector sequence includes: Extract the gaze feature information and spatial location information of each object from the target prompt words; Using a pre-trained feature recognition model, the gaze feature information of each object is identified to obtain the first gaze direction vector corresponding to each object; Based on the spatial location information of each object, the first line of sight direction vector corresponding to each object is sorted to obtain the first line of sight direction vector sequence.

3. The method according to claim 1, characterized in that, The step of using a pre-trained gaze analysis model to identify human eyes in the reference image and obtain a second gaze direction vector sequence includes: The eye region within each face detection box is determined from the reference image; For each face detection bounding box, a pre-trained gaze analysis model is used to identify the eye region within the face detection bounding box and determine the second gaze direction vector corresponding to the face detection bounding box. Based on the spatial position information of each face detection box in the reference image, the second gaze direction vector corresponding to each eye region is sorted to obtain the second gaze direction vector sequence.

4. The method according to claim 3, characterized in that, The step of using a pre-trained gaze analysis model to identify the eye region within the face detection box and determine the second gaze direction vector corresponding to the face detection box includes: Determine the third gaze direction vector corresponding to each eye within the face detection box; The second gaze direction vector corresponding to the face detection box is determined based on the third gaze direction vector corresponding to each eye within the face detection box.

5. The method according to claim 4, characterized in that, The step of adjusting the reference image based on the first gaze direction vector sequence and the second gaze direction vector sequence to obtain the target image includes: Based on the first line-of-sight vector sequence and the second line-of-sight vector sequence, determine the included angle sequence corresponding to the reference image; Detect whether a target angle exists in the angle sequence; When the target angle is not present in the angle sequence, the reference image is used as the target image; When a target angle exists in the angle sequence, the reference image is adjusted according to the target angle to obtain the target image.

6. The method according to claim 5, characterized in that, The step of adjusting the reference image according to the target angle to obtain the target image includes: Based on the sequence position of the target angle in the first line of sight vector sequence, the target eye region to be adjusted is determined from the reference image; Using a pre-trained gaze adjustment model, the target eye region in the reference image is adjusted according to the target angle so that the angle corresponding to the adjusted target eye region is within a preset range, thereby obtaining the target image.

7. The method according to claim 5, characterized in that, The step of determining the included angle sequence corresponding to the reference image based on the first line-of-sight vector sequence and the second line-of-sight vector sequence includes: From the first line of sight direction vector sequence and the second line of sight direction vector sequence, determine the first line of sight direction vector and the second line of sight direction vector corresponding to each sequence position; Determine the included angle corresponding to each sequence position based on the first line of sight direction vector and the second line of sight direction vector corresponding to each sequence position; Based on the sequence position corresponding to each included angle, all included angles are sorted to obtain the included angle sequence corresponding to the reference image.

8. An image processing apparatus, characterized in that, The device includes: The acquisition unit is used to acquire target prompt words input by the user, wherein the target prompt words include gaze feature information; An input unit is used to input the target prompt word into a pre-trained text image model so that the pre-trained text image model outputs a reference image, the reference image including an eye region; The analysis unit is used to analyze the gaze feature information in the target prompt word using a pre-trained feature recognition model to obtain a first gaze direction vector sequence; The recognition unit is used to identify the eye region in the reference image using a pre-trained gaze analysis model to obtain a second gaze direction vector sequence. The adjustment unit is used to adjust the reference image according to the first line-of-sight vector sequence and the second line-of-sight vector sequence to obtain the target image.

9. An image processing device, characterized in that, include: At least one communication interface; At least one bus connected to the at least one communication interface; at least one processor connected to the at least one bus; At least one memory connected to the at least one bus, wherein the processor is configured to: Obtain target prompt words input by the user, wherein the target prompt words include gaze feature information; The target prompt word is input into a pre-trained text image model so that the pre-trained text image model outputs a reference image, which includes the eye region; Using a pre-trained feature recognition model, the gaze feature information in the target prompt is analyzed to obtain a first gaze direction vector sequence; Using a pre-trained gaze analysis model, the eye region in the reference image is identified to obtain a second gaze direction vector sequence; The reference image is adjusted based on the first line-of-sight vector sequence and the second line-of-sight vector sequence to obtain the target image.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the image processing method according to any one of claims 1 to 7.