Face information processing method, device and computer readable storage medium

Through the face information processing model based on Transformer, the problem that face-related tasks in the existing technology cannot be handled uniformly is solved, efficient and unified face information processing is achieved, and model performance is improved.

CN115457626BActive Publication Date: 2025-06-06SHENZHEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211022252.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-23
Publication Date
2025-06-06
Estimated Expiration
2042-08-23

AI Technical Summary

Technical Problem

In existing computer vision technology, downstream tasks related to face cannot be completed by a model in a unified training, resulting in high costs and difficult to improve model performance.

Method used

The face information processing model based on Transformer is adopted, and the face information processing model is obtained, and the face image and attribute text are inputted to the preset model to generate the target face image. On this basis, the face image and attribute text are processed to complete multitasking.

Benefits of technology

It realizes unified processing of face-related tasks, improves model performance and efficiency, and reduces costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115457626B_ABST
    Figure CN115457626B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, device and computer-readable storage medium for processing face information, wherein the method comprises: obtaining a face synthesis text, and then inputting the face synthesis text into a preset model to obtain a target face image corresponding to the face synthesis text, wherein the preset model is a trained Transformer-based face information processing model. The present invention can output a target face image through a preset model based on the face synthesis text, so that face-related tasks are completed by a preset model, thereby improving model performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of face information processing technology, and in particular to a face information processing method, device and computer-readable storage medium. Background Art

[0002] As an important field of computer vision, face has generated a large number of downstream tasks, such as face generation, age prediction, expression classification, face attributes, etc., which have broad research and application value.

[0003] In computer vision, tasks such as face generation, age prediction, expression classification, and face attributes are usually modeled as separate tasks, requiring separate model design, data collection, model training, and deployment. This model is costly and it is difficult to leverage the associations between tasks to improve model performance.

[0004] The above contents are only used to assist in understanding the technical solution of the present invention and do not constitute an admission that the above contents are prior art. Summary of the invention

[0005] The main purpose of the present invention is to provide a facial information processing method, device and computer-readable storage medium, aiming to solve the technical problem that downstream tasks related to faces in existing computer vision cannot be completed by unified training of a model.

[0006] To achieve the above object, the present invention provides a method for processing face information, the method comprising the following steps:

[0007] Get face synthesis text;

[0008] The face synthesis text is input into a preset model to obtain a target face image corresponding to the face synthesis text, wherein the preset model is a trained Transformer-based face information processing model.

[0009] Furthermore, the preset model includes a data input layer, a deep learning layer and a data output layer connected in sequence, the data input layer includes a first word segmenter, the deep learning layer includes an Embedding layer and a Transformer module, and the step of inputting the face synthesis text into the preset model to obtain a target face image corresponding to the face synthesis text includes:

[0010] The first word segmenter converts the face synthesis text into a first text feature sequence, and inputs the first text feature sequence into a deep learning layer;

[0011] Determining a first target vector corresponding to the first text feature sequence in the Embedding layer;

[0012] The first target vector is input into a Transformer module to obtain a first output feature output by the Transformer module, and the first output feature is input into a data output layer to obtain a target face image.

[0013] Furthermore, the step of determining the first target vector corresponding to the first text feature sequence in the Embedding layer includes:

[0014] In the Embedding layer, the first text feature sequence is converted into a first feature vector, position encoding is performed on the first feature vector to determine a first position encoding vector, paragraph encoding is performed on the first feature vector to determine a first paragraph encoding vector, and character encoding is performed on the first feature vector to determine a first character encoding vector;

[0015] The first position encoding vector, the first paragraph encoding vector, the first character encoding vector, and the first feature vector are added together to determine a first target vector.

[0016] Further, the Transformer module includes at least two Transformer layers connected in sequence, the last Transformer layer in the Transformer module includes a softmax layer, the data output layer includes a VQ-GAN module, and the step of inputting the first target vector into the Transformer module to obtain a first output feature output by the Transformer module, and inputting the first output feature into the data output layer to obtain a target face image includes:

[0017] Inputting the first target vector into the Transformer module;

[0018] Obtaining the probability of each first candidate vector through the softmax layer, obtaining a preset number of first target candidate vectors with the highest probability from the first candidate vectors, and randomly obtaining second target candidate vectors from the first target candidate vectors, and obtaining first output features corresponding to the second target candidate vectors;

[0019] The first output feature is input into the data output layer, and the first output feature is converted through the VQ-GAN module to obtain a target face image.

[0020] Furthermore, the face information processing method further includes:

[0021] Get face image and face attribute text;

[0022] The face image and the face attribute text are input into a preset model to obtain the target attribute text.

[0023] Furthermore, the preset model includes a data input layer, a deep learning layer and a data output layer connected in sequence, the data input layer includes a first word segmenter and a VQ-GAN module, the deep learning layer includes an Embedding layer and a Transformer module, and the step of inputting the face image and the face attribute text into the preset model to obtain a target face image corresponding to the face synthesis text includes:

[0024] The first word segmenter converts the face attribute text into a second text feature sequence, the VQ-GAN module converts the face image into an image feature sequence, and connects the second text feature sequence and the image feature sequence to determine a target feature sequence;

[0025] Determining a second target vector corresponding to the target feature sequence in the Embedding layer;

[0026] The second target vector is input into the Transformer module to obtain the second output feature output by the Transformer module, and the second output feature is input into the data output layer to obtain the target attribute text.

[0027] Furthermore, the step of determining the second target vector corresponding to the target feature sequence in the Embedding layer includes:

[0028] In the Embedding layer, the target feature sequence is converted into a second feature vector, position encoding is performed on the second feature vector to determine a second position encoding vector, paragraph encoding is performed on the second feature vector to determine a second paragraph encoding vector, and character encoding is performed on the second feature vector to determine a second character encoding vector;

[0029] The second position encoding vector, the second paragraph encoding vector, the second character encoding vector, and the second feature vector are added together to determine a second target vector.

[0030] Further, the Transformer module includes at least two Transformer layers connected in sequence, the last Transformer layer in the Transformer module includes a softmax layer, the data output layer includes a second word segmenter, and the step of inputting the second target vector into the Transformer module to obtain a second output feature output by the Transformer module, and inputting the second output feature into the data output layer to obtain the target attribute text includes:

[0031] Inputting the second target vector into the Transformer module;

[0032] Obtaining the probability of each second candidate vector through the softmax layer, obtaining the third target candidate vector with the highest probability among the second candidate vectors, and obtaining the second output feature corresponding to the third target candidate vector;

[0033] Inputting the second output feature into the data output layer;

[0034] The second tokenizer converts the second output feature to obtain target attribute text.

[0035] In addition, to achieve the above-mentioned purpose, the present invention also provides a face information processing device, which includes: a memory, a processor, and a face information processing program stored on the memory and executable on the processor, and the face information processing program implements the steps of the aforementioned face information processing method when executed by the processor.

[0036] In addition, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, on which a face information processing program is stored, and when the face information processing program is executed by a processor, the steps of the aforementioned face information processing method are implemented.

[0037] The present invention obtains a face synthesis text and then inputs the face synthesis text into a preset model to obtain a target face image corresponding to the face synthesis text, wherein the preset model is a trained Transformer-based face information processing model. The present invention can output a target face image through a preset model based on the face synthesis text, so that face-related tasks are completed by a preset model, thereby improving model performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 It is a structural schematic diagram of a face information processing device in a hardware operating environment involved in an embodiment of the present invention;

[0039] Figure 2Schematic diagram of the process of the first embodiment of the face information processing method of the present invention.

[0040] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0041] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.

[0042] like Figure 1 As shown, Figure 1 It is a structural diagram of a face information processing device in a hardware operating environment involved in an embodiment of the present invention.

[0043] The facial information processing device of the embodiment of the present invention can be a PC, or it can be a smart phone, a tablet computer, an e-book reader, an MP3 (Moving Picture Experts Group Audio Layer III) player, an MP4 (Moving Picture Experts Group Audio Layer IV) player, a portable computer, or other portable terminal devices with display function.

[0044] like Figure 1 As shown, the face information processing device may include: a processor 1001, such as a CPU, a network interface 1004, a user interface 1003, a memory 1005, and a communication bus 1002. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display), an input unit such as a keyboard (Keyboard), and the user interface 1003 may also include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory, or a stable memory (non-volatile memory), such as a disk memory. The memory 1005 may also be a storage device independent of the aforementioned processor 1001.

[0045] Optionally, the face information processing device may also include a camera, an RF (Radio Frequency) circuit, a sensor, an audio circuit, a WiFi module, and the like. Among them, sensors such as light sensors, motion sensors, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor, wherein the ambient light sensor may adjust the brightness of the display screen according to the brightness of the ambient light, and the proximity sensor may turn off the display screen and / or backlight when the face information processing device is moved to the ear. As a type of motion sensor, the gravity acceleration sensor can detect the magnitude of acceleration in each direction (generally three axes), and can detect the magnitude and direction of gravity when stationary, which can be used for applications that recognize the posture of the face information processing device (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; of course, the face information processing device may also be equipped with other sensors such as gyroscopes, barometers, hygrometers, thermometers, infrared sensors, etc., which will not be repeated here.

[0046] Those skilled in the art will understand that Figure 1 The terminal structure shown in the figure does not constitute a limitation on the terminal, and may include more or less components than shown in the figure, or combine certain components, or arrange the components differently.

[0047] like Figure 1 As shown, the memory 1005 as a computer storage medium may include an operating system, a network communication module, a user interface module, and a face information processing program.

[0048] exist Figure 1 In the terminal shown, the network interface 1004 is mainly used to connect to the background server and communicate data with the background server; the user interface 1003 is mainly used to connect to the client (user end) and communicate data with the client; and the processor 1001 can be used to call the face information processing program stored in the memory 1005.

[0049] In this embodiment, the face information processing device includes: a memory 1005, a processor 1001, and a face information processing program stored on the memory 1005 and executable on the processor 1001, wherein the processor 1001 calls the face information processing program stored in the memory 1005 and executes the steps of the face information processing method in the following embodiments.

[0050] The present invention also provides a method for processing face information, referring to Figure 2 , Figure 2 Schematic diagram of the process of the first embodiment of the method of the present invention.

[0051] In this embodiment, the face information processing method includes the following steps:

[0052] Step S101, obtaining face synthesis text;

[0053] In this embodiment, a face synthesis text is obtained, wherein the content of the face synthesis text may be a face generation field and a description of the face information, for example: "Generate image" is the face generation field, and "a woman with long black hair and big eyes" is the description of the face information.

[0054] Step S102, inputting the face synthesis text into a preset model to obtain a target face image corresponding to the face synthesis text, wherein the preset model is a trained Transformer-based face information processing model.

[0055] It should also be noted that the trained Transformer-based face information processing model is trained using labeled face images, face description information corresponding to the face images, and labeling results; and the Transformer-based face information processing model includes a sequentially connected data input layer, a deep learning layer, and a data output layer. The deep learning layer includes an Embedding layer and at least two sequentially connected Transformer layers, each of which includes a self-attention layer, and the self-attention mechanism used by the self-attention layer is:

[0056]

[0057] In the above formula, Q, K, V are query matrix, key matrix and value matrix respectively, M is the attention mask matrix, d k is the dimension of the input. In order to improve the performance of the model, a multi-head attention mechanism is adopted, that is, multiple W Q , W K , W V The matrix generates multiple query matrices, key matrices, and value matrices, and then outputs multiple eigenvalues ​​according to the above formula, and then concatenates the multiple eigenvalues ​​and multiplies them by a matrix parameter to output the final feature. The segmented mask enables the input part of the model to use bidirectional attention, and the output part is from left to right, so as to obtain better generation capabilities. The value of M can be 0 or negative infinity. When M is 0, the attention mechanism is turned on, and when M is negative infinity, the attention mechanism is turned off.

[0058] In this embodiment, the face synthesis text is input into the trained Transformer-based face information processing model to obtain a target face image corresponding to the face synthesis text. For example, the content of the input face synthesis text is "generate an image and a woman with long black hair and big eyes". After being processed by the trained Transformer-based face information processing model, the target face image corresponding to the face synthesis text is generated, and the target face image is a woman with long black hair and big eyes.

[0059] The facial information processing method proposed in this embodiment obtains a synthetic facial text and then inputs the synthetic facial text into a preset model to obtain a target facial image corresponding to the synthetic facial text, wherein the preset model is a trained Transformer-based facial information processing model, which can output the target facial image through the preset model according to the synthetic facial text, so that face-related tasks are completed by a preset model, thereby improving model performance.

[0060] Based on the first embodiment, a second embodiment of the face information processing method of the present invention is proposed. In this embodiment, step S102 includes:

[0061] Step S201, the first word segmenter converts the face synthesis text into a first text feature sequence, and inputs the first text feature sequence into a deep learning layer;

[0062] Step S202, determining a first target vector corresponding to the first text feature sequence in the Embedding layer;

[0063] Step S203: input the first target vector into a Transformer module to obtain a first output feature output by the Transformer module, and input the first output feature into a data output layer to obtain a target face image.

[0064] It should be noted that the preset model includes a data input layer, a deep learning layer and a data output layer connected in sequence, the data input layer includes a first word segmenter, and the deep learning layer includes an Embedding layer and a Transformer module.

[0065] This embodiment proposes that the face synthesis text is converted into a first text feature sequence through a first word segmenter in a data input layer, wherein the first text feature sequence represents a discrete value, and the first word segmenter can select WordPiece, and then the first text feature sequence is input into a deep learning layer, wherein the deep learning layer includes an Embedding layer and a Transformer module including at least two Transformer layers connected in sequence, and the number of Transformer layers can be increased according to the requirements of the training model.

[0066] Next, the first text feature sequence is converted into the first feature vector in the Embedding layer. The first target vector is then used as the input of the first Transformer layer, and the inputs of the remaining Transformer layers are all the outputs of the previous Transformer layer. The first output feature of the last Transformer layer is obtained and input into the data output layer, where the data output layer includes a VQ-GAN module, which is used to convert discrete values ​​into images.

[0067] Furthermore, in one embodiment, step S202 includes:

[0068] Step a, converting the first text feature sequence into a first feature vector in the Embedding layer, performing position encoding on the first feature vector to determine a first position encoding vector, performing paragraph encoding on the first feature vector to determine a first paragraph encoding vector, and performing character encoding on the first feature vector to determine a first character encoding vector;

[0069] Step b: adding the first position encoding vector, the first paragraph encoding vector, the first character encoding vector and the first feature vector to determine a first target vector.

[0070] In this embodiment, the acquired first text feature sequence is converted into a first feature vector in the Embedding layer, and the first feature vector is position-encoded to determine a first position encoding vector, then the first feature vector is paragraph-encoded to determine a first paragraph encoding vector, the first feature vector is character-encoded to determine a first character encoding vector, and then the first position encoding vector, the first paragraph encoding vector, the first character encoding vector and the first feature vector are added together to determine a first target vector.

[0071] Furthermore, in one embodiment, step S203 includes:

[0072] Step c, inputting the first target vector into the Transformer module;

[0073] Step d, obtaining the probability of each first candidate vector through the softmax layer, obtaining a preset number of first target candidate vectors with the highest probability from the first candidate vectors, and randomly obtaining a second target candidate vector from the first target candidate vectors, and obtaining a first output feature corresponding to the second target candidate vector;

[0074] Step e: inputting the first output feature into the data output layer, and converting the first output feature through the VQ-GAN module to obtain a target face image.

[0075] It should be noted that the Transformer module includes at least two Transformer layers connected in sequence, the last Transformer layer in the Transformer module includes a softmax layer, and the data output layer includes a VQ-GAN module.

[0076] In this embodiment, the first target vector is used as the input of the first Transformer layer in the Transformer module, and the inputs of the remaining Transformer layers are all the outputs of the previous Transformer layer. The probability of each first candidate vector is obtained in the softmax layer. According to the random sampling method, a preset number of first target candidate vectors with a high probability are obtained in the first candidate vector, and the second target candidate vector is randomly obtained in the first target candidate vector. For example, the first candidate vector includes an a vector with a probability of 0.3, a b vector with a probability of 0.1, a c vector with a probability of 0.2, and a d vector with a probability of 0.4. The preset number is 2, then the a vector and the d vector are obtained, and randomly selected from the a vector and the b vector. Finally, the second output feature corresponding to the second target candidate vector is used as the first output feature of the last Transformer layer. Then, the first output feature is obtained, and the first output feature is input to the data output layer. The first output feature is converted by the VQ-GAN module of the data output layer, and the target face image can be obtained.

[0077] The face information processing method proposed in this embodiment converts the face synthesis text into a first text feature sequence through the first word segmenter, inputs the first text feature sequence into the deep learning layer, then determines the first target vector corresponding to the first text feature sequence in the Embedding layer, and finally inputs the first target vector into the Transformer module to obtain the first output feature output by the Transformer module, and inputs the first output feature into the data output layer to obtain the target face image. The method can output the target face image through a preset model based on the face synthesis text, so that face-related tasks are completed by a preset model, thereby improving the model performance.

[0078] Based on the above embodiments, a third embodiment of the face information processing method of the present invention is proposed. In this embodiment, the face information processing method further includes:

[0079] Step 301, obtaining a face image and face attribute text;

[0080] In this embodiment, a face image and face attribute text are obtained. For example, the attribute text can be set to "age prediction", "expression recognition", "face attribute classification", "skin color recognition", etc.

[0081] Step 302: input the face image and the face attribute text into a preset model to obtain target attribute text.

[0082] In this embodiment, the facial attribute text and the facial image are input into a trained Transformer-based facial information processing model to obtain the target attribute text. For example, the input attribute text is "age prediction", and then the facial image is analyzed and processed based on the Transformer-based facial information processing model to obtain the target attribute text, which is the specific value of the predicted age.

[0083] The facial information processing method proposed in this embodiment obtains a facial image and facial attribute text, and then inputs the facial image and the facial attribute text into a preset model to obtain a target attribute text. The target attribute text can be output through the preset model based on the facial image and the facial attribute text, so that face-related tasks are completed by a preset model, thereby improving model performance.

[0084] Based on the third embodiment, a fourth embodiment of the face information processing method of the present invention is proposed. In this embodiment, step S302 includes:

[0085] Step 401, the first word segmenter converts the face attribute text into a second text feature sequence, the VQ-GAN module converts the face image into an image feature sequence, and connects the second text feature sequence and the image feature sequence to determine a target feature sequence;

[0086] Step 402, determining a second target vector corresponding to the target feature sequence in the Embedding layer;

[0087] Step 403: input the second target vector into the Transformer module to obtain the second output feature output by the Transformer module, and input the second output feature into the data output layer to obtain the target attribute text.

[0088] It should be noted that the preset model includes a data input layer, a deep learning layer and a data output layer connected in sequence. The data input layer includes a first word segmenter and a VQ-GAN module, and the deep learning layer includes an Embedding layer and a Transformer module.

[0089] This embodiment proposes that the facial attribute text is converted into a second text feature sequence through the first word segmenter, the VQ-GAN module converts the facial image into an image feature sequence, and the second text feature sequence and the image feature sequence are connected to determine the target feature sequence. For example, the second text feature sequence and the image feature sequence can be connected in sequence according to the order of conversion.

[0090] Next, a second target vector corresponding to the target feature sequence is determined in the Embedding layer. The second target vector is input into the Transformer module to obtain a second output feature output by the Transformer module, and the second output feature is input into the data output layer to obtain the target attribute text. The deep learning layer includes the Embedding layer and the Transformer module includes at least two layers of Transformer layers connected in sequence, and the number of Transformer layers can be increased according to the requirements of the training model.

[0091] Furthermore, in one embodiment, step S402 includes:

[0092] Step f, converting the target feature sequence into a second feature vector in the Embedding layer, performing position encoding on the second feature vector to determine a second position encoding vector, performing paragraph encoding on the second feature vector to determine a second paragraph encoding vector, and performing character encoding on the second feature vector to determine a second character encoding vector;

[0093] Step g: adding the second position encoding vector, the second paragraph encoding vector, the second character encoding vector and the second feature vector to determine a second target vector.

[0094] In this embodiment, the target feature sequence is converted into a second feature vector in the Embedding layer, the second feature vector is position-encoded to determine a second position coding vector, the second feature vector is paragraph-encoded to determine a second paragraph coding vector, and the second feature vector is character-encoded to determine a second character coding vector; the second position coding vector, the second paragraph coding vector, the second character coding vector and the second feature vector are added to determine a second target vector.

[0095] Furthermore, in one embodiment, step S402 includes:

[0096] Step h, inputting the second target vector into the Transformer module;

[0097] Step i, obtaining the probability of each second candidate vector through the softmax layer, obtaining the third target candidate vector with the highest probability among the second candidate vectors, and obtaining the second output feature corresponding to the third target candidate vector;

[0098] Step i: input the second output feature into the data output layer, and convert the second output feature through the second word segmenter to obtain the target attribute text.

[0099] It should be noted that the Transformer module includes at least two Transformer layers connected in sequence, the last Transformer layer in the Transformer module includes a softmax layer, and the data output layer includes a second tokenizer, wherein the second tokenizer is used to convert text discrete values ​​into text.

[0100] In this embodiment, the target feature sequence is converted into a second feature vector in the Embedding layer, the second feature vector is position-encoded to determine a second position coding vector, the second feature vector is paragraph-encoded to determine a second paragraph coding vector, the second feature vector is character-encoded to determine a second character coding vector, and then the second position coding vector, the second paragraph coding vector, the second character coding vector and the second feature vector are added to determine the second target vector.

[0101] Next, the second target vector is used as the input of the first Transformer layer, and the input of the remaining Transformer layers is the output of the previous Transformer layer.

[0102] Next, the probability of each second candidate vector is obtained in the softmax layer, and according to the greedy search algorithm, the third target candidate vector with the highest probability is obtained from the second candidate vectors. For example, the second candidate vectors include an a vector with a probability of 0.3, a b vector with a probability of 0.1, a c vector with a probability of 0.2, and a d vector with a probability of 0.4. The preset number is 2, then the d vector is obtained. Finally, the fourth output feature corresponding to the third target candidate vector is used as the third output feature of the last Transformer layer.

[0103] Finally, the third output feature of the last Transformer layer is obtained, the third output feature is input into the data output layer, and the third output feature is transformed by the second word segmenter to obtain the target attribute text.

[0104] The face information processing method proposed in this embodiment is that the first word segmenter converts the face attribute text into a second text feature sequence, the VQ-GAN module converts the face image into an image feature sequence, connects the second text feature sequence and the image feature sequence to determine the target feature sequence, and then determines the second target vector corresponding to the target feature sequence in the Embedding layer, and then inputs the second target vector into the Transformer module to obtain the second output feature output by the Transformer module, and inputs the second output feature into the data output layer to obtain the target attribute text. The target attribute text can be output through a preset model according to the face image and the face attribute text, so that face-related tasks are completed by a preset model, thereby improving the model performance.

[0105] In addition, an embodiment of the present invention also proposes a facial information processing device, which includes: a memory, a processor, and a facial information processing program stored on the memory and executable on the processor, wherein the facial information processing program implements the steps of the facial information processing method described above when executed by the processor.

[0106] In addition, an embodiment of the present invention further proposes a computer-readable storage medium, on which a face information processing program is stored. When the face information processing program is executed by a processor, the steps of the face information processing method described above are implemented.

[0107] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or system. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or system including the element.

[0108] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0109] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes a number of instructions for a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present invention.

[0110] The above are only preferred embodiments of the present invention, and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. A method for processing face information, It is characterized in that The face information processing method comprises the following steps: Get face synthesis text; The face synthesis text is input into a preset model to obtain a target face image corresponding to the face synthesis text, wherein the preset model is a trained Transformer-based face information processing model, and the preset model includes a data input layer, a deep learning layer, and a data output layer connected in sequence. At this time, the data input layer includes a first word segmenter, and the deep learning layer includes an Embedding layer and a Transformer module. The step of inputting the face synthesis text into the preset model to obtain a target face image corresponding to the face synthesis text includes: The first word segmenter converts the face synthesis text into a first text feature sequence, and inputs the first text feature sequence into the deep learning layer; in the Embedding layer, a first target vector corresponding to the first text feature sequence is determined; the first target vector is input into the Transformer module to obtain a first output feature output by the Transformer module, and the first output feature is input into the data output layer to obtain a target face image; The face information processing method further includes: Acquire a face image and face attribute text; input the face image and the face attribute text into a preset model to obtain a target attribute text. At this time, the data input layer includes a first word segmenter and a VQ-GAN module. The step of inputting the face image and the face attribute text into the preset model to obtain the target attribute text includes: the first word segmenter converts the face attribute text into a second text feature sequence, the VQ-GAN module converts the face image into an image feature sequence, connects the second text feature sequence and the image feature sequence to determine a target feature sequence; determines a second target vector corresponding to the target feature sequence in the Embedding layer; inputs the second target vector into a Transformer module to obtain a second output feature output by the Transformer module, and inputs the second output feature into the data output layer to obtain the target attribute text.

2. The face information processing method according to claim 1, It is characterized in that The step of determining the first target vector corresponding to the first text feature sequence in the Embedding layer includes: In the Embedding layer, the first text feature sequence is converted into a first feature vector, position encoding is performed on the first feature vector to determine a first position encoding vector, paragraph encoding is performed on the first feature vector to determine a first paragraph encoding vector, and character encoding is performed on the first feature vector to determine a first character encoding vector; The first position encoding vector, the first paragraph encoding vector, the first character encoding vector, and the first feature vector are added together to determine a first target vector.

3. The face information processing method according to claim 2, It is characterized in that The Transformer module includes at least two sequentially connected Transformer layers, the last Transformer layer in the Transformer module includes a softmax layer, the data output layer includes a VQ-GAN module, and the step of inputting the first target vector into the Transformer module to obtain a first output feature output by the Transformer module, and inputting the first output feature into the data output layer to obtain a target face image includes: Inputting the first target vector into the Transformer module; Obtaining the probability of each first candidate vector through the softmax layer, obtaining a preset number of first target candidate vectors with the highest probability from the first candidate vectors, and randomly obtaining second target candidate vectors from the first target candidate vectors, and obtaining first output features corresponding to the second target candidate vectors; The first output feature is input into the data output layer, and the first output feature is converted through the VQ-GAN module to obtain a target face image.

4. The face information processing method according to claim 1, It is characterized in that The step of determining the second target vector corresponding to the target feature sequence in the Embedding layer includes: In the Embedding layer, the target feature sequence is converted into a second feature vector, position encoding is performed on the second feature vector to determine a second position encoding vector, paragraph encoding is performed on the second feature vector to determine a second paragraph encoding vector, and character encoding is performed on the second feature vector to determine a second character encoding vector; The second position encoding vector, the second paragraph encoding vector, the second character encoding vector, and the second feature vector are added together to determine a second target vector.

5. The face information processing method according to claim 4, It is characterized in that The Transformer module includes at least two sequentially connected Transformer layers, the last Transformer layer in the Transformer module includes a softmax layer, the data output layer includes a second word segmenter, and the step of inputting the second target vector into the Transformer module to obtain a second output feature output by the Transformer module, and inputting the second output feature into the data output layer to obtain the target attribute text includes: Inputting the second target vector into the Transformer module; Obtaining the probability of each second candidate vector through the softmax layer, obtaining the third target candidate vector with the highest probability among the second candidate vectors, and obtaining the second output feature corresponding to the third target candidate vector; The second output feature is input into the data output layer, and the second output feature is converted by the second word segmenter to obtain the target attribute text.

6. A facial information processing device, It is characterized in that The facial information processing device includes: a memory, a processor, and a facial information processing program stored in the memory and executable on the processor. When the facial information processing program is executed by the processor, the steps of the facial information processing method as described in any one of claims 1 to 5 are implemented.

7. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores a face information processing program, and when the face information processing program is executed by the processor, the steps of the face information processing method as described in any one of claims 1 to 5 are implemented.