Data processing method and device, electronic equipment and computer readable storage medium
This method of generating facial images through multi-round interactions utilizes target description information and historical information sets, combined with facial parameters from the initial and previous rounds for detailed adjustments. This solves the problem of users struggling to efficiently generate facial images that meet their expectations, and achieves efficient and accurate generation of customized facial images.
Patent Information
- Application Number
- CN202510180039.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2026-02-13
- Estimated Expiration
- 2045-02-18
AI Technical Summary
In existing technologies, when users customize game characters or generate facial images, it is difficult to efficiently and accurately meet their personalized needs. In particular, non-professional users spend a lot of time and effort adjusting facial parameters, and the generated character details are difficult to adjust, making it difficult to meet user expectations.
Through multiple rounds of interaction, initial facial parameters are generated based on target description information and historical information sets. Then, by combining facial attribute adjustment information and facial parameters from the previous round, detailed adjustments are made to generate a facial image that meets the user's expectations.
It improves the efficiency and accuracy of facial image customization, and can accurately generate facial images that meet user expectations through multiple rounds of adjustments, reducing the difficulty of operation and time cost for users.
Smart Images

Figure CN119971503B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, in particular to a data processing method and device, electronic equipment and computer readable storage medium. BACKGROUND
[0002] With the vigorous development of electronic games, customizing game characters to meet the personalized needs of users has become a key part of the game. In the process of customizing game characters, users usually manually adjust complex facial parameters such as bone positions and makeup colors to create characters with specific appearances and styles. This method is extremely time-consuming and labor-intensive, especially for non-professional users, and it is more difficult to customize game characters by adjusting facial parameters.
[0003] In order to reduce the cost and difficulty of character customization, in related technologies, a common way is to use natural language processing (NLP) to automatically generate corresponding characters according to the text description provided by the user (such as "cool boy", "more handsome", "more cute", etc.). In some ways, the user can further input instructions to adjust the character. Although this method simplifies the character customization process to some extent, the character generated based on the user's input instructions often has unsatisfactory details, resulting in a character that does not meet the user's expectations.
[0004] Therefore, how to accurately and efficiently customize a character face that meets the user's expectations has become a problem to be solved. SUMMARY
[0005] The present application provides a data processing method, device, electronic equipment and computer readable storage medium, which can make detailed adjustments to the face image based on the initial face parameters and the target face parameters of the previous round, thereby generating a face image that meets the user's expectations and improving the efficiency and accuracy of face image customization. The specific scheme is as follows:
[0006] In a first aspect, the embodiments of the present application provide a data processing method, which comprises:
[0007] According to the target description information corresponding to the target round, the initial face parameters corresponding to the target round are generated; wherein the target description information is used to describe the adjustment of the face image in the target round;
[0008] According to the target description information, the face attribute adjustment information corresponding to the target round is determined; wherein the face attribute adjustment information is used to indicate the adjustment state of each face attribute in the target round;
[0009] According to the face attribute adjustment information, the initial face parameters corresponding to the target round and the target face parameters corresponding to the previous round of the target round, the target face image corresponding to the target round is generated.
[0010] In a second aspect, an embodiment of the present application provides a data processing apparatus, the apparatus comprising:
[0011] a first generating unit configured to generate initial face parameters corresponding to a target round according to target description information corresponding to the target round, wherein the target description information is used to describe an adjustment to be performed on the face image in the target round;
[0012] a determining unit configured to determine face attribute adjustment information corresponding to the target round according to the target description information, wherein the face attribute adjustment information is used to indicate an adjustment state of each face attribute in the target round;
[0013] a second generating unit configured to generate a target face image corresponding to the target round according to the face attribute adjustment information, the initial face parameters corresponding to the target round, and target face parameters corresponding to a previous round of the target round.
[0014] In a third aspect, the present application further provides an electronic device, comprising:
[0015] a processor; and
[0016] a memory configured to store a data processing program, and the electronic device, after being powered on and running the program by the processor, executes the method of the first aspect.
[0017] In a fourth aspect, an embodiment of the present application further provides a computer readable storage medium, which stores a data processing program, and the program, when being run by a processor, executes the method of the first aspect.
[0018] Compared with the prior art, the present application has the following advantages:
[0019] The data processing method provided by the present application comprises the following steps: generating initial face parameters corresponding to a target round according to target description information corresponding to the target round, wherein the target description information is used to describe an adjustment to be performed on the face image in the target round; determining face attribute adjustment information corresponding to the target round according to the target description information, wherein the face attribute adjustment information is used to indicate an adjustment state of each face attribute in the target round; and generating a target face image corresponding to the target round according to the face attribute adjustment information, the initial face parameters corresponding to the target round, and target face parameters corresponding to a previous round of the target round.
[0020] As can be seen, the initial facial parameters generated through the target description information can reflect the facial features of the facial image to be generated. The facial attribute adjustment information determined by the target description information accurately indicates the adjustment status of each facial attribute in the target round. Thus, based on the facial attribute adjustment information, the initial facial parameters of the target round, and the target facial parameters of the previous round, detailed adjustments can be made in the target round to generate a detailed adjusted target facial image. Therefore, the data processing method provided in this application embodiment can perform detailed adjustments on the facial image based on the initial facial parameters and the target facial parameters of the previous round, thereby generating a facial image that meets the user's expectations and improving the efficiency and accuracy of facial image customization. Attached Figure Description
[0021] Figure 1 This is a data processing system diagram provided in an embodiment of the present application for implementing a data processing method;
[0022] Figure 2 This is a flowchart of the data processing method provided in this application;
[0023] Figure 3 This is a flowchart detailing the algorithm of the data processing method provided in the embodiments of this application;
[0024] Figure 4 This is a structural block diagram of an example of the data processing apparatus provided in the embodiments of this application;
[0025] Figure 5 This is a structural block diagram of an example of an electronic device for data processing provided in an embodiment of this application. Detailed Implementation
[0026] Many specific details are set forth in the following description to provide a full understanding of this application. However, this application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this application; therefore, this application is not limited to the specific embodiments disclosed below.
[0027] It should be noted that the terms "first", "second", "third", etc. in the claims, specification and drawings of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. The data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include", "have" and their variants are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0028] It should be understood that in the embodiments of the present application, "at least one" means one or more, and "multiple" means two or more. "And / or" is only a description of the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can mean that A exists alone, A and B exist together, and B exists alone. The character " / " generally represents that the associated objects before and after it are in an "or" relationship. "Including A, B and / or C" means including any one or any two or three of A, B and C.
[0029] It should be understood that in the embodiments of the present application, "B corresponding to A", "B corresponding to A", "A corresponding to B" or "B corresponding to A" means that B is associated with A, and B can be determined according to A. Determining B according to A does not mean that B is determined only according to A, but also can be determined according to A and / or other information.
[0030] Before the embodiments of the present application are described in detail, the prior art is further described first.
[0031] In role-playing games (RPG), augmented reality (AR) and virtual reality (VR) games, etc., creating a game character that meets specific needs for users has become an important research direction. With the continuous development of the game industry and the progress of technology, users' demand for creating personalized characters that meet specific aesthetics or styles is growing. However, implementing personalized character customization often involves complex parameter adjustments, including but not limited to adjustments of facial bone positions, skin textures, makeup colors, etc., which not only requires users to master professional knowledge, but also is quite time-consuming and laborious.
[0032] In order to reduce the labor cost and time cost in character customization, the industry often uses the following methods to perform character customization:
[0033] One is to customize the character based on a reference image, which allows the user to upload a reference image, and the system will automatically adjust the facial features of the character according to the reference image to generate a character with facial features highly similar to the reference image.
[0034] The second is to customize the character based on a text description, which allows the user to input a text description, and the system will generate a character according to the text description, such as "cool boy" or "cuter", to generate a character that meets the text description.
[0035] However, the above methods can only generate a character at a time, and lack the ability to iterate the generated character, which means that once the initially generated character does not fully meet the user's expectations, the user will have difficulty making subtle adjustments, thereby affecting the final result. And in the above-mentioned way, it is difficult for the user to make specific modifications to the character, especially when pursuing details. This makes it particularly difficult to achieve complex or abstract style descriptions, especially for ordinary players who have no professional background.
[0036] Based on the above reasons, in order to be able to flexibly adjust the facial image through multiple rounds of dialogue and generate a facial image that meets the user's expectations, thereby improving the efficiency and accuracy of facial image customization, the first embodiment of the present application provides a data processing method, which is applied to an electronic device. The electronic device can be a desktop computer, a notebook computer, a mobile phone, a tablet computer, an electronic watch, etc., or other electronic devices capable of data processing, and the embodiments of the present application are not specifically limited.
[0037] The data processing method provided by the embodiments of the present application can be applied to the face pinching scene in the game, so that a facial image that meets the user's expectations can be generated in the game; the data processing method can also be applied to the scene of chatting with an intelligent agent model in human-computer interaction, so that a facial image that meets the user's expectations can be generated in human-computer interaction.
[0038] In an optional embodiment, when the data processing method is running on a terminal device, the terminal device can include a display screen and a processor, the display screen is used to present an interactive picture and receive the description information input by the user. The interactive picture can include an information interaction area, and the user can input the description information in the information interaction area. The target facial image generated in each round can be displayed in the information interaction area. The processor is used to store the data processing program, run the program, generate the interactive picture, respond to the instruction, and control the display of the interactive picture on the display screen. When the user operates the interactive picture through the display screen, the interactive picture can control the content of the terminal device locally by responding to the received operation instruction. The way the terminal device provides the graphical user interface to the user can include multiple ways, for example, the graphical user interface can be rendered and displayed on the display screen of the terminal device, or the graphical user interface can be presented through holographic projection.
[0039] In an optional embodiment, when the digital processing method is run on a server, the method can be implemented and executed based on a cloud system. The cloud system is based on cloud computing. The cloud system includes a server and a client device. The running subject of the data processing program and the client interface presentation subject are separated, the storage and running of the data processing method are completed on the server. The client interface presentation is completed on the client, and the client is mainly used for data receiving, sending and client interface presentation. For example, the client can be a display device close to the user side with data transmission function, such as a mobile terminal, a television, a computer, a palm computer, a personal digital assistant, a head-mounted display device, etc., but the terminal device for data processing is the server in the cloud. During the use of the application corresponding to the data processing program, the user instructs the client to send instructions to the server, the server encodes and compresses the client interface and other data according to the instructions, returns the client through the network, and finally decodes and outputs the client interface through the client.
[0040] It should be noted that in the embodiments of the present application, the execution subject of the data processing method can be a terminal device or a server, wherein the terminal device can be a local terminal device or a client device in the cloud system mentioned above. The embodiments of the present application do not limit the type of execution subject.
[0041] For example, in combination with the above description, Figure 1 A data processing system 100 for implementing the data processing method is shown, which can include at least one terminal 101, at least one server 102, and a network 103. The terminal 101 held by the user can be connected to the server 102 through the network. The terminal is any device with computing hardware that can support and execute the software application tool corresponding to the intelligent physical therapy device.
[0042] In the above data processing system 100, the terminal 101 is used to install and run the application program corresponding to the above data processing program. In some cases, the application program can not be installed in advance in the terminal 101, and the user can directly access the application program through a browser or other client. During the process in which the user enters the application program through the application program, the terminal 101 and the server 102 can interact with each other, the terminal 101 sends various information to the server 102, the server 102 determines the display data of the terminal 101 according to the received information, and sends the display data to the terminal 101, so as to display the display data sent by the server 102 to the user through the terminal 101.
[0043] In possible application scenarios, different terminals 101 can be served by different servers 102, and the servers 102 corresponding to different terminals 101 can be the same server.
[0044] In addition, when the data processing system 100 includes multiple terminals, multiple servers, and multiple networks, different terminals can be connected to each other through different networks and through different servers.
[0045] The terminal 101 can have one or more multi-touch screens for sensing and obtaining the input of a user through a touch or sliding operation performed at one or more points of the touch display screen. The terminal 101 can also be connected to a keyboard and / or a mouse and / or a gamepad, etc., so that the user can perform interface operations through the keyboard and / or the mouse and / or the gamepad.
[0046] The network can be a wireless network or a wired network. For example, the wireless network can be a wireless local area network (WLAN), a local area network (LAN), a cellular network, a 2G network, a 3G network, a 4G network, a 5G network, etc. In addition, different terminals can also use their own Bluetooth networks or hotspot networks to connect to other terminals or servers, etc. In addition, the system 100 can include multiple databases coupled to different servers, and the data generated by the application of the above data processing program can be continuously stored in the databases.
[0047] It should be noted that, Figure 1 The data processing system diagram shown is only an example, and the data processing system 100 described in the embodiments of the present application is used to more clearly illustrate the technical solutions of the embodiments of the present application, and does not constitute a limitation on the technical solutions provided by the embodiments of the present application.
[0048] The technical solutions of the present application will be described in detail below through specific embodiments. It should be noted that the following specific embodiments can be combined with each other, and the same or similar concepts or processes can not be described in detail in some embodiments.
[0049] In the following, Figure 2 and Figure 3 The data processing method provided by the embodiments of the present application is introduced.
[0050] As Figure 2 shown, the flowchart of the data processing method provided by the present application includes the following steps S103-S105.
[0051] It should be noted that before step S103, the data processing method provided by the present application can further include the following steps S101 and S102:
[0052] Step S101: obtaining first description information input by a target round for a face image to be generated, and a historical information set corresponding to the target round.
[0053] In the embodiments of the present application, when generating the expected face image, multiple rounds of interaction can be performed, and the target round refers to any round except the first round in the multiple rounds of interaction. The first description information refers to information input by the target round for the face image to be generated, and the first description information represents the adjustment requirement of the face image to be generated in the target round. The first description information can be text form or voice form, which is not limited in the present application.
[0054] The historical information set corresponding to the target round can be an information set composed of historical information corresponding to one or more historical rounds before the target round. For example, the target round is the fifth round, and the historical information set corresponding to the target round can be an information set composed of historical information corresponding to the fourth round, or an information set composed of historical information corresponding to the first round to the fourth round.
[0055] Optionally, the historical information corresponding to the historical round can at least include description information input by the historical round for the face image to be generated.
[0056] Step S102: generating target description information corresponding to the target round according to the first description information and the historical information set, wherein the target description information corresponding to the target round is used to describe the adjustment of the face image to be performed in the target round.
[0057] It should be noted that, in addition to generating the target description information corresponding to the target round according to the first description information and the historical information set, the target editing strength corresponding to the target round can also be generated. The target editing strength corresponding to the target round represents the influence degree of the target description information corresponding to the target round on the continuous type of face attributes when generating the target face parameters corresponding to the target round.
[0058] This step is used to analyze the first description information input by the target round and the historical information set corresponding to the target round, and obtain the target description information used to describe the adjustment of the face image to be performed in the target round.
[0059] It should be noted that the first description information is the adjustment requirement of the user input in the target round for the generated face image, the first description information is usually colloquial information, and the first description information does not necessarily include the face attribute to be adjusted or the adjustment method of the face attribute to be adjusted; the target description information is information obtained by analyzing the first description information on the basis of the historical information set, and is used to feedback the adjustment to be performed on the face image in the target round on the basis of the combination of the information in the historical information set and the first description information input text, and the target description information includes the face attribute to be adjusted and the adjustment method of the face attribute to be adjusted.
[0060] For example, assuming that the historical information in a certain historical round in the historical information set includes "nose bridge heightening", and the first description information of the target round is "nose bridge adjustment is not enough", the target description information of the target round is "nose bridge heightening", indicating that the nose bridge of the face image needs to be further heightened on the basis of the last round. Assuming that the historical information in the last round in the historical information set is "eye enlargement", and the first description information of the target round is "enlargement again", the target description information of the target round is "eye enlargement", indicating that the eye of the face image needs to be further enlarged on the basis of the last round.
[0061] It should be noted that when the first description information includes the face attribute to be adjusted and the adjustment method of the face attribute to be adjusted, the target description information of the target round is usually consistent with the first description information. Assuming that the first description information of the target round is "eye enlargement", the target description information is "eye enlargement", indicating that the eye of the face image needs to be further enlarged on the basis of the last round.
[0062] In this step, the first description information and the historical information set are analyzed, and the target description information corresponding to the target round is generated, and the target editing strength is also generated, which is used to represent the influence degree of the target description information corresponding to the target round on the continuous type face attribute when generating the target face parameter corresponding to the target round.
[0063] In the embodiment of the application, the target editing strength generated for the target round can be located in the target interval [0, 1].
[0064] When the target editing strength is 0, it indicates that the parameter corresponding to the continuous type of facial attribute does not change, and the parameter corresponding to the continuous type of facial attribute in the target facial parameter corresponding to the target round continues to use the parameter corresponding to the continuous type of facial attribute in the target facial parameter of the last round. When the target editing strength is 1, it indicates that the parameter corresponding to the continuous type of facial attribute is completely affected by the target description information, and the parameter corresponding to the continuous type of facial attribute in the target facial parameter corresponding to the target round is the parameter corresponding to the continuous type of facial attribute in the initial facial parameter of the target round.
[0065] When the target editing strength is between 0 and 1, it indicates that the parameter corresponding to the continuous type of facial attribute is affected by both the parameter corresponding to the continuous type of facial attribute in the target facial parameter of the last round and the target description information, and the parameter corresponding to the continuous type of facial attribute in the target facial parameter corresponding to the target round is the parameter obtained by interpolating and fusing the parameter corresponding to the continuous type of facial attribute in the initial facial parameter of the target round and the parameter corresponding to the continuous type of facial attribute in the target facial parameter of the last round.
[0066] In this way, through the target editing strength, the parameters of the continuous type can be finely adjusted, and the accuracy of face image customization is further improved.
[0067] The specific influence mode of the target editing strength on the continuous type of facial attribute will be described in detail in subsequent steps.
[0068] In this step, the target description information and the target editing strength are generated through the first description information and the historical information set, the historical rounds of the target round are fully considered, the generated target description information includes the facial attributes to be adjusted and the adjustment mode of the facial attributes to be adjusted, the generated target editing strength combines the understanding of the context, and accurately indicates the influence degree of the target description information on the continuous type of facial attribute, thereby providing a good foundation for accurately and efficiently generating a face image meeting the user's expectation.
[0069] Step S103: generating an initial facial parameter corresponding to the target round according to target description information corresponding to the target round; wherein the target description information is used to describe the adjustment to be performed on the face image in the target round.
[0070] The initial facial parameters can be used to describe facial features corresponding to the target description information, and the facial features can include facial attributes corresponding to each facial attribute. The facial attributes can include facial attributes of five facial bone dimensions, facial attributes of makeup position and shape dimensions, and facial attributes of makeup color dimensions, etc. The facial attributes of the five facial bone dimensions can include, but are not limited to, forehead, temple, chin, cheekbone, cheek, apple muscle, lower jaw, eyebrow, eye, nose, mouth, and ear. The facial attributes of the makeup position and shape dimensions can include, but are not limited to, blush, eyebrow, eyeliner, eye shadow, eyelash, eyelid, lip, skin, facial feature, scar, iris, upper beard, and lower beard. The facial attributes of the makeup color dimensions can include, but are not limited to, blush color, facial line color, eyebrow color, eye color, eye shadow color, eyeliner color, lip color, and beard color.
[0071] In an optional embodiment, the step of generating the initial facial parameters corresponding to the target round according to the target description information can be implemented by the following steps:
[0072] Generating an initial facial image corresponding to the target round according to the target description information;
[0073] Determining the initial facial parameters corresponding to the initial facial image.
[0074] In this embodiment, the initial facial image corresponding to the target round can be generated according to the target description information first. The initial facial image can reflect the facial features corresponding to the target description information, that is, the facial attributes to be adjusted included in the target description information are adjusted according to the adjustment manner included in the target description information. Then, the initial facial image is converted into the initial facial parameters in the form of parameters.
[0075] It should be noted that the initial facial image of the target round can be an image that has no association with the target facial image corresponding to the historical round. For example, the target description information of the target round is that the eyes are a little larger, and the initial facial image of the target round is an image in which the eyes are larger than the preset size. Other facial attributes such as eyebrows, nose, and mouth cannot be guaranteed to be exactly the same as the target facial image corresponding to the historical round.
[0076] Step S104: determining facial attribute adjustment information corresponding to the target round according to the target description information; wherein the facial attribute adjustment information is used to indicate the adjustment state of each facial attribute in the target round.
[0077] The face attribute adjustment information is used to indicate an adjustment state of each face attribute in the target round, which can represent whether the corresponding face attribute is adjusted or not. Specifically, the face attribute adjustment information can include sub-adjustment information corresponding to each face attribute respectively; the sub-adjustment information includes first adjustment information or second adjustment information, the first adjustment information is used to indicate that the corresponding face attribute is in a first state to be adjusted in the target round, and the second adjustment information is used to indicate that the corresponding face attribute is in a second state not to be adjusted in the target round.
[0078] For each face attribute, the sub-adjustment information corresponding to the face attribute can be a binary value, when the binary value is 0, it represents that the corresponding face attribute is not adjusted in the target round, and when the binary value is 1, it represents that the corresponding face attribute is to be adjusted in the target round, that is, the first adjustment information is 1, and the second adjustment information is 0.
[0079] For example, the face attribute adjustment information includes five bits, the first bit is the hair color, the second bit is the eye color, the third bit is the skin color, the fourth bit is the smile degree, and the fifth bit is the eyebrow sparseness degree. Assuming that the face attribute adjustment information of the target round is 01001, it represents that the hair color needs to be adjusted, the eye color and the skin color do not need to be adjusted, the smile degree needs to be adjusted, and the eyebrow sparseness degree does not need to be adjusted in the target round.
[0080] In this step, by analyzing the target description information, classifying according to the face attributes, and obtaining the face attribute adjustment information of whether each face attribute is modified, it can be determined which face attributes in the target round will be modified relative to the last round and which face attributes will not be modified relative to the last round, so as to accurately generate the target face image of the target round.
[0081] Step S105: generating a target face image corresponding to the target round according to the face attribute adjustment information, the initial face parameters corresponding to the target round, and the target face parameters corresponding to the last round of the target round.
[0082] After determining the face attribute adjustment information corresponding to the target round and the initial face parameters corresponding to the target round, the target face parameters corresponding to the last round of the target round can be combined to generate the target face image corresponding to the target round.
[0083] In an optional implementation, the above step S105 can be realized by the following steps:
[0084] determining the target face parameters corresponding to the target round according to the face attribute adjustment information, the initial face parameters corresponding to the target round, and the target face parameters corresponding to the last round;
[0085] According to the target face parameter corresponding to the target round, a target face image corresponding to the target round is generated.
[0086] It should be noted that each round of non-first round has an initial face parameter and a target face parameter, and the target face parameter of each round is generated based on the face attribute adjustment information of the current round and the initial face parameter of the current round and the target face parameter of the previous round.
[0087] In this way, according to the face attribute adjustment information corresponding to the target round, the initial face parameter corresponding to the target round, and the target face parameter corresponding to the previous round, the target face parameter corresponding to the target round can be determined.
[0088] Since the initial face parameter of the target round is used to describe the face features corresponding to the target description information, and the face attribute adjustment information of the target round is used to indicate the adjustment state of each face attribute in the target round, in combination with the target face parameter of the previous round, the target face parameter after detail adjustment can be generated in the target round.
[0089] Optionally, after generating the target face parameter after detail adjustment, the target face parameter can be rendered by a game engine to generate a target face image after detail adjustment.
[0090] Optionally, after generating the target face parameter after detail adjustment, the target face parameter can be converted into a target face image by an image generation model.
[0091] It should be noted that after generating the target face image, the target face image can be displayed in a graphical user interface to enable a user to determine whether the target face image of the target round meets expectations.
[0092] The data processing method provided by the embodiments of the present application includes the following steps: generating an initial face parameter corresponding to a target round according to target description information corresponding to the target round; wherein the target description information is used to describe the adjustment to be performed on a face image in the target round; determining face attribute adjustment information corresponding to the target round according to the target description information; wherein the face attribute adjustment information is used to indicate the adjustment state of each face attribute in the target round; and generating a target face image corresponding to the target round according to the face attribute adjustment information, the initial face parameter corresponding to the target round, and the target face parameter corresponding to the previous round of the target round.
[0093] As can be seen, the initial facial parameters generated through the target description information can reflect the facial features of the facial image to be generated. The facial attribute adjustment information determined by the target description information accurately indicates the adjustment status of each facial attribute in the target round. Thus, based on the facial attribute adjustment information, the initial facial parameters of the target round, and the target facial parameters of the previous round, detailed adjustments can be made in the target round to generate a detailed adjusted target facial image. Therefore, the data processing method provided in this application embodiment can perform detailed adjustments on the facial image based on the initial facial parameters and the target facial parameters of the previous round, thereby generating a facial image that meets the user's expectations and improving the efficiency and accuracy of facial image customization.
[0094] In one optional implementation, the aforementioned historical information set includes historical information corresponding to at least one historical round prior to the target round. For a given historical round, the corresponding historical information includes the description information of the facial image input to be generated in that historical round, the target description information corresponding to that historical round, and the target editing intensity corresponding to that historical round.
[0095] As shown in the following formula (1), it is the expression for the historical information set in the data processing method provided in the embodiments of this application:
[0096] H i ={τ i-m ,(t i-m ,s i-m ),…,τ i-1 ,(t i-1 ,s i-1 )} Formula (1)
[0097] In the above formula (1), m is the number of selected historical rounds, and H i Let τ be the set of historical information for the i-th round, 1≤m<i, and τ i-m For the im-th round, the input description information is the facial image to be generated, t i-m For the target description information corresponding to the im-th round, s i-m Let τ be the target editing intensity corresponding to the im-th round. i-1 For the (i-1)th round, the input description information for the facial image to be generated, t i-1 For the target description information corresponding to the (i-1)th round, s i-1 The target editing intensity corresponds to the (i-1)th round.
[0098] For example, when i is 10 and m is 5, the historical information set corresponding to the 10th round includes the description information of the facial image input to be generated in the 5th round, the target description information corresponding to the 5th round, the target editing intensity corresponding to the 5th round, the description information of the facial image input to be generated in the 6th round, the target description information corresponding to the 6th round, the target editing intensity corresponding to the 6th round, the description information of the facial image input to be generated in the 7th round, the target description information corresponding to the 7th round, the target editing intensity corresponding to the 7th round, the description information of the facial image input to be generated in the 8th round, the target description information corresponding to the 8th round, the target editing intensity corresponding to the 8th round, the description information of the facial image input to be generated in the 9th round, the target description information corresponding to the 9th round, and the target editing intensity corresponding to the 9th round.
[0099] The following describes in detail the determination of the target facial parameter corresponding to the target round:
[0100] In an optional implementation, the step of "determining the target facial parameter corresponding to the target round according to the facial attribute adjustment information, the initial facial parameter corresponding to the target round, and the target facial parameter corresponding to the previous round" can specifically include the following steps:
[0101] According to the facial attribute adjustment information, a first facial attribute in which the adjustment state is a first state is determined; the first state indicates that the first facial attribute is a facial attribute to be adjusted in the target round.
[0102] According to a first sub-parameter corresponding to the first facial attribute in the initial facial parameter corresponding to the target round and a second sub-parameter corresponding to the first facial attribute in the target facial parameter corresponding to the previous round, a third sub-parameter corresponding to the first facial attribute in the target round is determined.
[0103] The second sub-parameter in the target facial parameter corresponding to the previous round is replaced by the third sub-parameter to obtain the target facial parameter corresponding to the target round.
[0104] In this implementation, each facial attribute can be divided, and specifically, the first facial attribute in which the adjustment state is the first state and the second facial attribute in which the adjustment state is the second state can be divided. The first state indicates that the first facial attribute is a facial attribute to be adjusted in the target round, and the second state indicates that the second facial attribute is a facial attribute that does not undergo adjustment in the target round. As known from the foregoing description, for each facial attribute, the sub-adjustment information corresponding to the facial attribute can be a binary value. In this way, the facial attribute with a binary value of 1 can be determined as the first facial attribute, and the facial attribute with a binary value of 2 can be determined as the second facial attribute.
[0105] Then, a first sub-parameter corresponding to the first facial attribute in the initial facial parameter corresponding to the target round and a second sub-parameter corresponding to the first facial attribute in the target facial parameter corresponding to the previous round can be obtained, so as to determine a third sub-parameter of the first facial attribute corresponding to the target round according to the first sub-parameter and the second sub-parameter, the third sub-parameter being a sub-parameter corresponding to the first facial attribute in the target facial parameter corresponding to the target round.
[0106] In this way, the second sub-parameter in the target facial parameter corresponding to the previous round can be replaced by the third sub-parameter, so as to obtain the target facial parameter corresponding to the target round.
[0107] It should be noted that, since the second facial attribute is a facial attribute that does not undergo adjustment in the target round, the sub-parameter corresponding to the second facial attribute in the target facial parameter corresponding to the target round still uses the sub-parameter corresponding to the second facial attribute in the target facial parameter corresponding to the previous round, so that the second facial attribute does not undergo adjustment in the target round.
[0108] In this way, by determining the first facial attribute that undergoes change in the target round, only the first facial attribute in the target facial parameter of the previous round is updated, and the second facial attribute in the target facial parameter of the previous round is kept unchanged, so as to realize the detail adjustment of the facial image.
[0109] In an optional implementation, the step of "determining a third sub-parameter of the first facial attribute corresponding to the target round according to a first sub-parameter corresponding to the first facial attribute in an initial facial parameter corresponding to the target round and a second sub-parameter corresponding to the first facial attribute in a target facial parameter corresponding to a previous round" can include the following steps S201 and S202:
[0110] Step S201: determining a first weight corresponding to the first sub-parameter corresponding to the first facial attribute in the initial facial parameter corresponding to the target round and a second weight corresponding to the second sub-parameter corresponding to the first facial attribute in the target facial parameter corresponding to the previous round based on a parameter type of the first facial attribute; wherein the sum of the first weight and the second weight is 1;
[0111] Step S202: weighting the first sub-parameter and the second sub-parameter according to the first weight and the second weight to obtain the third sub-parameter of the first facial attribute corresponding to the target round.
[0112] It should be noted that the face attributes introduced above can be divided into continuous type and discrete type according to the parameter type. The discrete type of face attribute refers to the face attribute with a fixed number of parameters, which is usually used to describe the characteristics with explicit classification or selection; the continuous type of face parameter refers to the face attribute with parameters that can change arbitrarily within a certain range, which is usually used to describe the characteristics that can be measured by specific values and can change smoothly within a certain range.
[0113] For example, the discrete type of face attribute can include makeup type, expression state, skin color tone, etc. The parameters of the makeup type can include natural makeup, heavy makeup, smoky makeup, bare makeup, etc. Each makeup style is a clear category, rather than a smooth change process; the expression state can include smile, frown, surprise, anger, etc., each of which is a clear type; the skin color tone can include white, yellow, black, etc. The tone of the category.
[0114] For example, the continuous type of face attribute can include eye spacing, nose height, apple muscle size, etc. The parameters of the continuous type of face attribute change continuously within a certain range, and there is no obvious boundary.
[0115] In this embodiment, the first weight corresponding to the first sub-parameter of the first face attribute in the initial face parameter corresponding to the target round can be determined based on the parameter type of the first face attribute, and then the difference between 1 and the first weight can be taken as the second weight corresponding to the second sub-parameter of the first face attribute in the target face parameter corresponding to the last round.
[0116] The first weight is the influence degree of the first sub-parameter on the third sub-parameter in the target face parameter corresponding to the target round to be generated, and the second weight is the influence degree of the second sub-parameter on the third sub-parameter in the target face parameter corresponding to the target round to be generated.
[0117] Then, the first sub-parameter and the second sub-parameter can be weighted and summed according to the first weight and the second weight, so as to obtain the third sub-parameter of the first face attribute corresponding to the target round.
[0118] As shown in the following formula (2), it is a calculation formula of the third sub-parameter in the data processing method provided by the embodiment of the application:
[0119] x i =x t *m w +x i-1 *(1-m w ) Formula (2)
[0120] In the above formula (2), xi is a third sub-parameter corresponding to the first facial attribute in the initial facial parameter corresponding to the target round, x t is a first sub-parameter corresponding to the first facial attribute in the initial facial parameter corresponding to the target round, m w is a first weight, x i-1 is a second sub-parameter corresponding to the first facial attribute in the target facial parameter corresponding to the previous round, 1-m w is a second weight.
[0121] For the first facial attribute of the discrete type and the first facial attribute of the continuous type, the calculation manner of the first weight corresponding to different parameter types is different. The calculation manner of the first weight under different parameter types is introduced as follows:
[0122] Optionally, the step of "determining the first weight corresponding to the first sub-parameter corresponding to the first facial attribute in the initial facial parameter corresponding to the target round based on the parameter type of the first facial attribute" can include the following steps:
[0123] If the parameter type of the first facial attribute is the discrete type, the first weight corresponding to the first sub-parameter corresponding to the first facial attribute in the initial facial parameter corresponding to the target round is determined as 1.
[0124] If the parameter type of the first facial attribute is the discrete type, such as the makeup type, the expression state, the skin color tone and the like introduced in the foregoing, the first weight corresponding to the first sub-parameter corresponding to the first facial attribute in the initial facial parameter corresponding to the target round can be determined as 1, and correspondingly, the second weight corresponding to the second sub-parameter corresponding to the first facial attribute in the target facial parameter corresponding to the previous round is 0.
[0125] That is to say, for the first facial attribute to be adjusted in the target round, in the case where the parameter type of the first facial attribute is the discrete type, the third sub-parameter corresponding to the first facial attribute in the target facial parameter corresponding to the target round is the first sub-parameter corresponding to the first facial attribute in the initial facial parameter corresponding to the target round.
[0126] Optionally, the step of "determining the first weight corresponding to the first sub-parameter corresponding to the first facial attribute in the initial facial parameter corresponding to the target round based on the parameter type of the first facial attribute" can include the following steps:
[0127] If the parameter type of the first facial attribute is the continuous type, the normalized position data corresponding to the first sub-parameter is determined.
[0128] According to the normalized position data, the first weight corresponding to the first sub-parameter is determined.
[0129] If the first facial attribute is of a continuous type, such as the aforementioned double eye distance, nose bridge height, apple muscle size, etc., the normalized position data of the first sub-parameter corresponding to the first facial attribute in the initial facial parameter corresponding to the target round can be determined first, and the normalized position data is used to indicate the relative change of the first facial attribute in the target round.
[0130] In an optional implementation manner, the normalized position data can be determined in the following manner:
[0131] According to the data relationship between the first sub-parameter and the second sub-parameter, position reference data corresponding to the first sub-parameter is determined.
[0132] According to the position reference data, the first sub-parameter and the second sub-parameter, the normalized position data corresponding to the first sub-parameter is determined.
[0133] In the implementation manner, the data relationship between the first sub-parameter corresponding to the first facial attribute in the initial facial parameter corresponding to the target round and the second sub-parameter corresponding to the first facial attribute in the target facial parameter corresponding to the previous round can be determined first, and the data relationship is used to represent the change of the first facial attribute in the target round relative to the previous round. The data relationship can include a first data relationship that the first sub-parameter is greater than the second sub-parameter, or a second data relationship that the first sub-parameter is less than the second sub-parameter. The first data relationship is used to represent that the change of the first facial attribute in the target round relative to the previous round is that the parameter becomes larger, and the second data relationship is used to represent that the change of the first facial attribute in the target round relative to the previous round is that the parameter becomes smaller.
[0134] In the embodiment, according to the data relationship between the first sub-parameter and the second sub-parameter, position reference data corresponding to the first sub-parameter is determined, and the position reference data is reference data of the normalized position data.
[0135] In different data relationships, the step of “determining position reference data corresponding to the first sub-parameter according to the data relationship between the first sub-parameter and the second sub-parameter” can include the following two cases:
[0136] Case one, if the first sub-parameter is greater than the second sub-parameter, the maximum value of the first facial attribute in each target facial parameter corresponding to each round is determined as the position reference data corresponding to the first sub-parameter.
[0137] In the case one, the data relationship between the first sub-parameter and the second sub-parameter is the first data relationship, in which case, the maximum value of the first facial attribute in the target facial parameters corresponding to each round can be determined as the position reference data corresponding to the first sub-parameter. For example, if there are totally 10 rounds, the maximum value of the first facial attribute in the target facial parameters corresponding to the 10 rounds can be determined as the position reference data corresponding to the first sub-parameter.
[0138] In the case two, if the first sub-parameter is less than the second sub-parameter, the minimum value of the first facial attribute in the target facial parameters corresponding to each round can be determined as the position reference data corresponding to the first sub-parameter.
[0139] In the case two, the data relationship between the first sub-parameter and the second sub-parameter is the second data relationship, in which case, the minimum value of the first facial attribute in the target facial parameters corresponding to each round can be determined as the position reference data corresponding to the first sub-parameter. For example, if there are totally 10 rounds, the minimum value of the first facial attribute in the target facial parameters corresponding to the 10 rounds can be determined as the position reference data corresponding to the first sub-parameter.
[0140] After the position reference data corresponding to the first sub-parameter is determined, the normalized position data corresponding to the first sub-parameter can be determined according to the position reference data.
[0141] Specifically, the normalized position data can be determined by the following steps:
[0142] determining a first difference between the position reference data and the second sub-parameter;
[0143] determining a second difference between the first sub-parameter and the second sub-parameter;
[0144] determining the absolute value of the ratio of the first difference and the second difference as the normalized position data corresponding to the first sub-parameter.
[0145] In the embodiment, the first difference between the position reference data and the second sub-parameter can be determined first. Specifically, if the first sub-parameter is greater than the second sub-parameter, the first difference between the maximum value of the first facial attribute in the target facial parameters corresponding to each round and the second sub-parameter can be determined; if the first sub-parameter is less than the second sub-parameter, the first difference between the minimum value of the first facial attribute in the target facial parameters corresponding to each round and the second sub-parameter can be determined.
[0146] Then, a second difference value between a first sub-parameter corresponding to the first facial attribute in the initial facial parameter corresponding to the target round and a second sub-parameter corresponding to the first facial attribute in the target facial parameter corresponding to the last round is determined.
[0147] Finally, a ratio between the first difference value and the second difference value is determined, and an absolute value corresponding to the ratio is determined as the normalized position data corresponding to the first parameter. It should be noted that when the first sub-parameter is greater than the second sub-parameter, the normalized position data is used to represent a relative position of the first sub-parameter relative to a maximum value of the first facial attribute in the target facial parameters corresponding to the rounds and the second sub-parameter; when the first sub-parameter is less than the second sub-parameter, the normalized position data is used to represent a relative position of the first sub-parameter relative to a minimum value of the first facial attribute in the target facial parameters corresponding to the rounds and the second sub-parameter.
[0148] As shown in the following formula (3), is a calculation formula of the normalized position data corresponding to the data processing method provided by the embodiments of the present application:
[0149]
[0150] In the above formula (3), c is the normalized position data corresponding to the first sub-parameter, x max is a maximum value of the first facial attribute in the target facial parameters corresponding to the rounds, x min is a minimum value of the first facial attribute in the target facial parameters corresponding to the rounds, x i-1 is the second sub-parameter corresponding to the first facial attribute in the target facial parameter corresponding to the last round, and x t is the first sub-parameter corresponding to the first facial attribute in the initial facial parameter corresponding to the target round.
[0151] It can be seen that when the first sub-parameter is greater than the second sub-parameter, the first sub-parameter x i is closer to the maximum value x max of the first facial attribute in the target facial parameters corresponding to the rounds, the normalized position data c is closer to 1; the first sub-parameter x t is closer to the second sub-parameter x i-1 , the normalized position data c is greater; when the first sub-parameter is less than the second sub-parameter, the first sub-parameter x t is closer to the maximum value x min of the first facial attribute in the target facial parameters corresponding to the rounds, the normalized position data c is closer to 1; the first sub-parameter is closer to the second sub-parameter x i-1 , the normalized position data c is greater.
[0152] In this way, the first sub-parameter corresponding normalized position data can be accurately determined based on the data relationship between the first sub-parameter and the second sub-parameter, so that the first sub-parameter corresponding normalized position data can be accurately and efficiently determined.
[0153] After the first sub-parameter corresponding normalized position data is determined, the first sub-parameter corresponding first weight can be determined in the following manner.
[0154] The product of the normalized position data and the target editing intensity is determined as the first sub-parameter corresponding first weight.
[0155] In this embodiment, since the parameter type of the first facial attribute is continuous type, and the target editing intensity corresponding to the target round represents the influence degree of the target description information corresponding to the target round on the continuous type facial attribute when generating the target facial parameter corresponding to the target round, the first sub-parameter corresponding normalized position data can be influenced by the target editing intensity, so as to obtain the first sub-parameter corresponding first weight, and thus the third sub-parameter corresponding to the first facial attribute in the target facial parameter corresponding to the target round can be determined through the first weight.
[0156] As shown in formula (4), the formula for determining the first sub-parameter corresponding first weight through the target editing intensity corresponding to the target round and the first sub-parameter corresponding normalized position data in the data processing method provided by the embodiment of the application is as follows:
[0157] m w =c*s i Formula (4)
[0158] In the above formula (4), m w is the first sub-parameter corresponding first weight, c is the first sub-parameter corresponding normalized position data, and s i is the target editing intensity corresponding to the target round.
[0159] The first weight is obtained by influencing the first sub-parameter corresponding normalized position data through the target editing intensity, so that the third sub-parameter generated through the first weight is influenced by the target editing intensity, thereby realizing fine control of the third sub-parameter. Since the third sub-parameter is the sub-parameter corresponding to the first facial attribute in the target facial parameter corresponding to the target round, the generated target facial parameter corresponding to the target round is a facial parameter after detailed adjustment.
[0160] It should be noted that in the embodiments of the present application, the steps of generating the target description information, the target editing strength, the initial face image, the initial face parameter, and the face attribute adjustment information can be implemented through a machine model. The use and training of the optional model involved in the embodiments of the present application are described in detail as follows:
[0161] In an optional embodiment, the step of generating the target description information and the target editing strength corresponding to the target round according to the first description information and the historical information set can be implemented by the following steps:
[0162] The first description information and the historical information set are input into the pre-trained language model to output the target description information and the target editing strength corresponding to the target round through the pre-trained language model.
[0163] In the embodiment, the first description information and the historical information set can be input into the pre-trained language model, so that the first description information and the historical information set are parsed through the pre-trained language model to obtain the target description information and the target editing strength corresponding to the target round.
[0164] Optionally, the pre-trained language model can be trained in the following manner:
[0165] At least one sample corpus is obtained; wherein the sample corpus includes sample current description information, a sample historical information set, and sample target information, the sample target information including sample target description information and sample target editing strength;
[0166] The sample corpus is input into the language model to be trained, so that the language model outputs corresponding predicted target information according to the sample current description information and the sample historical description information; wherein the predicted target information includes predicted target description information and predicted target editing strength;
[0167] The model parameters of the language model are adjusted so that the difference between the predicted target information and the sample target information is less than a first threshold value, thereby obtaining the pre-trained language model.
[0168] In the embodiment, the GPT-4 or other human-computer interaction model can be used to simulate user dialogue interaction to generate one or more initial sample corpora each containing multiple rounds of dialogue interaction (such as 5 rounds, 6 rounds, or 7 rounds, etc.). For each initial sample corpus, the reply content with a quality lower than a preset quality is deleted to obtain a sample corpus. The sample corpus can include sample current description information, a sample historical information set, and sample target information, the sample target information including sample target description information and sample target editing strength.
[0169] Then, the sample corpus is input into the language model to be trained to output corresponding predicted target information including predicted target description information and predicted target editing intensity through the language model to be trained. Then, the model parameters of the language model to be trained can be adjusted based on a training strategy that the difference between the predicted target information and the sample target information is less than a first threshold, so as to obtain the pre-trained language model.
[0170] In this way, the difference between the predicted target information predicted by the pre-trained language model and the sample target information included in the sample corpus is small, so that in the inference stage, the pre-trained language model can accurately predict the target description information and the target editing intensity corresponding to the target round based on the first description information of the target round and the historical information set corresponding to the target round.
[0171] Optionally, in the embodiment, the training configuration parameters of the language model to be trained are as follows: two A100 graphics cards, Batchsize is 8, learning rate is 5e-6, AdamW optimizer is used to train 2000 steps. Specifically, the training task of the language model to be trained can be run in parallel on two A100 graphics cards. A100 is a high-performance computing GPU that can accelerate the training process. Batch size refers to the number of data samples processed before updating the model parameters, and Batchsize of 8 means that 8 sample corpora will be processed in each iteration during the training process. The learning rate determines the step size of parameter update, and the learning rate of 5e-6 means that in the gradient descent process, the parameters will move a very small step in the direction of reducing the loss function each time, which can ensure that the training process is more stable. AdamW more accurately realizes the L2 regularization effect, which helps to prevent overfitting and improve the generalization performance, and is crucial for the effective convergence of the model. Training 2000 steps means that the entire training process will perform 2000 parameter updates.
[0172] In an optional embodiment, the above step "generating an initial face image corresponding to the target round according to the target description information" can be implemented by the following steps:
[0173] inputting the target description information into the pre-trained image diffusion model to obtain an initial noisy image and a noise value corresponding to each time node through the image diffusion model;
[0174] removing the noise value corresponding to each time node in the initial noisy image in turn to obtain the initial face image corresponding to the target round.
[0175] In this embodiment, the target description information can be input into the pre-trained image diffusion model, so as to analyze the target description information through the pre-trained image diffusion model to obtain an initial noise image, which is a noise image corresponding to a first time node. Then, a first noise value corresponding to the first time node can be predicted based on the initial noise image, and the first noise value can be removed from the initial noise image to obtain a noise image corresponding to a second time node. Then, a second noise value corresponding to the second time node can be predicted based on the noise image corresponding to the second time node, and the second noise value can be removed from the noise image corresponding to the second time node to obtain a noise image corresponding to a third time node. In this way, noise values corresponding to each time node can be obtained, and the noise values corresponding to each time node can be removed in sequence, so that an initial face image corresponding to a target round can be restored. Optionally, the pre-trained image diffusion model can be trained in the following manner:
[0176] at least one sample data pair is constructed, wherein the sample data pair includes a sample face image and label information corresponding to the sample face image;
[0177] For each time node, a preset noise value is sampled through the sample face image to obtain a sample noise image;
[0178] The label information corresponding to the sample noise image is input into the image diffusion model to be trained, so that a predicted noise value is output through the image diffusion model;
[0179] The model parameters of the image diffusion model are adjusted so that the difference between the predicted noise value and the preset noise value belonging to the same time node is less than a second threshold value, thereby obtaining a pre-trained image diffusion model.
[0180] In this embodiment, one or more sample face images can be collected first, and the sample face images can be obtained by random rendering through a game engine. Then, the sample face images can be labeled, and for each sample face image, label information such as makeup, style, and attribute information of each facial attribute can be labeled, so as to construct at least one sample data pair, each sample data pair including a sample face image and label information corresponding to the sample face image.
[0181] Afterwards, the original sample face image is used to sample a corresponding preset noise value at different sampling time points (i.e., time nodes), which can be randomly obtained from a preset distribution (such as a standard normal distribution), thereby obtaining a sample noisy image corresponding to each time node. Afterwards, the label information corresponding to the sample noisy image can be input into the image diffusion model to be trained, so that the image diffusion model to be trained can predict the predicted noise value corresponding to the sample noisy image. Afterwards, based on the training strategy that the difference between the predicted noise value and the preset noise value belonging to the same time node is less than a second threshold, the model parameters of the image diffusion model to be trained are adjusted, thereby obtaining a pre-trained image diffusion model.
[0182] In this way, the difference between the predicted noise value predicted by the pre-trained image diffusion model for each time node and the true noise value is small, so that in the inference stage, the pre-trained language model can predict the noise value corresponding to each time node based on the target description information, thereby accurately determining the initial face image.
[0183] As shown in formula (5), the optimization function used by the image diffusion model in the data processing method provided by the embodiment of the application during training is as follows:
[0184]
[0185] In the above formula (5), L θ is an optimization objective function, t is a time node (i.e., a sampling time), x t is a sample noisy image, c T is label information, ∈ is an actual preset noise value, E (t, ∈, c T ) represents an expected value of the time node t, the preset noise value ∈ and the label information c θ , ω (t) is a weight function related to the time node t, ∈ θ (x t , t, c T ) represents a predicted noise value at the time node t, the sample noisy image x t and the label information c T , through the formula, the model parameters θ of the image diffusion model can be effectively optimized.
[0186] Optionally, in the embodiment, the training configuration parameters of the image diffusion model to be trained are as follows: four A100 graphics cards, Batchsize is 16, learning rate is 5e-6, AdamW optimizer is used to train 6000 steps.
[0187] In an alternative embodiment, the step of "determining initial facial parameters corresponding to the initial facial image" can be implemented by the following steps:
[0188] inputting the initial facial image into the pre-trained parameter conversion model to output the corresponding initial facial parameters through the pre-trained parameter conversion model.
[0189] In this embodiment, the initial facial image can be input into the pre-trained parameter conversion model, so that the initial facial image is converted into the corresponding initial facial parameters through the pre-trained parameter conversion model.
[0190] Optionally, the pre-trained parameter conversion model can be trained in the following way:
[0191] obtaining at least one sample facial parameter and rendering the sample facial parameter into a sample facial image;
[0192] inputting the sample facial parameter into the parameter conversion model to be trained to output a predicted facial image through the parameter conversion model;
[0193] adjusting the parameters of the parameter conversion model to make the difference between the predicted facial image and the sample facial image within a preset range, thereby obtaining the pre-trained parameter conversion model.
[0194] When training the parameter conversion model to be trained, first, at least one sample facial parameter can be obtained, and the sample facial parameter is rendered into a sample facial image.
[0195] In an alternative specific implementation, the step of "obtaining at least one sample facial parameter" includes:
[0196] obtaining at least one initial sample facial parameter;
[0197] applying random disturbance to each facial attribute of the initial sample facial parameter to obtain a disturbed sample facial parameter.
[0198] In a specific implementation, at least one initial sample facial parameter can be obtained, for example, 50,000 sets of face parameters can be collected in a game, which are the sample facial parameters, and then random disturbance is applied to each facial attribute of the initial facial parameter to obtain the sample facial parameter. For example, the size of the eyes ranges from 0 to 1, a disturbance value is randomly and uniformly sampled from -0.1 to 0.1, and the disturbance value is applied to the sub-parameter of the size of the eyes in the initial facial parameter to obtain the sample facial parameter after disturbance of the size of the eyes.
[0199] After obtaining the sample facial parameters, the sample facial parameters can be rendered into a sample facial image by a game engine. Then, the sample facial parameters are input into the parameter conversion model to be trained to output a corresponding predicted facial image by the parameter conversion model to be trained. Then, the model parameters of the parameter conversion model to be trained can be adjusted based on a training strategy that the difference between the predicted facial image and the sample facial image is within a preset range, so as to obtain the pre-trained parameter conversion model.
[0200] In this way, the pre-trained parameter conversion model has a small difference between the predicted facial image predicted based on the sample facial parameters and the sample facial image corresponding to the sample facial parameters, so that in the inference stage, the pre-trained parameter conversion model can accurately convert the initial facial image into initial facial parameters.
[0201] Optionally, in the embodiment, the training configuration parameters of the parameter conversion model to be trained are as follows: four A30 graphics cards, a Batchsize of 16, a learning rate of 1e-4, and an AdamW optimizer is used to train 100000 steps.
[0202] In an optional implementation, the step of generating the facial attribute adjustment information corresponding to the target round according to the target description information can be implemented by the following steps:
[0203] The target description information is input into the pre-trained facial attribute classification model to output the facial attribute adjustment information corresponding to the target round by the pre-trained facial attribute classification model.
[0204] In the embodiment, the target description information can be input into the pre-trained facial attribute classification model, so that the target description information is classified according to facial attributes by the pre-trained facial attribute classification model, thereby obtaining the facial attribute adjustment information corresponding to the target round.
[0205] Optionally, the pre-trained facial attribute classification model can be trained by the following steps:
[0206] At least one sample description information is obtained; wherein the sample description information is labeled with sample facial attribute adjustment information;
[0207] The sample description information is input into the facial attribute classification model to be trained to output predicted facial attribute adjustment information by the facial attribute classification model;
[0208] The model parameters of the facial attribute classification model are adjusted to make the difference between the predicted facial attribute adjustment information and the sample facial attribute adjustment information less than a third threshold value, thereby obtaining the pre-trained facial attribute classification model.
[0209] In this embodiment, first, at least one sample description information can be obtained, for example, 10,000 pieces of sample description information can be generated by ChatGPT, and the sample description information is labeled with sample facial attribute adjustment information, for example, the sample description information is that the eyes are a little bigger, and in the case that each sub-adjustment information in the facial attribute adjustment information is a binary value, the sub-adjustment information corresponding to the eyes in the sample facial attribute adjustment information corresponding to the sample description information is 1, and the sub-adjustment information corresponding to the rest of the facial attributes is 0.
[0210] Then, the sample description information can be input into the facial attribute classification model to be trained to output corresponding predicted facial attribute adjustment information through the facial attribute classification model to be trained. Then, based on the training strategy that the difference between the predicted facial attribute adjustment information and the sample facial attribute adjustment information is less than a third threshold, the model parameters of the facial attribute classification model to be trained are adjusted, thereby obtaining a pre-trained facial attribute classification model.
[0211] In this way, the difference between the predicted facial attribute adjustment information predicted by the pre-trained facial attribute classification model based on the sample description information and the real sample facial attribute adjustment information labeled by the sample description information is small, so that in the inference stage, the pre-trained facial attribute classification model can accurately classify based on the target description information to obtain the facial attribute adjustment information corresponding to the target round.
[0212] Optionally, in this embodiment, the training configuration parameters of the parameter conversion model to be trained are as follows: one 2080 graphics card, Batchsize is set to 64, learning rate is set to 3e-5, AdamW optimizer is used to train 2000 steps.
[0213] It should be noted that the above-mentioned language model to be trained, image diffusion model to be trained, parameter conversion model to be trained and training configuration parameters of the parameter conversion model to be trained are only optional examples, and in actual application, the training configuration parameters can be set according to actual situation, and the present application does not specifically limit the training configuration parameters.
[0214] For example, the training configuration parameters of the language model to be trained are as follows: one 2080 graphics card, Batchsize is set to 64, learning rate is set to 3e-5, AdamW optimizer is used to train 2000 steps. Figure 3As shown, it is the algorithm detail flow chart of the data processing method provided by the embodiment of the application. First, the first description information 20 corresponding to the target round and the historical information set 21 corresponding to the target round are input into the pre-trained language model 22, the pre-trained language model 22 outputs the target description information 23 corresponding to the target round and the target editing intensity 24, then the target description information 23 is input into the pre-trained image diffusion model 25, the pre-trained image diffusion model 25 outputs the initial face image 26 corresponding to the target round, then the initial face image 26 corresponding to the target round is input into the pre-trained parameter conversion model 27, the pre-trained parameter conversion model 27 outputs the initial face parameter 28 corresponding to the target round, then the target description information 23 is input into the pre-trained face attribute classification model 29, the pre-trained face attribute classification model 29 outputs the face attribute adjustment information 30 corresponding to the target round, in this way, according to the initial face parameter 28 of the target round, the target face parameter 31 corresponding to the last round, the face attribute adjustment information 30 corresponding to the target round, and in combination with the target editing intensity 24 corresponding to the target round, the target face parameter 32 corresponding to the target round can be obtained.
[0215] Corresponding to the data processing method provided by the first embodiment of the application, the second embodiment of the application also provides a data processing device, as shown in Figure 4 As shown, the data processing device 400 comprises:
[0216] The first generation unit 401 is configured to generate initial face parameters corresponding to a target round according to target description information corresponding to the target round, wherein the target description information is used to describe the adjustment of a face image in the target round to be performed.
[0217] The determination unit 402 is configured to determine face attribute adjustment information corresponding to the target round according to the target description information, wherein the face attribute adjustment information is used to indicate the adjustment state of each face attribute in the target round.
[0218] The second generation unit 403 is configured to generate a target face image corresponding to the target round according to the face attribute adjustment information, the initial face parameters corresponding to the target round, and the target face parameters corresponding to the last round of the target round.
[0219] Optionally, the data processing device 400 further comprises a third generation unit, and the third generation unit is configured to:
[0220] Obtain the first description information input by the target round for the face image to be generated, and the historical information set corresponding to the target round.
[0221] generate target description information corresponding to the target round and a target editing strength corresponding to the target round according to the first description information and the set of historical information, wherein the target editing strength corresponding to the target round represents a degree of influence of the target description information corresponding to the target round on the continuous type of facial attributes when generating a target facial parameter corresponding to the target round.
[0222] Optionally, the set of historical information includes historical information corresponding to at least one historical round before the target round, and the historical information includes description information input by the historical round for the face image to be generated, target description information corresponding to the historical round, and a target editing strength corresponding to the historical round.
[0223] Optionally, the first generation unit 401 is specifically configured to:
[0224] generate an initial face image corresponding to the target round according to the target description information corresponding to the target round;
[0225] determine an initial facial parameter corresponding to the initial face image.
[0226] Optionally, the second generation unit 403 is specifically configured to:
[0227] determine a target facial parameter corresponding to the target round according to the facial attribute adjustment information, the initial facial parameter corresponding to the target round, and the target facial parameter corresponding to the previous round;
[0228] generate a target face image corresponding to the target round according to the target facial parameter corresponding to the target round.
[0229] Optionally, the second generation unit 403 is specifically configured to:
[0230] determine a first facial attribute in which an adjustment state is a first state according to the facial attribute adjustment information, wherein the first state represents that the first facial attribute is a facial attribute to be adjusted in the target round;
[0231] determine a third sub-parameter of the first facial attribute corresponding to the target round according to a first sub-parameter of the first facial attribute in the initial facial parameter corresponding to the target round and a second sub-parameter of the first facial attribute in the target facial parameter corresponding to the previous round;
[0232] replace the second sub-parameter in the target facial parameter corresponding to the previous round with the third sub-parameter to obtain the target facial parameter corresponding to the target round.
[0233] Optionally, the second generation unit 403 is specifically configured to:
[0234] determine, based on a parameter type of the first facial attribute, a first weight corresponding to a first sub-parameter corresponding to the first facial attribute in the initial facial parameter corresponding to the target round, and a second weight corresponding to a second sub-parameter corresponding to the first facial attribute in the target facial parameter corresponding to the previous round; wherein a sum of the first weight and the second weight is 1;
[0235] weight the first sub-parameter and the second sub-parameter according to the first weight and the second weight to obtain a third sub-parameter of the first facial attribute corresponding to the target round.
[0236] Optionally, the second generation unit 403 is specifically configured to:
[0237] If the parameter type of the first facial attribute is a discrete type, the first weight corresponding to the first sub-parameter corresponding to the first facial attribute in the initial facial parameter corresponding to the target round is determined to be 1.
[0238] Optionally, the second generation unit 403 is specifically configured to:
[0239] If the parameter type of the first facial attribute is a continuous type, the first sub-parameter corresponding to the normalized position data is determined.
[0240] According to the normalized position data, the first weight corresponding to the first sub-parameter is determined.
[0241] Optionally, the second generation unit 403 is specifically configured to:
[0242] According to the data relationship between the first sub-parameter and the second sub-parameter, the position reference data corresponding to the first sub-parameter is determined.
[0243] According to the position reference data, the first sub-parameter and the second sub-parameter, the normalized position data corresponding to the first sub-parameter is determined.
[0244] Optionally, the second generation unit 403 is specifically configured to:
[0245] determine a first difference between the position reference data and the second sub-parameter;
[0246] determine a second difference between the first sub-parameter and the second sub-parameter;
[0247] determine, as the normalized position data corresponding to the first sub-parameter, an absolute value corresponding to a ratio of the first difference to the second difference.
[0248] Optionally, the second generation unit 403 is specifically configured to:
[0249] If the first sub-parameter is greater than the second sub-parameter, a maximum value of the first facial attribute in each target facial parameter corresponding to each round is determined as the position reference data corresponding to the first sub-parameter.
[0250] Optionally, the second generation unit 403 is specifically configured to:
[0251] If the first sub-parameter is less than the second sub-parameter, the minimum value of the first facial attribute in each target facial parameter corresponding to each round is determined as the position reference data corresponding to the first sub-parameter.
[0252] Optionally, the second generation unit 403 is specifically configured to:
[0253] The product of the normalized position data and the target editing intensity is determined as the first weight corresponding to the first sub-parameter.
[0254] Optionally, the second generation unit 403 is specifically configured to:
[0255] The first description information and the historical information set are input into the pre-trained language model to output the target description information and the target editing intensity corresponding to the target round through the pre-trained language model.
[0256] Optionally, the data processing apparatus 400 further comprises a training unit, and the training unit is configured to obtain the pre-trained language model by the following manner:
[0257] At least one sample corpus is obtained; wherein the sample corpus comprises sample current description information, a sample historical information set, and sample target information, the sample target information comprises sample target description information and sample target editing intensity;
[0258] The sample corpus is input into the language model to be trained, so that the language model outputs corresponding predicted target information according to the sample current description information and the sample historical description information; wherein the predicted target information comprises predicted target description information and predicted target editing intensity;
[0259] The model parameters of the language model are adjusted so that the difference between the predicted target information and the sample target information is less than a first threshold value, and the pre-trained language model is obtained.
[0260] Optionally, the second generation unit 403 is specifically configured to:
[0261] The target description information is input into the pre-trained image diffusion model to obtain an initial noisy image and a noise value corresponding to each time node through the image diffusion model;
[0262] The noise value corresponding to each time node is sequentially removed in the initial noisy image to obtain an initial facial image corresponding to the target round.
[0263] Optionally, the training unit is configured to obtain the pre-trained image diffusion model by the following manner:
[0264] Construct at least one sample data pair; wherein the sample data pair comprises a sample facial image and label information corresponding to the sample facial image;
[0265] For each time node, sample a preset noise value from the sample facial image to obtain a sample noisy image;
[0266] Input the label information corresponding to the sample noisy image into the image diffusion model to be trained, so as to output a predicted noise value through the image diffusion model;
[0267] Adjust the model parameters of the image diffusion model, so that the difference between the predicted noise value and the preset noise value belonging to the same time node is less than a second threshold value, to obtain a pre-trained image diffusion model.
[0268] Optionally, the second generation unit 403 is specifically configured to:
[0269] Input the initial facial image into the pre-trained parameter conversion model to output the corresponding initial facial parameters through the pre-trained parameter conversion model.
[0270] Optionally, the training unit is configured to train the pre-trained parameter conversion model in the following manner:
[0271] Obtain at least one sample facial parameter, and render the sample facial parameter into a sample facial image;
[0272] Input the sample facial parameter into the parameter conversion model to be trained to output a predicted facial image through the parameter conversion model;
[0273] Adjust the parameters of the parameter conversion model, so that the difference between the predicted facial image and the sample facial image is within a preset range, to obtain a pre-trained parameter conversion model.
[0274] Optionally, the training unit is specifically configured to:
[0275] Obtain at least one initial sample facial parameter;
[0276] Apply random disturbance to each facial attribute of the initial sample facial parameter to obtain a sample facial parameter.
[0277] Optionally, the second generation unit 403 is specifically configured to:
[0278] Input the target description information into the pre-trained facial attribute classification model to output the facial attribute adjustment information corresponding to the target round through the pre-trained facial attribute classification model.
[0279] Optionally, the training unit is configured to train the pre-trained facial attribute classification model in the following manner:
[0280] Obtaining at least one sample description information; wherein the sample description information is labeled with sample facial attribute adjustment information;
[0281] Inputting the sample description information into the facial attribute classification model to be trained to output predicted facial attribute adjustment information through the facial attribute classification model;
[0282] Adjusting the model parameters of the facial attribute classification model to make the difference between the predicted facial attribute adjustment information and the sample facial attribute adjustment information less than a third threshold value, and obtaining a pre-trained facial attribute classification model.
[0283] Corresponding to the data processing method provided by the first embodiment of the present application, the third embodiment of the present application further provides an electronic device for data processing.
[0284] As shown in Figure 5 , it is a structural block diagram of an example of the electronic device for data processing provided by the embodiments of the present application.
[0285] In the present embodiment, an optional hardware structure of the electronic device 500 can be as shown in Figure 5 , which includes at least one processor 501, at least one memory 502 and at least one communication bus 505; the memory 502 contains programs 503 and data 504.
[0286] The bus 505 can be a communication device for transmitting data between components inside the electronic device 500, such as an internal bus (for example, a CPU-memory bus, where the processor is the central processing unit, CPU), an external bus (for example, a universal serial bus port, a peripheral component interconnect express port), etc.
[0287] In addition, the electronic device further includes at least one network interface 506 and at least one peripheral interface 507. The network interface 506 provides wired or wireless communication related to an external network 508 (for example, the Internet, an intranet, a local area network, a mobile communication network, etc.); in some embodiments, the network interface 506 can include any number of network interface controllers (English: network interface controller, NIC), radio frequency (English: Radio Frequency, RF) modules, transponders, transceivers, modems, routers, gateways, any combination of wired network adapters, wireless network adapters, Bluetooth adapters, infrared adapters, near field communication (English: Near Field Communication, NFC) adapters, cellular network chips, etc.
[0288] The peripheral interface 507 is used to connect with peripherals, and the peripherals can be peripherals 1 in the figureFigure 5 509 in the middle), peripheral 2 ( Figure 5 510 in the middle) and peripheral 3 ( Figure 5 (511 in the original text). Peripherals are peripheral devices, which may include, but are not limited to, cursor control devices (such as mice, touchpads, or touchscreens), keyboards, displays (such as cathode ray tube displays, liquid crystal displays), monitors or light-emitting diode displays, video input devices (such as cameras or input interfaces coupled to video files), etc.
[0289] The processor 501 may be a CPU, an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application.
[0290] Memory 502 may include high-speed RAM (Random Access Memory) memory, and may also include non-volatile memory, such as at least one disk storage device.
[0291] Specifically, processor 501 calls the program and data stored in memory 502 to execute the following steps:
[0292] Based on the target description information corresponding to the target round, initial facial parameters corresponding to the target round are generated; wherein, the target description information is used to describe the adjustments to be made to the facial image in the target round;
[0293] Based on the target description information, facial attribute adjustment information corresponding to the target round is determined; wherein, the facial attribute adjustment information is used to indicate the adjustment status of each facial attribute in the target round;
[0294] Based on facial attribute adjustment information, the initial facial parameters corresponding to the target round, and the target facial parameters corresponding to the previous round of the target round, a target facial image corresponding to the target round is generated.
[0295] Corresponding to the data processing method provided in the first embodiment of this application, the fourth embodiment of this application provides a computer-readable storage medium storing a program for a data processing method, which is executed by a processor to perform the following steps:
[0296] Based on the target description information corresponding to the target round, initial facial parameters corresponding to the target round are generated; wherein, the target description information is used to describe the adjustments to be made to the facial image in the target round;
[0297] According to the target description information, determine face attribute adjustment information corresponding to the target round; wherein the face attribute adjustment information is used to indicate the adjustment state of each face attribute in the target round.
[0298] According to the face attribute adjustment information, the initial face parameter corresponding to the target round, and the target face parameter corresponding to the last round of the target round, generate a target face image corresponding to the target round.
[0299] It should be noted that the detailed description of the device, the electronic equipment and the computer readable storage medium provided by the second embodiment, the third embodiment and the fourth embodiment of the present application can refer to the related description of the first embodiment of the present application, which will not be repeated here.
[0300] Although the present application is disclosed with the preferred embodiments as above, it is not intended to limit the present application, and any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of the present application, therefore the protection scope of the present application should be limited by the scope defined by the claims of the present application.
[0301] In a typical configuration, a node device in a blockchain includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0302] The memory can include non-persistent memory in computer readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer readable media.
[0303] 1. Computer readable media includes permanent and non-permanent, removable and non-removable media can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other properties of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tape, magnetic tape magnetic disk storage or other magnetic storage medium or any other non-transmission medium that can be used to store information that can be accessed by a computing device. According to the definition in this paper, computer readable media does not include non-transitory computer readable media (transitory media), such as modulated data signals and carriers.
[0304] 2. Those skilled in the art will appreciate that embodiments of the present application can be readily used as a method, a system or a computer program product. Accordingly, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Furthermore, the present application can take the form of a computer program product on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, etc.) embodying computer readable program code thereon for use by or in connection with an instruction execution system. Program code embodied on a computer-usable storage medium can be transmitted using any apparatus of a transmission medium (wireline, wireless, optical fiber, cable, etc.) or embodiment combining software and hardware aspects that can be accessed directly from the storage medium or stored on a storage medium that can later be accessed by the instruction execution system.
[0305] Although the present application has been disclosed in connection with the preferred embodiments thereof, it should be understood that many modifications, variations and additions therein can be made by those skilled in the art and it is intended to cover all such modifications, variations and additions included within the spirit and scope of the application. Therefore, the scope of the present application should be determined not by the preferred embodiments but by the appended claims and their equivalents.
Claims
1. A data processing method, characterized by, The method comprises: inputting target description information corresponding to a target round into a pre-trained image diffusion model, and generating an initial face image corresponding to the target round through the pre-trained image diffusion model; wherein the target description information is used to describe the adjustment of the face image in the target round; determining initial face parameters corresponding to the initial face image; inputting the target description information into a pre-trained face attribute classification model, and outputting face attribute adjustment information corresponding to the target round through the pre-trained face attribute classification model; wherein the face attribute adjustment information is used to indicate the adjustment state of each face attribute in the target round; generating a target face image corresponding to the target round according to the face attribute adjustment information, the initial face parameters corresponding to the target round, and target face parameters corresponding to a previous round of the target round.
2. The method of claim 1, wherein, The generating of the target face image corresponding to the target round according to the face attribute adjustment information, the initial face parameters corresponding to the target round, and the target face parameters corresponding to the previous round of the target round comprises: determining target face parameters corresponding to the target round according to the face attribute adjustment information, the initial face parameters corresponding to the target round, and the target face parameters corresponding to the previous round; generating a target face image corresponding to the target round according to the target face parameters corresponding to the target round.
3. The method of claim 2, wherein, The determining of the target face parameters corresponding to the target round according to the face attribute adjustment information, the initial face parameters corresponding to the target round, and the target face parameters corresponding to the previous round comprises: determining a first face attribute in which the adjustment state is a first state according to the face attribute adjustment information; wherein the first state represents that the first face attribute is a face attribute to be adjusted in the target round; determining a third sub-parameter of the first face attribute corresponding to the target round according to a first sub-parameter of the first face attribute in the initial face parameters corresponding to the target round, and a second sub-parameter of the first face attribute in the target face parameters corresponding to the previous round; replacing the second sub-parameter in the target face parameters corresponding to the previous round with the third sub-parameter to obtain the target face parameters corresponding to the target round.
4. The method of claim 3, wherein, The determining of the third sub-parameter of the first face attribute corresponding to the target round according to the first sub-parameter of the first face attribute in the initial face parameters corresponding to the target round, and the second sub-parameter of the first face attribute in the target face parameters corresponding to the previous round comprises: determining a first weight corresponding to the first sub-parameter of the first face attribute in the initial face parameters corresponding to the target round, and a second weight corresponding to the second sub-parameter of the first face attribute in the target face parameters corresponding to the previous round based on the parameter type of the first face attribute; wherein the sum of the first weight and the second weight is 1. weighting the first sub-parameter and the second sub-parameter according to the first weight and the second weight, to obtain a third sub-parameter corresponding to the first facial attribute in the target round.
5. The method of claim 4, wherein, The method further includes: If the parameter type of the first facial attribute is a discrete type, determining that the first weight corresponding to the first sub-parameter corresponding to the first facial attribute in the initial facial parameter corresponding to the target round is 1.
6. The method of claim 4, wherein, The method further includes: If the parameter type of the first facial attribute is a continuous type, determining normalized position data corresponding to the first sub-parameter. According to the normalized position data, determining the first weight corresponding to the first sub-parameter.
7. The method of claim 6, wherein, The method further includes: According to a data relationship between the first sub-parameter and the second sub-parameter, determining position reference data corresponding to the first sub-parameter. According to the position reference data, the first sub-parameter, and the second sub-parameter, determining the normalized position data corresponding to the first sub-parameter.
8. The method of claim 7, wherein, The method further includes: Determining a first difference value between the position reference data and the second sub-parameter. Determining a second difference value between the first sub-parameter and the second sub-parameter. Determining an absolute value corresponding to a ratio of the first difference value and the second difference value as the normalized position data corresponding to the first sub-parameter.
9. The method of claim 7, wherein, The method further includes: If the first sub-parameter is greater than the second sub-parameter, determining a maximum value of the first facial attribute in each target facial parameter corresponding to each round as the position reference data corresponding to the first sub-parameter.
10. The method of claim 7, wherein, The method further includes: If the first sub-parameter is less than the second sub-parameter, determining a minimum value of the first facial attribute in each target facial parameter corresponding to each round as the position reference data corresponding to the first sub-parameter.
11. The method of claim 1, wherein, The method further includes: Before determining the initial facial parameter corresponding to the target round, the method further includes: obtaining first description information input by a target round for a to-be-generated facial image, and a historical information set corresponding to the target round; generate target description information and target editing strength corresponding to the target round according to the first description information and the set of historical information; wherein the target editing strength corresponding to the target round represents an influence degree of the target description information corresponding to the target round on a continuous type of facial attribute when generating a target facial parameter corresponding to the target round.
12. The method of claim 11, wherein, The set of historical information includes historical information corresponding to at least one historical round before the target round, and the historical information includes description information input by the historical round for a generated facial image, target description information corresponding to the historical round, and target editing strength corresponding to the historical round.
13. The method of claim 6, wherein, The first weight corresponding to the first sub-parameter is determined according to the normalized position data, comprising: The product of the normalized position data and the target editing strength corresponding to the target round is determined as the first weight corresponding to the first sub-parameter.
14. The method of claim 11, wherein, The target description information and the target editing strength corresponding to the target round are generated according to the first description information and the set of historical information, comprising: The first description information and the set of historical information are input into a pre-trained language model to output the target description information and the target editing strength corresponding to the target round through the pre-trained language model.
15. The method of claim 14, wherein, The pre-trained language model is trained in the following way: At least one sample corpus is obtained; wherein the sample corpus includes sample current description information, a set of sample historical information, and sample target information, the sample target information includes sample target description information and sample target editing strength; The sample corpus is input into a language model to be trained, so that the language model outputs corresponding predicted target information according to the sample current description information and the sample historical description information; wherein the predicted target information includes predicted target description information and predicted target editing strength; Adjust the model parameters of the language model to make the difference between the predicted target information and the sample target information less than a first threshold value, and obtain a pre-trained language model.
16. The method of claim 1, wherein, The initial facial image corresponding to the target round is generated through the pre-trained image diffusion model, comprising: An initial noisy image and a noise value corresponding to each time node are obtained through the image diffusion model; The noise value corresponding to each time node is sequentially removed in the initial noisy image to obtain the initial facial image corresponding to the target round.
17. The method of claim 16, wherein, The pre-trained image diffusion model is trained in the following way: At least one sample data pair is constructed; wherein the sample data pair includes a sample facial image and label information corresponding to the sample facial image; For each time node, a preset noise value is sampled from the sample facial image to obtain a sample noisy image; The label information corresponding to the sample noisy image is input into an image diffusion model to be trained to output a predicted noise value through the image diffusion model; Adjust a model parameter of the image diffusion model so that a difference between the predicted noise value and the preset noise value belonging to the same time node is less than a second threshold value, to obtain a pre-trained image diffusion model.
18. The method of claim 1, wherein, The determining the initial face parameter corresponding to the initial face image comprises: inputting the initial face image into a pre-trained parameter conversion model to output the corresponding initial face parameter through the pre-trained parameter conversion model.
19. The method of claim 18, wherein, The pre-trained parameter conversion model is obtained by the following method: obtain at least one sample face parameter, and render the sample face parameter into a sample face image; input the sample face parameter into a parameter conversion model to be trained to output a predicted face image through the parameter conversion model; adjust the parameter of the parameter conversion model so that the difference between the predicted face image and the sample face image is within a preset range, to obtain a pre-trained parameter conversion model.
20. The method of claim 19, wherein, The obtaining at least one sample face parameter comprises: obtain at least one initial sample face parameter; apply random disturbance to each face attribute of the initial sample face parameter to obtain a sample face parameter.
21. The method of claim 1, wherein, The pre-trained face attribute classification model is obtained by the following method: obtain at least one sample description information; wherein the sample description information is labeled with sample face attribute adjustment information; input the sample description information into a face attribute classification model to be trained to output predicted face attribute adjustment information through the face attribute classification model; adjust the model parameter of the face attribute classification model so that the difference between the predicted face attribute adjustment information and the sample face attribute adjustment information is less than a third threshold value, to obtain a pre-trained face attribute classification model.
22. A data processing apparatus, characterized in that, The device comprises: a first generation unit configured to input target description information corresponding to a target round into a pre-trained image diffusion model, generate an initial face image corresponding to the target round through the pre-trained image diffusion model, and determine an initial face parameter corresponding to the initial face image; wherein the target description information is used to describe the adjustment of the face image in the target round; and a determination unit configured to input the target description information into a pre-trained face attribute classification model, and output face attribute adjustment information corresponding to the target round through the pre-trained face attribute classification model; wherein the face attribute adjustment information is used to indicate the adjustment state of each face attribute in the target round; a second generation unit configured to generate a target face image corresponding to the target round according to the face attribute adjustment information, the initial face parameter corresponding to the target round, and a target face parameter corresponding to a previous round of the target round.
23. An electronic device, comprising: comprise: a processor; and a memory for storing a data processing program, after the electronic device is powered on and the program is run through the processor, the method of any one of claims 1-21 is executed.
24. A computer-readable storage medium, characterized in that, A data processing program is stored, which is run by the processor to execute the method of any one of claims 1-21.
Citation Information
Patent Citations
Face recognition method
CN104200194A
Speaking face video generation method and device based on multi-modal information control
CN117456587A