Data processing method and device, electronic equipment and computer readable storage medium
By adjusting details based on the initial facial parameters and the target facial parameters of the previous round, the problem of difficulty in accurately and efficiently customizing the facial images of the game characters in the prior art is solved, and more efficient and precise facial image customization is achieved.
Patent Information
- Application Number
- CN202510180039.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-18
AI Technical Summary
It is difficult to accurately and efficiently customize the facial images of game characters that meet users' expectations, especially among non-professional users. When customizing the characters by adjusting facial parameters, it is difficult and the results are unsatisfactory.
By adjusting details based on the initial facial parameters and the target facial parameters of the previous round, a facial image that meets the user's expectations is generated. The specific steps include generating initial facial parameters based on the target description information of the target round, determining the facial attribute adjustment information, and adjusting the details in combination with the target face parameters of the previous round.
Improves the efficiency and accuracy of facial image customization to ensure that the generated character facial images meet users' expectations.
Smart Images

Figure CN119971503A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a data processing method, device, electronic device and computer-readable storage medium. Background Art
[0002] With the booming development of electronic games, customizing game characters to meet the personalized needs of users has become a key part of the game. In the process of customizing game characters, users usually manually adjust complex facial parameters such as bone position and makeup color to create characters with specific appearance and style. This method is extremely time-consuming and labor-intensive, especially for non-professional users, it is more difficult to customize game characters by adjusting facial parameters.
[0003] In order to reduce the cost and difficulty of character customization, a common method in related technologies is to use natural language processing (NLP) to automatically generate corresponding characters based on text descriptions provided by users (for example, "cool boy", "more handsome", "cuter", etc.). In some methods, users can further input instructions to adjust the characters. Although this method simplifies the character customization process to a certain extent, the characters generated based on the instructions input by the user are often unsatisfactory in detail adjustment, resulting in the generated characters not meeting user expectations.
[0004] Therefore, how to accurately and efficiently customize the character's face to meet user expectations has become an urgent problem to be solved. Summary of the invention
[0005] The present application provides a data processing method, device, electronic device and computer-readable storage medium, which can adjust the details of a facial image based on the initial facial parameters and the target facial parameters of the previous round, thereby generating a facial image that meets the user's expectations and improving the efficiency and accuracy of facial image customization. The specific scheme is as follows:
[0006] In a first aspect, an embodiment of the present application provides a data processing method, the method comprising:
[0007] Generating initial facial parameters corresponding to the target round according to target description information corresponding to the target round; wherein the target description information is used to describe the adjustment to be performed on the facial image in the target round;
[0008] Determine facial attribute adjustment information corresponding to the target round according to the target description information; wherein the facial attribute adjustment information is used to indicate the adjustment status of each facial attribute in the target round;
[0009] A target facial image corresponding to the target round is generated according to the facial attribute adjustment information, the initial facial parameters corresponding to the target round, and the target facial parameters corresponding to the previous round of the target round.
[0010] In a second aspect, an embodiment of the present application provides a data processing device, the device comprising:
[0011] A first generating unit, configured to generate initial facial parameters corresponding to the target round according to target description information corresponding to the target round; wherein the target description information is used to describe the adjustment to be performed on the facial image in the target round;
[0012] A determination unit, configured to determine facial attribute adjustment information corresponding to the target round according to the target description information; wherein the facial attribute adjustment information is used to indicate an adjustment state of each facial attribute in the target round;
[0013] The second generating unit is used to generate a target facial image corresponding to the target round according to the facial attribute adjustment information, the initial facial parameters corresponding to the target round and the target facial parameters corresponding to the previous round of the target round.
[0014] In a third aspect, the present application further provides an electronic device, including:
[0015] Processor; and
[0016] The memory is used to store a data processing program. After the electronic device is powered on and the program is run by the processor, the method of the first aspect is executed.
[0017] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium storing a data processing program, which is executed by a processor to perform the method of the first aspect.
[0018] Compared with the prior art, this application has the following advantages:
[0019] The data processing method provided by the embodiment of the present application includes the following steps: generating initial facial parameters corresponding to the target round according to target description information corresponding to the target round; wherein the target description information is used to describe the adjustment to be performed on the facial image in the target round; determining facial attribute adjustment information corresponding to the target round according to the target description information; wherein the facial attribute adjustment information is used to indicate the adjustment status of each facial attribute in the target round; and generating a target facial image corresponding to the target round according to the facial attribute adjustment information, the initial facial parameters corresponding to the target round, and the target facial parameters corresponding to the previous round of the target round.
[0020] It can be seen that the initial facial parameters generated by the target description information can feedback the facial features of the facial image to be generated, and the facial attribute adjustment information determined by the target description information accurately indicates the adjustment status of each facial attribute in the target round. In this way, based on the facial attribute adjustment information and the initial facial parameters of the target round, combined with the target facial parameters of the previous round, detail adjustments can be made in the target round to generate a target facial image after detail adjustment. Therefore, the data processing method provided in the embodiment of the present application can make detail adjustments to the facial image based on the initial facial parameters and the target facial parameters of the previous round, thereby generating a facial image that meets the user's expectations and improving the efficiency and accuracy of facial image customization. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 is a data processing system diagram for implementing a data processing method provided in an embodiment of the present application;
[0022] Figure 2 is a flow chart of the data processing method provided by the present application;
[0023] Figure 3 is a flowchart of the algorithm details of the data processing method provided in the embodiment of the present application;
[0024] Figure 4 is a structural block diagram of an example of a data processing device provided in an embodiment of the present application;
[0025] Figure 5 It is a structural block diagram of an example of an electronic device for data processing provided in an embodiment of the present application. DETAILED DESCRIPTION
[0026] Many specific details are described in the following description to facilitate a full understanding of the present application. However, the present application can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the connotation of the present application, so the present application is not limited by the specific implementation disclosed below.
[0027] It should be noted that the terms "first", "second", "third", etc. in the claims, description and drawings of the present application are used to distinguish similar objects and are not used to describe a specific order or sequence. The data used in this way are interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including", "having" and their variants are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0028] It should be understood that in the embodiments of the present application, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" is merely a way to describe the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the previous and subsequent associated objects are in an "or" relationship. "Including A, B and / or C" means including any one, any two, or any three of A, B, and C.
[0029] It should be understood that in the embodiments of the present application, "B corresponding to A", "B corresponding to A", "A corresponds to B", or "B corresponds to A" means that B is associated with A, and B can be determined according to A. Determining B according to A does not mean determining B only according to A, and B can also be determined according to A and / or other information.
[0030] Before describing the implementation methods of the present application in detail, the prior art will be further described first.
[0031] In role-playing games (RPG), augmented reality (AR) and virtual reality (VR), creating game characters that meet specific needs for users has become an important research direction. With the continuous development of the game industry and the advancement of technology, users have a growing demand for creating personalized characters that meet specific aesthetics or styles. However, achieving personalized character customization often involves complex parameter adjustments, including but not limited to the adjustment of parameters such as facial bone position, skin texture, and makeup color, which not only requires users to master professional knowledge, but is also quite time-consuming and laborious.
[0032] In order to reduce the manpower and time costs of role customization, the industry often uses the following methods to customize roles:
[0033] The first is to customize characters based on reference images. This method allows users to upload a reference image, and the system will automatically adjust the character's facial features based on the reference image to generate a character that is highly similar to the facial features of the reference image.
[0034] The second is to customize characters based on text descriptions. This method allows users to enter a text description, and the system will generate a character based on the text description, such as "cool boy" or "cuter", to generate a character that matches the text description.
[0035] However, the above method can only generate characters once, and lacks the ability to iterate the generated characters, which means that once the initially generated characters fail to fully meet the user's expectations, it is difficult for the user to make subtle adjustments, thus affecting the final effect. In addition, it is difficult for users to make specific modifications to the characters in the above method, especially when pursuing details. This makes it particularly difficult to achieve complex or abstract style descriptions, especially for ordinary players without professional knowledge backgrounds.
[0036] Based on the above reasons, in order to be able to flexibly adjust the details of facial images through multiple rounds of conversations, thereby generating facial images that meet user expectations and improving the efficiency and accuracy of facial image customization, the first embodiment of the present application provides a data processing method, which is applied to an electronic device. The electronic device can be a desktop computer, a laptop computer, a mobile phone, a tablet computer, an electronic watch, etc., or other electronic devices capable of data processing, which is not specifically limited in the embodiments of the present application.
[0037] The data processing method provided in the embodiment of the present application can be applied to the face-pinching scene in the game, so that a facial image that meets the user's expectations can be generated in the game; the data processing method can also be applied to the scene of chatting with an intelligent body model in human-computer interaction, so that a facial image that meets the user's expectations can be generated in human-computer interaction.
[0038] In an optional embodiment, when the data processing method is run on a terminal device, the terminal device may include a display screen and a processor, and the display screen is used to present an interactive screen and receive description information input by a user. The interactive screen may include an information interaction area, in which the user can input description information, and in which the target facial image generated in each round can be displayed. The processor is used to store a data processing program, run the program, generate an interactive screen, respond to instructions, and control the display of the interactive screen on the display screen. When the user operates the interactive screen through the display screen, the interactive screen can control the local content of the terminal device by responding to the received operation instructions. The terminal device may provide a graphical user interface to the user in a variety of ways, for example, it may be rendered and displayed on the display screen of the terminal device, or the graphical user interface may be presented through holographic projection.
[0039] In an optional embodiment, when the digital processing method is run on a server, the method can be implemented and executed based on a cloud system. The cloud system is based on cloud computing. The cloud system includes a server and a client device. The operating body of the data processing program and the client interface presentation body are separated, and the storage and operation of the data processing method are completed on the server. The client interface presentation is completed on the client, and the client is mainly used for receiving and sending data and presenting the client interface. For example, the client can be a display device with a data transmission function close to the user side, such as a mobile terminal, a television, a computer, a handheld computer, a personal digital assistant, a head-mounted display device, etc., but the terminal device for data processing is a server in the cloud. During the application period corresponding to the data processing program, the user instructs the client to send instructions to the server, and the server encodes and compresses the client interface and other data according to the instructions, and returns them to the client through the network. Finally, the client decodes and outputs the client interface.
[0040] It should be noted that in the embodiments of the present application, the execution subject of the data processing method can be a terminal device or a server, wherein the terminal device can be a local terminal device or a client device in the aforementioned cloud system. The embodiments of the present application do not limit the type of the execution subject.
[0041] For example, in combination with the above introduction, Figure 1 A data processing system 100 for implementing a data processing method provided by an embodiment of the present application is shown, and the data processing system 100 may include at least one terminal 101, at least one server 102, and a network 103. The terminal 101 held by the user may be connected to the server 102 via a network. The terminal is any device with computing hardware that can support and execute software application tools corresponding to the intelligent physical therapy device.
[0042] In the above data processing system 100, the terminal 101 is used to install and run the application corresponding to the above data processing program. In some cases, the application may not be installed in advance in the terminal 101, and the user may directly access the application through a client such as a browser. In the process of the user accessing the application through the application, data can be exchanged between the terminal 101 and the server 102. The terminal 101 sends various information to the server 102. The server 102 determines the display data of the terminal 101 according to the received information, and sends the display data to the terminal 101, so that the display data sent by the server 102 can be displayed to the user through the terminal 101.
[0043] In a possible application scenario, different terminals 101 may be served by different servers 102 , and the servers 102 corresponding to different terminals 101 may be the same server.
[0044] In addition, when the data processing system 100 includes multiple terminals, multiple servers, and multiple networks, different terminals can be connected to each other through different networks and different servers.
[0045] Among them, the terminal 101 can have one or more multi-touch screens for sensing and obtaining the input of the user through touch or sliding operations performed at multiple points on one or more touch display screens. The terminal 101 can also be connected to a keyboard and / or a mouse and / or a game controller, so that the user can perform interface operations through the keyboard and / or mouse and / or a game controller.
[0046] The network may be a wireless network or a wired network, such as a wireless network such as a wireless local area network (WLAN), a local area network (LAN), a cellular network, a 2G network, a 3G network, a 4G network, a 5G network, etc. In addition, different terminals may also use their own Bluetooth network or hotspot network to connect to other terminals or to a server, etc. In addition, the system 100 may include multiple databases, which are coupled to different servers, and the data generated by the application corresponding to the above data processing program may be continuously stored in the database.
[0047] It should be noted that Figure 1 The data processing system diagram shown is merely an example. The data processing system 100 described in the embodiment of the present application is intended to more clearly illustrate the technical solution of the embodiment of the present application, and does not constitute a limitation on the technical solution provided in the embodiment of the present application.
[0048] The technical solution of the present application is described in detail below through specific embodiments. It should be noted that the following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0049] The following, combined Figure 2 and Figure 3 The data processing method provided in the embodiment of the present application is introduced.
[0050] like Figure 2 As shown, it is a flow chart of the data processing method provided by the present application, including the following steps S103 to S105.
[0051] It should be noted that, before step S103, the data processing method provided by the present application may further include the following steps S101 and S102:
[0052] Step S101: Acquire first description information inputted for a facial image to be generated in a target round, and a set of historical information corresponding to the target round.
[0053] In an embodiment of the present application, when generating an expected facial image, multiple rounds of interaction may be performed, and the target round refers to any round other than the first round in the multiple rounds of interaction. The first description information refers to the information input for the facial image to be generated in the target round, and the first description information represents the adjustment requirements for the facial image to be generated in the target round. The first description information may be description information in text form or in voice form, and the present application does not limit this.
[0054] The historical information set corresponding to the target round may be an information set consisting of historical information corresponding to one or more historical rounds before the target round. For example, if the target round is the 5th round, the historical information set corresponding to the target round may be an information set consisting of historical information corresponding to the 4th round, or may be an information set consisting of historical information corresponding to the 1st to 4th rounds.
[0055] Optionally, the historical information corresponding to the historical round may at least include description information input for the facial image to be generated in the historical round.
[0056] Step S102: Generate target description information corresponding to the target round according to the first description information and the historical information set; wherein the target description information corresponding to the target round is used to describe the adjustment to be performed on the facial image in the target round.
[0057] It should be noted that, in addition to generating the target description information corresponding to the target round according to the first description information and the historical information set, the editing strength corresponding to the target round can also be generated. The target editing strength corresponding to the target round represents the degree of influence of the target description information corresponding to the target round on the continuous type of facial attributes when generating the target facial parameters corresponding to the target round.
[0058] This step is used to parse the first description information currently input in the target round and the historical information set corresponding to the target round to obtain target description information for describing the adjustment to be performed on the facial image in the target round.
[0059] It should be noted that the above-mentioned first description information is the adjustment requirement for the facial image to be generated input by the user in the target round. The first description information is usually colloquial information, and the first description information does not necessarily include the facial attributes to be adjusted or the adjustment method of the facial attributes to be adjusted; the target description information is the information obtained by parsing the first description information based on the historical information set, and is used to feedback the adjustments to be made to the facial image in the target round based on the combination of the information in the historical information set and the input text of the first description information. The target description information includes the facial attributes to be adjusted and the adjustment method of the facial attributes to be adjusted.
[0060] For example, assuming that the historical information of a certain historical round in the historical information set includes "adjust the bridge of the nose", and the first description information of the target round is "the bridge of the nose is not adjusted enough", then the target description information of the target round is "adjust the bridge of the nose", indicating that the bridge of the nose of the facial image needs to be further adjusted higher on the basis of the previous round. Assuming that the historical information of the previous round in the historical information set is "enlarge the eyes", and the first description information of the target round is "enlarge again", then the target description information of the target round is "enlarge the eyes", indicating that the eyes of the facial image need to be further adjusted higher on the basis of the previous round.
[0061] It should be noted that when the first description information includes the facial attributes to be adjusted and the adjustment method of the facial attributes to be adjusted, the target description information of the target round is usually consistent with the first description information. Assuming that the first description information of the target round is "eye enlargement", the target description information is "eye enlargement", indicating that the eyes of the facial image need to be further enlarged on the basis of the previous round.
[0062] In this step, the first description information and the historical information set are parsed, and while generating the target description information corresponding to the target round, a target editing strength can also be generated. The target editing strength is used to characterize the degree of influence of the target description information corresponding to the target round on the continuous type of facial attributes when generating the target facial parameters corresponding to the target round.
[0063] In the embodiment of the present application, the target editing strength generated for the target round may be within the target interval [0, 1].
[0064] When the target editing strength is 0, it means that the parameters corresponding to the facial attributes of the continuous type do not change, and the parameters corresponding to the facial attributes of the continuous type in the target facial parameters corresponding to the target round continue to use the parameters corresponding to the facial attributes of the continuous type in the target facial parameters of the previous round. When the target editing strength is 1, it means that the parameters corresponding to the facial attributes of the continuous type are completely affected by the target description information, and the parameters corresponding to the facial attributes of the continuous type in the target facial parameters corresponding to the target round are the parameters corresponding to the facial attributes of the continuous type in the initial facial parameters of the target round.
[0065] When the target editing intensity is at an intermediate value between 0 and 1, it means that the parameters corresponding to the continuous type of facial attributes are affected by both the parameters corresponding to the continuous type of facial attributes in the target facial parameters of the previous round and the target description information. Then, the parameters corresponding to the continuous type of facial attributes in the target facial parameters corresponding to the target round are the parameters obtained by interpolating and fusion the parameters corresponding to the continuous type of facial attributes in the initial facial parameters of the target round and the parameters corresponding to the continuous type of facial attributes in the target facial parameters of the previous round.
[0066] In this way, through the target editing intensity, the parameters of the continuous type can be finely adjusted to further improve the accuracy of facial image customization.
[0067] The specific way in which the target editing strength affects the continuous type of facial attributes will be introduced in detail in the subsequent steps.
[0068] In this step, target description information and target editing strength are generated through the first description information and the historical information set, and the historical rounds of the target rounds are fully considered. The generated target description information includes the facial attributes to be adjusted and the adjustment method of the facial attributes to be adjusted. The generated target editing strength combines the understanding of the context and accurately indicates the degree of influence of the target description information on the continuous type of facial attributes, which provides a good foundation for the accurate and efficient generation of facial images that meet user expectations.
[0069] Step S103: generating initial facial parameters corresponding to the target round according to the target description information corresponding to the target round; wherein the target description information is used to describe the adjustment to be performed on the facial image in the target round.
[0070] The above initial facial parameters can be used to describe the facial features corresponding to the target description information, and the facial features may include features corresponding to various facial attributes. Various facial attributes may include facial attributes of the facial features and bones dimension, facial attributes of the makeup position and shape dimension, and facial attributes of the makeup color dimension. Among them, the facial attributes of the facial features and bones dimension may include but are not limited to forehead, temple, chin, cheekbone, cheek, apple muscle, jaw, eyebrow, eye, nose, mouth, ear; the facial attributes of the makeup position and shape dimension may include but are not limited to blush, eyebrows, eyeliner, eye shadow, eyelashes, eyelids, lips, skin, facial features, scars, iris, upper beard, lower beard; the facial attributes of the makeup color dimension may include but are not limited to blush color, facial line color, eyebrow color, eye color, eye shadow color, eyeliner color, lip color, beard color, etc.
[0071] In an optional implementation, the step of generating the initial facial parameters corresponding to the target round according to the target description information may be implemented by the following steps:
[0072] Generate an initial facial image corresponding to the target round according to the target description information;
[0073] Initial facial parameters corresponding to the initial facial image are determined.
[0074] In this embodiment, an initial facial image corresponding to the target round can be generated according to the target description information. The initial facial image can feedback the facial features corresponding to the target description information, that is, the facial attributes to be adjusted included in the target description information in the initial facial image are adjusted according to the adjustment method included in the target description information; thereafter, the initial facial image is converted into initial facial parameters in parameter form.
[0075] It should be noted that the initial facial image of the target round can be an image that has no association with the target facial image corresponding to the historical round. For example, if the target description information of the target round is that the eyes are larger, then the initial facial image of the target round is an image with eyes larger than the preset size. Other facial attributes such as eyebrows, nose, and mouth cannot be guaranteed to be exactly the same as the target facial image corresponding to the historical round.
[0076] Step S104: determining facial attribute adjustment information corresponding to the target round according to the target description information; wherein the facial attribute adjustment information is used to indicate the adjustment status of each facial attribute in the target round.
[0077] The above-mentioned facial attribute adjustment information is used to indicate the adjustment status of each facial attribute in the target round, and the adjustment status can represent whether the corresponding facial attribute has been adjusted. Specifically, the facial attribute adjustment information can include sub-adjustment information corresponding to each facial attribute; wherein the sub-adjustment information includes first adjustment information or second adjustment information, the first adjustment information is used to indicate that the corresponding facial attribute is in a first state to be adjusted in the target round, and the second adjustment information is used to indicate that the corresponding facial attribute is in a second state that is not adjusted in the target round.
[0078] For each facial attribute, the corresponding sub-adjustment information can be a binary value. When the binary value is 0, it indicates that the corresponding facial attribute will not be adjusted in the target round. When the binary value is 1, it indicates that the corresponding facial attribute will be adjusted in the target round. That is, the first adjustment information is 1 and the second adjustment information is 0.
[0079] Exemplarily, the facial attribute adjustment information includes five bits, the first bit is hair color, the second bit is eye color, the third bit is skin color, the fourth bit is smile degree, and the fifth bit is eyebrow sparseness. Assuming that the facial attribute adjustment information of the target round is 01001, it means that in the target round, the hair color needs to be adjusted, the eye color and skin color are not adjusted, the smile degree is adjusted, and the eyebrow sparseness is not adjusted.
[0080] In this step, by parsing the target description information, classifying it according to facial attributes, and obtaining facial attribute adjustment information on whether each facial attribute has been modified, it is possible to clarify which facial attributes in the target round will be modified compared to the previous round, and which facial attributes will not be modified compared to the previous round, thereby accurately generating the target facial image of the target round.
[0081] Step S105: Generate a target facial image corresponding to the target round according to the facial attribute adjustment information, the initial facial parameters corresponding to the target round, and the target facial parameters corresponding to the previous round of the target round.
[0082] After determining the facial attribute adjustment information corresponding to the target round and the initial facial parameters corresponding to the target round, the target facial image corresponding to the target round can be generated in combination with the target facial parameters corresponding to the previous round of the target round.
[0083] In an optional implementation, the above step S105 can be implemented by the following steps:
[0084] Determining target facial parameters corresponding to the target round according to the facial attribute adjustment information, the initial facial parameters corresponding to the target round, and the target facial parameters corresponding to the previous round;
[0085] According to the target facial parameters corresponding to the target round, a target facial image corresponding to the target round is generated.
[0086] It should be noted that for each round other than the first round, there are initial facial parameters and target facial parameters. The target facial parameters of each round are based on the facial attribute adjustment information of this round and are generated by the initial facial parameters of this round and the target facial parameters of the previous round.
[0087] In this way, the target facial parameters corresponding to the target round can be determined based on the facial attribute adjustment information corresponding to the target round, the initial facial parameters corresponding to the target round and the target facial parameters corresponding to the previous round.
[0088] Since the initial facial parameters of the target round are used to describe the facial features corresponding to the target description information, the facial attribute adjustment information of the target round is used to indicate the adjustment status of each facial attribute in the target round. In this way, combined with the target facial parameters of the previous round, detail adjustments can be made in the target round to generate target facial parameters after detail adjustments.
[0089] Optionally, after the target facial parameters after detail adjustment are generated, the target facial parameters may be rendered by a game engine to generate a target facial image after detail adjustment.
[0090] Optionally, after generating the target facial parameters after detail adjustment, the target facial parameters may be converted into a target facial image through an image generation model.
[0091] It should be noted that after the target facial image is generated, the target facial image may be displayed in a graphical user interface so that the user can determine whether the target facial image of the target round meets expectations.
[0092] The data processing method provided by the embodiment of the present application includes the following steps: generating initial facial parameters corresponding to the target round according to target description information corresponding to the target round; wherein the target description information is used to describe the adjustment to be performed on the facial image in the target round; determining facial attribute adjustment information corresponding to the target round according to the target description information; wherein the facial attribute adjustment information is used to indicate the adjustment status of each facial attribute in the target round; and generating a target facial image corresponding to the target round according to the facial attribute adjustment information, the initial facial parameters corresponding to the target round, and the target facial parameters corresponding to the previous round of the target round.
[0093] It can be seen that the initial facial parameters generated by the target description information can feedback the facial features of the facial image to be generated, and the facial attribute adjustment information determined by the target description information accurately indicates the adjustment status of each facial attribute in the target round. In this way, based on the facial attribute adjustment information and the initial facial parameters of the target round, combined with the target facial parameters of the previous round, detail adjustments can be made in the target round to generate a target facial image after detail adjustment. Therefore, the data processing method provided in the embodiment of the present application can make detail adjustments to the facial image based on the initial facial parameters and the target facial parameters of the previous round, thereby generating a facial image that meets the user's expectations and improving the efficiency and accuracy of facial image customization.
[0094] In an optional embodiment, the above-mentioned historical information set includes historical information corresponding to at least one historical round before the above-mentioned target round. For a historical round, its corresponding historical information includes the description information of the historical round for the facial image to be generated, the target description information corresponding to the historical round, and the target editing intensity corresponding to the historical round.
[0095] As shown in the following formula (1), it is an expression of the historical information set in the data processing method provided in the embodiment of the present application:
[0096] H i ={τ i-m ,(t i-m ,s i-m ),…,τ i-1 ,(t i-1 ,s i-1 )} Formula (1)
[0097] In the above formula (1), m is the number of historical rounds selected, H i is the historical information set of the i-th round, 1≤m<i, τ i-m is the description information of the facial image to be generated in the im-th round, t i-m is the target description information corresponding to the im-th round, s i-m is the target editing strength corresponding to the im-th round, τ i-1 is the description information of the facial image to be generated in the i-1th round, t i-1 is the target description information corresponding to the i-1th round, s i-1 is the target editing intensity corresponding to the i-1th round.
[0098] Exemplarily, when i is 10 and m is 5, the historical information set corresponding to the 10th round includes the description information of the facial image input to be generated in the 5th round, the target description information corresponding to the 5th round, the target editing strength corresponding to the 5th round, the description information of the facial image input to be generated in the 6th round, the target description information corresponding to the 6th round, the target editing strength corresponding to the 6th round, the description information of the facial image input to be generated in the 7th round, the target description information corresponding to the 7th round, the target editing strength corresponding to the 7th round, the description information of the facial image input to be generated in the 8th round, the target description information corresponding to the 8th round, the target editing strength corresponding to the 8th round, the description information of the facial image input to be generated in the 9th round, the target description information corresponding to the 9th round, and the target editing strength corresponding to the 9th round.
[0099] The following is a detailed introduction to determining the target facial parameters corresponding to the target round:
[0100] In an optional implementation, the above step of “determining the target facial parameters corresponding to the target round according to the facial attribute adjustment information, the initial facial parameters corresponding to the target round, and the target facial parameters corresponding to the previous round” may specifically include the following steps:
[0101] Determine, according to the facial attribute adjustment information, a first facial attribute whose adjustment state is a first state among the facial attributes; wherein the first state indicates that the first facial attribute is a facial attribute to be adjusted in the target round;
[0102] Determine a third sub-parameter corresponding to the first facial attribute in the target round according to the first sub-parameter corresponding to the first facial attribute in the initial facial parameters corresponding to the target round and the second sub-parameter corresponding to the first facial attribute in the target facial parameters corresponding to the previous round;
[0103] The second sub-parameter in the target facial parameters corresponding to the previous round is replaced by the third sub-parameter to obtain the target facial parameters corresponding to the target round.
[0104] In this embodiment, each facial attribute can be divided, specifically, into a first facial attribute whose adjustment state is a first state and a second facial attribute whose adjustment state is a second state, wherein the first state represents that the first facial attribute is a facial attribute to be adjusted in the target round, and the second state represents that the second facial attribute is a facial attribute that is not adjusted in the target round. As can be seen from the above introduction, for each facial attribute, the corresponding sub-adjustment information can be a binary value, so that a facial attribute with a binary value of 1 can be determined as a first facial attribute, and a facial attribute with a binary value of 2 can be determined as a second facial attribute.
[0105] Afterwards, the first sub-parameter corresponding to the first facial attribute in the initial facial parameters corresponding to the target round and the second sub-parameter corresponding to the first facial attribute in the target facial parameters corresponding to the previous round can be obtained, so as to determine the third sub-parameter corresponding to the first facial attribute in the target round based on the first sub-parameter and the second sub-parameter, and the third sub-parameter is the sub-parameter corresponding to the first facial attribute in the target facial parameters corresponding to the target round.
[0106] In this way, the second sub-parameter in the target facial parameters corresponding to the previous round can be replaced by the third sub-parameter, thereby obtaining the target facial parameters corresponding to the target round.
[0107] It should be noted that since the second facial attribute is a facial attribute that does not undergo adjustment in the target round, the sub-parameter corresponding to the second facial attribute in the target facial parameters corresponding to the target round still uses the sub-parameter corresponding to the second facial attribute in the target facial parameters corresponding to the previous round, thereby maintaining the second facial attribute without adjustment in the target round.
[0108] This setting method determines the first facial attribute that has changed in the target round, only updates the first facial attribute in the target facial parameters of the previous round, and keeps the second facial attribute in the target facial parameters of the previous round unchanged, thereby achieving detailed adjustment of the facial image.
[0109] In an optional implementation, the above step of “determining a third sub-parameter corresponding to the first facial attribute in the target round according to the first sub-parameter corresponding to the first facial attribute in the initial facial parameters corresponding to the target round and the second sub-parameter corresponding to the first facial attribute in the target facial parameters corresponding to the previous round” may include the following steps S201 and S202:
[0110] Step S201: Based on the parameter type of the first facial attribute, determine a first weight corresponding to a first sub-parameter corresponding to the first facial attribute in the initial facial parameters corresponding to the target round, and a second weight corresponding to a second sub-parameter corresponding to the first facial attribute in the target facial parameters corresponding to the previous round; wherein the sum of the first weight and the second weight is 1;
[0111] Step S202: weighting the first sub-parameter and the second sub-parameter according to the first weight and the second weight to obtain a third sub-parameter corresponding to the first facial attribute in the target round.
[0112] It should be noted that the facial attributes described above can be divided into continuous and discrete types according to the parameter type. Discrete facial attributes refer to facial attributes with a fixed number of parameters, which are usually used to describe features with clear classification or selectivity; continuous facial parameters refer to facial attributes with parameters that can vary arbitrarily within a certain range, which are usually used to describe features that can be measured by specific values and can vary smoothly within a certain range.
[0113] For example, discrete facial attributes may include makeup type, expression state, skin tone, etc. Among them, the parameters of makeup type may include natural makeup, heavy makeup, smoky makeup, nude makeup, etc. Each makeup style is a clear category, rather than a smooth changing process; expression state may include smile, frown, surprise, anger, etc., each expression state is a clear type; skin tone may include white, yellow, black, etc.
[0114] For example, continuous facial attributes may include eye distance, nose bridge height, apple cheek size, etc. The parameters of continuous facial attributes vary continuously within a certain range without obvious boundaries.
[0115] In this implementation, the first weight corresponding to the first sub-parameter corresponding to the first facial attribute in the initial facial parameters corresponding to the target round can be determined based on the parameter type of the first facial attribute, and then the difference between 1 and the first weight can be used as the second weight corresponding to the second sub-parameter corresponding to the first facial attribute in the target facial parameters corresponding to the previous round.
[0116] The first weight is the influence of the first sub-parameter on the third sub-parameter in the target facial parameters corresponding to the target round to be generated, and the second weight is the influence of the second sub-parameter on the third sub-parameter in the target facial parameters corresponding to the target round to be generated.
[0117] Afterwards, the first sub-parameter and the second sub-parameter may be weighted and summed according to the first weight and the second weight, thereby obtaining a third sub-parameter corresponding to the first facial attribute in the target round.
[0118] As shown in the following formula (2), it is the calculation formula of the third sub-parameter in the data processing method provided in the embodiment of the present application:
[0119] x i =x t *m w +x i-1 *(1-m w ) Formula (2)
[0120] In the above formula (2), xi is the third sub-parameter corresponding to the first facial attribute in the target round, x t is the first sub-parameter corresponding to the first facial attribute in the initial facial parameters corresponding to the target round, m w is the first weight, x i-1 is the second sub-parameter corresponding to the first facial attribute in the target facial parameter corresponding to the previous round, 1-m w is the second weight.
[0121] For discrete first facial attributes and continuous first facial attributes, different parameter types correspond to different calculation methods of the first weight. The following introduces the calculation methods of the first weight under different parameter types:
[0122] Optionally, the above step of “determining, based on the parameter type of the first facial attribute, a first weight corresponding to a first sub-parameter corresponding to the first facial attribute in the initial facial parameters corresponding to the target round” may include the following steps:
[0123] If the parameter type of the first facial attribute is a discrete type, it is determined that a first weight corresponding to a first sub-parameter corresponding to the first facial attribute in the initial facial parameters corresponding to the target round is 1.
[0124] If the parameter type of the first facial attribute is a discrete type, such as the makeup type, expression state, skin tone, etc. mentioned above, the first weight corresponding to the first sub-parameter corresponding to the first facial attribute in the initial facial parameters corresponding to the target round can be determined to be 1, and accordingly, the second weight corresponding to the second sub-parameter corresponding to the first facial attribute in the target facial parameters corresponding to the previous round is 0.
[0125] That is to say, for the first facial attribute to be adjusted in the target round, when the parameter type of the first facial attribute is a discrete type, the third sub-parameter corresponding to the first facial attribute in the target facial parameters corresponding to the target round is the first sub-parameter corresponding to the first facial attribute in the initial facial parameters corresponding to the target round.
[0126] Optionally, the above step of “determining, based on the parameter type of the first facial attribute, a first weight corresponding to a first sub-parameter corresponding to the first facial attribute in the initial facial parameters corresponding to the target round” may include the following steps:
[0127] If the parameter type of the first facial attribute is a continuous type, determining normalized position data corresponding to the first sub-parameter;
[0128] A first weight corresponding to the first sub-parameter is determined according to the normalized position data.
[0129] If the first facial attribute is of a continuous type, such as the distance between the eyes, the height of the bridge of the nose, the size of the apple cheeks, etc. mentioned above, the normalized position data of the first sub-parameter corresponding to the first facial attribute in the initial facial parameters corresponding to the target round can be determined first. The normalized position data is used to indicate the relative change of the first facial attribute in the target round.
[0130] In an optional specific implementation, the normalized position data may be determined in the following manner:
[0131] Determine the position reference data corresponding to the first sub-parameter according to the data relationship between the first sub-parameter and the second sub-parameter;
[0132] The normalized position data corresponding to the first sub-parameter is determined according to the position reference data, the first sub-parameter and the second sub-parameter.
[0133] In this implementation, a data relationship between a first sub-parameter corresponding to a first facial attribute in the initial facial parameters corresponding to the target round and a second sub-parameter corresponding to the first facial attribute in the target facial parameters corresponding to the previous round can be determined first. The data relationship is used to characterize the change of the first facial attribute in the target round relative to the previous round. The data relationship may include: a first data relationship in which the first sub-parameter is greater than the second sub-parameter, or a second data relationship in which the first sub-parameter is less than the second sub-parameter. The first data relationship is used to characterize the change of the first facial attribute in the target round relative to the previous round as a parameter increase, and the second data relationship is used to characterize the change of the first facial attribute in the target round relative to the previous round as a parameter decrease.
[0134] In this implementation manner, the position reference data corresponding to the first sub-parameter may be determined according to the data relationship between the first sub-parameter and the second sub-parameter, and the position reference data is reference data for normalized position data.
[0135] Under different data relationships, the above step of "determining the position reference data corresponding to the first sub-parameter according to the data relationship between the first sub-parameter and the second sub-parameter" may include the following two situations:
[0136] Case 1: if the first sub-parameter is greater than the second sub-parameter, the maximum value of the first facial attribute in each target facial parameter corresponding to each round is determined as the position reference data corresponding to the first sub-parameter.
[0137] In case 1, the data relationship between the first sub-parameter and the second sub-parameter is the above-mentioned first data relationship. In this case, the maximum value of the first facial attribute in each target facial parameter corresponding to each round can be determined, and the maximum value is determined as the position reference data corresponding to the first sub-parameter. For example, if a total of 10 rounds are currently performed, the maximum value of the first facial attribute in the sub-parameters of each target facial parameter corresponding to the 10 rounds can be determined first, and the maximum value is the position reference data corresponding to the first sub-parameter.
[0138] Case 2: if the first sub-parameter is smaller than the second sub-parameter, the minimum value of the first facial attribute among the target facial parameters corresponding to each round is determined as the position reference data corresponding to the first sub-parameter.
[0139] In case 2, the data relationship between the first sub-parameter and the second sub-parameter is the second data relationship described above. In this case, the minimum value of the first facial attribute in each target facial parameter corresponding to each round can be determined, and the minimum value is determined as the position reference data corresponding to the first sub-parameter. For example, if a total of 10 rounds are currently performed, the minimum value of the first facial attribute in the sub-parameters of each target facial parameter corresponding to the 10 rounds can be determined first, and the minimum value is the position reference data corresponding to the first sub-parameter.
[0140] After the position reference data corresponding to the first sub-parameter is determined, the normalized position data corresponding to the first sub-parameter may be determined based on the position parameter data.
[0141] Specifically, the normalized position data can be determined by the following steps:
[0142] determining a first difference between the position reference data and the second sub-parameter;
[0143] determining a second difference between the first sub-parameter and the second sub-parameter;
[0144] An absolute value corresponding to the ratio of the first difference to the second difference is determined as normalized position data corresponding to the first sub-parameter.
[0145] In this embodiment, the first difference between the position reference data and the second sub-parameter can be determined first. Specifically, if the first sub-parameter is greater than the second sub-parameter, the first difference between the maximum value of the first facial attribute in each target facial parameter corresponding to each round and the second sub-parameter can be determined; if the first sub-parameter is less than the second sub-parameter, the first difference between the minimum value of the first facial attribute in each target facial parameter corresponding to each round and the second sub-parameter can be determined.
[0146] Afterwards, a second difference between a first sub-parameter corresponding to the first facial attribute in the initial facial parameters corresponding to the target round and a second sub-parameter corresponding to the first facial attribute in the target facial parameters corresponding to the previous round is determined.
[0147] Finally, the ratio between the first difference and the second difference is determined, and the absolute value corresponding to the ratio is determined as the normalized position data corresponding to the first parameter. It should be noted that when the first sub-parameter is greater than the second sub-parameter, the normalized position data is used to characterize the relative position of the first sub-parameter relative to the maximum value of the first facial attribute in each target facial parameter corresponding to each round and the second sub-parameter; when the first sub-parameter is less than the second sub-parameter, the normalized position data is used to characterize the relative position of the first sub-parameter relative to the minimum value of the first facial attribute in each target facial parameter corresponding to each round and the second sub-parameter.
[0148] As shown in the following formula (3), it is the calculation formula corresponding to the normalized position data in the data processing method provided in the embodiment of the present application:
[0149]
[0150] In the above formula (3), c is the normalized position data corresponding to the first sub-parameter, x max is the maximum value of the first facial attribute among the target facial parameters corresponding to each round, x min The minimum value of the first facial attribute among the target facial parameters corresponding to each round, x i-1 is the second sub-parameter corresponding to the first facial attribute in the target facial parameter corresponding to the previous round, x t It is the first sub-parameter corresponding to the first facial attribute in the initial facial parameters corresponding to the target round.
[0151] It can be seen that when the first sub-parameter is greater than the second sub-parameter, the first sub-parameter x i The closer to the maximum value x of the first facial attribute in each target facial parameter corresponding to each round max , the closer the normalized position data c is to 1; the first sub-parameter x t The closer to the second subparameter x i-1 , the larger the normalized position data c is; when the first sub-parameter is smaller than the second sub-parameter, the first sub-parameter x t The closer to the maximum value x of the first facial attribute in each target facial parameter corresponding to each round min , the closer the normalized position data c is to 1; the first sub-parameter The closer to the second subparameter x i-1 , the larger the normalized position data c is.
[0152] This setting method can accurately determine the position reference data corresponding to the first sub-parameter based on the data relationship between the first sub-parameter and the second sub-parameter, thereby accurately and efficiently determining the normalized position data corresponding to the first sub-parameter.
[0153] After the normalized position data corresponding to the first sub-parameter is determined, the first weight corresponding to the first sub-parameter may be determined in the following manner;
[0154] The product of the normalized position data and the target editing strength is determined as the first weight corresponding to the first sub-parameter.
[0155] In the present embodiment, since the parameter type of the first facial attribute is a continuous type, and the target editing strength corresponding to the target round represents the degree of influence of the target description information corresponding to the target round on the continuous type of facial attributes when generating the target facial parameters corresponding to the target round, the normalized position data corresponding to the first sub-parameter can be influenced by the target editing strength to obtain the first weight corresponding to the first sub-parameter. In this way, the third sub-parameter corresponding to the first facial attribute in the target facial parameters corresponding to the target round can be determined through the first weight.
[0156] As shown in formula (4), in the data processing method provided in the embodiment of the present application, the calculation formula for determining the first weight corresponding to the first sub-parameter by the target editing strength corresponding to the target round and the normalized position data corresponding to the first sub-parameter is:
[0157] m w =c*s i Formula (4)
[0158] In the above formula (4), m w is the first weight corresponding to the first sub-parameter, c is the normalized position data corresponding to the first sub-parameter, s i Edit the intensity for the target corresponding to the target round.
[0159] The normalized position data corresponding to the first sub-parameter is influenced by the target editing strength to obtain the first weight. In this way, the third sub-parameter generated by the first weight is influenced by the target editing strength, thereby achieving refined control of the third sub-parameter. Since the third sub-parameter is the sub-parameter corresponding to the first facial attribute in the target facial parameter corresponding to the target round, the target facial parameter generated corresponding to the target round is the facial parameter after detail adjustment.
[0160] It should be noted that, in the embodiments of the present application, the steps of generating target description information, target editing strength, initial facial image, initial facial parameters and facial attribute adjustment information can all be implemented by a machine model. The use and training of the optional models involved in the embodiments of the present application are described in detail below:
[0161] In an optional implementation, the step of generating target description information and target editing strength corresponding to the target round according to the first description information and the historical information set may be implemented by the following steps:
[0162] The first description information and the historical information set are input into a pre-trained language model to output target description information and target editing strength corresponding to the target round through the pre-trained language model.
[0163] In this implementation, the first description information and the historical information set may be input into a pre-trained language model, so that the first description information and the historical information set are parsed by the pre-trained language model to obtain target description information and target editing strength corresponding to the target round.
[0164] Optionally, the above pre-trained language model can be trained in the following way:
[0165] Acquire at least one sample corpus; wherein the sample corpus includes sample current description information, sample history information set and sample target information, and the sample target information includes sample target description information and sample target editing strength;
[0166] Inputting the sample corpus into the language model to be trained, so that the language model outputs corresponding prediction target information according to the current description information of the sample and the historical description information of the sample; wherein the prediction target information includes the prediction target description information and the prediction target editing strength;
[0167] The model parameters of the language model are adjusted so that the difference between the predicted target information and the sample target information is less than a first threshold, thereby obtaining a pre-trained language model.
[0168] In this embodiment, a human-computer interaction model such as GPT-4 can be used to simulate user dialogue interaction, generate one or more initial sample corpora each containing multiple rounds of dialogue interaction (such as 5 rounds, 6 rounds, or 7 rounds, etc.), and for each initial sample corpus, delete the reply content with a quality lower than the preset quality to obtain a sample corpus. The sample corpus may include the current description information of the sample, the sample history information set, and the sample target information, and the sample target information includes the sample target description information and the sample target editing strength.
[0169] Afterwards, the sample corpus is input into the language model to be trained, so that the language model to be trained outputs corresponding prediction target information, wherein the prediction target information includes prediction target description information and prediction target editing strength. Afterwards, based on the training strategy that the difference between the prediction target information and the sample target information is less than a first threshold, the model parameters of the language model to be trained can be adjusted to obtain a pre-trained language model.
[0170] In this way, the difference between the predicted target information predicted by the pre-trained language model and the sample target information included in the sample corpus is small. In this way, in the reasoning stage, the pre-trained language model can accurately predict the target description information and target editing intensity corresponding to the target round based on the first description information of the target round and the historical information set corresponding to the target round.
[0171] Optionally, in this embodiment, the training configuration parameters of the language model to be trained are as follows: two A100 graphics cards, Batchsize of 8, learning rate of 5e-6, and training for 2000 steps using the AdamW optimizer. Specifically, the training task of the language model to be trained can be run in parallel on two A100 graphics cards. A100 is a high-performance computing GPU that can accelerate the training process. Batch size refers to the number of data samples processed before updating the model parameters. Batchsize of 8 means that 8 sample corpora will be processed in each iteration during the training process. The learning rate determines the step size of the parameter update. A learning rate of 5e-6 means that during the gradient descent process, each time the parameters are updated, they will move a very small step in the direction of reducing the loss function, which can ensure that the training process is more stable. AdamW more accurately implements the L2 regularization effect, helps prevent overfitting and improves generalization performance, and is crucial to whether the model can converge effectively. Training for 2000 steps means that the entire training process will perform 2000 parameter updates.
[0172] In an optional implementation, the above step of "generating an initial facial image corresponding to the target round according to the target description information" can be implemented by the following steps:
[0173] Input the target description information into the pre-trained image diffusion model to obtain the initial noisy image and the noise value corresponding to each time node through the image diffusion model;
[0174] The noise value corresponding to each time node is removed in turn from the initial noisy image to obtain the initial facial image corresponding to the target round.
[0175] In this embodiment, the target description information can be input into a pre-trained image diffusion model, so that the target description information is parsed by the pre-trained image diffusion model to obtain an initial noise image, which is also the noise image corresponding to the first time node. Then, based on the initial noise image, the first noise value corresponding to the first time node can be predicted. By removing the first noise value from the initial noise image, the noise image corresponding to the second time node can be obtained. Then, based on the noise image corresponding to the second time node, the second noise value corresponding to the second time node can be predicted. By removing the second noise value from the noise image corresponding to the second time node, the noise image corresponding to the third time node can be obtained. And so on, the noise value corresponding to each time node can be obtained, and the noise value corresponding to each time node can be removed in turn, so as to restore the initial facial image corresponding to the target round. Optionally, the above-mentioned pre-trained image diffusion model can be trained in the following way:
[0176] Constructing at least one sample data pair; wherein the sample data pair includes a sample facial image and annotation information corresponding to the sample facial image;
[0177] For each time node, a preset noise value is sampled from a sample facial image to obtain a sample noisy image;
[0178] Inputting the annotation information corresponding to the sample noisy image into the image diffusion model to be trained, so as to output a predicted noise value through the image diffusion model;
[0179] The model parameters of the image diffusion model are adjusted so that the difference between the predicted noise value and the preset noise value belonging to the same time node is smaller than the second threshold value, thereby obtaining a pre-trained image diffusion model.
[0180] In this embodiment, first, one or more sample facial images can be collected, and the sample facial images can be randomly rendered by a game engine. After that, the sample facial images can be annotated, and the annotation information such as makeup, style, and attribute information corresponding to each facial attribute of each sample facial image can be obtained, so as to construct at least one sample data pair, and each sample data pair includes a sample facial image and the annotation information corresponding to the sample facial image.
[0181] Afterwards, the original sample facial image is used to sample the corresponding preset noise values at different sampling moments (i.e., time nodes), and the preset noise values can be randomly drawn from a preset distribution (such as a standard normal distribution), thereby obtaining the sample noisy images corresponding to each time node. Afterwards, the annotation information corresponding to the sample noisy image can be input into the image diffusion model to be trained, so as to predict the predicted noise value corresponding to the sample noisy image through the image diffusion model to be trained. Afterwards, based on the training strategy that the difference between the predicted noise value and the preset noise value belonging to the same time node is less than the second threshold, the model parameters of the image diffusion model to be trained can be adjusted, thereby obtaining a pre-trained image diffusion model.
[0182] In this way, the difference between the predicted noise value predicted by the pre-trained image diffusion model for each time node and the actual noise value is small. In this way, in the inference stage, the pre-trained language model can predict the noise value corresponding to each time node based on the target description information, thereby accurately determining the initial facial image.
[0183] As shown in formula (5), it is the optimization function used in the training of the image diffusion model in the data processing method provided in the embodiment of the present application:
[0184]
[0185] In the above formula (5), L θ is the optimization objective function, t is the time node (i.e. sampling time), x t is the sample noise image, c T is the label information, ∈ is the actual preset noise value, Represents the time node t, the preset noise value ∈ and the annotation information c T The expected value of ω(t) is the weight function related to time node t, ∈ θ represents the image diffusion model, ∈ θ (x t ,t,c T ) indicates that at time node t, the sample noise image x t And the annotation information c T The predicted noise value under this condition can be effectively optimized through this formula.
[0186] Optionally, in this embodiment, the training configuration parameters of the image diffusion model to be trained are as follows: four A100 graphics cards, Batchsize of 16, learning rate of 5e-6, and training for 6000 steps using the AdamW optimizer.
[0187] In an optional implementation, the above step of “determining the initial facial parameters corresponding to the initial facial image” can be implemented by the following steps:
[0188] The initial facial image is input into the pre-trained parameter conversion model to output corresponding initial facial parameters through the pre-trained parameter conversion model.
[0189] In this implementation, the initial facial image may be input into a pre-trained parameter conversion model, so that the initial facial image may be converted into corresponding initial facial parameters through the pre-trained parameter conversion model.
[0190] Optionally, the above pre-trained parameter conversion model can be trained in the following way:
[0191] Obtain at least one sample facial parameter, and render the sample facial parameter into a sample facial image;
[0192] Inputting sample facial parameters into a parameter conversion model to be trained, so as to output a predicted facial image through the parameter conversion model;
[0193] The parameters of the parameter conversion model are adjusted so that the difference between the predicted facial image and the sample facial image is within a preset range, thereby obtaining a pre-trained parameter conversion model.
[0194] When training the parameter conversion model to be trained, first, at least one sample facial parameter may be obtained, and the sample facial parameter may be rendered as a sample facial image.
[0195] In an optional specific implementation, the step of "obtaining at least one sample facial parameter" includes:
[0196] obtaining at least one initial sample facial parameter;
[0197] Random perturbations are applied to the facial attributes of the initial sample facial parameters to obtain perturbed sample facial parameters.
[0198] In a specific implementation, at least one initial sample facial parameter can be obtained. For example, 50,000 sets of facial parameters can be collected in the game. The 50,000 sets of facial parameters are sample facial parameters. Then, random perturbations are applied to the facial attributes of the initial facial parameters to obtain sample facial parameters. For example, the range of eye size is 0 to 1, a perturbation value is randomly and uniformly sampled from -0.1 to 0.1, and the perturbation value is applied to the sub-parameter of eye size in the initial facial parameters to obtain sample facial parameters after eye size perturbation.
[0199] After obtaining the sample facial parameters, the sample facial parameters can be rendered into a sample facial image through a game engine. Afterwards, the sample facial parameters are input into a parameter conversion model to be trained, so as to output a corresponding predicted facial image through the parameter conversion model to be trained. Afterwards, based on a training strategy in which the difference between the predicted facial image and the sample facial image is within a preset range, the model parameters of the parameter conversion model to be trained can be adjusted, thereby obtaining a pre-trained parameter conversion model.
[0200] In this way, the difference between the predicted facial image predicted by the pre-trained parameter conversion model based on the sample facial parameters and the sample facial image corresponding to the sample facial parameters is small. In this way, in the inference stage, the pre-trained parameter conversion model can accurately convert the initial facial image into the initial facial parameters.
[0201] Optionally, in this embodiment, the training configuration parameters of the parameter conversion model to be trained are as follows: four A30 graphics cards, Batchsize of 16, learning rate of 1e-4, and training for 100,000 steps using the AdamW optimizer.
[0202] In an optional implementation, the step of generating facial attribute adjustment information corresponding to the target round according to the target description information may be implemented by the following steps:
[0203] The target description information is input into a pre-trained facial attribute classification model to output facial attribute adjustment information corresponding to the target round through the pre-trained facial attribute classification model.
[0204] In this implementation, the target description information may be input into a pre-trained facial attribute classification model, so that the target description information is classified according to facial attributes by the pre-trained facial attribute classification model, thereby obtaining facial attribute adjustment information corresponding to the target round.
[0205] Optionally, the pre-trained facial attribute classification model can be trained by the following steps:
[0206] Acquire at least one sample description information; wherein the sample description information is annotated with sample facial attribute adjustment information;
[0207] Inputting the sample description information into the facial attribute classification model to be trained, so as to output predicted facial attribute adjustment information through the facial attribute classification model;
[0208] The model parameters of the facial attribute classification model are adjusted so that the difference between the predicted facial attribute adjustment information and the sample facial attribute adjustment information is smaller than a third threshold, thereby obtaining a pre-trained facial attribute classification model.
[0209] In this implementation, first, at least one sample description information can be obtained. For example, ten thousand sample description information can be generated through ChatGPT. The sample description information is marked with sample facial attribute adjustment information. For example, the sample description information is that the eyes are bigger. When each sub-adjustment information in the facial attribute adjustment information is a binary value, the sub-adjustment information corresponding to the eyes in the sample facial attribute adjustment information corresponding to the sample description information is 1, and the sub-adjustment information corresponding to the remaining facial attributes is 0.
[0210] Afterwards, the sample description information can be input into the facial attribute classification model to be trained, so that the corresponding predicted facial attribute adjustment information can be outputted by the facial attribute classification model to be trained. Afterwards, based on the training strategy that the difference between the predicted facial attribute adjustment information and the sample facial attribute adjustment information is less than the third threshold, the model parameters of the facial attribute classification model to be trained can be adjusted, thereby obtaining a pre-trained facial attribute classification model.
[0211] In this way, the difference between the predicted facial attribute adjustment information predicted by the pre-trained facial attribute classification model based on the sample description information and the actual sample facial attribute adjustment information marked by the sample description information is small. In this way, in the reasoning stage, the pre-trained facial attribute classification model can accurately classify based on the target description information and obtain the facial attribute adjustment information corresponding to the target round.
[0212] Optionally, in this embodiment, the training configuration parameters of the parameter conversion model to be trained are as follows: a 2080 graphics card, Batchsize is set to 64, learning rate is set to 3e-5, and AdamW optimizer is used for training for 2000 steps.
[0213] It should be noted that the language model to be trained, the image diffusion model to be trained, the parameter conversion model to be trained, and the training configuration parameters of the parameter conversion model to be trained introduced above are only optional examples. In practical applications, they can be specifically set according to actual conditions. This application does not specifically limit the training configuration parameters.
[0214] like Figure 3As shown, it is an algorithm detail flow chart of the data processing method provided by the embodiment of the present application. First, the first description information 20 corresponding to the target round and the historical information set 21 corresponding to the target round are input into the pre-trained language model 22, and the pre-trained language model 22 outputs the target description information 23 and the target editing strength 24 corresponding to the target round. Then, the target description information 23 is input into the pre-trained image diffusion model 25, and the pre-trained image diffusion model 25 outputs the initial facial image 26 corresponding to the target round. Then, the initial facial image 26 corresponding to the target round is input into the pre-trained language model 22. In the parameter conversion model 27, the pre-trained parameter conversion model 27 outputs the initial facial parameters 28 corresponding to the target round. After that, the target description information 23 is input into the pre-trained facial attribute classification model 29, and the pre-trained facial attribute classification model 29 outputs the facial attribute adjustment information 30 corresponding to the target round. In this way, according to the initial facial parameters 28 of the target round, the target facial parameters 31 corresponding to the previous round, the facial attribute adjustment information 30 corresponding to the target round, and combined with the target editing strength 24 corresponding to the target round, the target facial parameters 32 corresponding to the target round can be obtained.
[0215] Corresponding to the data processing method provided in the first embodiment of the present application, the second embodiment of the present application further provides a data processing device, such as Figure 4 As shown, the data processing device 400 includes:
[0216] The first generating unit 401 is used to generate initial facial parameters corresponding to the target round according to target description information corresponding to the target round; wherein the target description information is used to describe the adjustment to be performed on the facial image in the target round;
[0217] A determination unit 402 is used to determine facial attribute adjustment information corresponding to the target round according to the target description information; wherein the facial attribute adjustment information is used to indicate the adjustment status of each facial attribute in the target round;
[0218] The second generating unit 403 is used to generate a target facial image corresponding to the target round according to the facial attribute adjustment information, the initial facial parameters corresponding to the target round and the target facial parameters corresponding to the previous round of the target round.
[0219] Optionally, the data processing device 400 further includes a third generating unit, and the third generating unit is configured to:
[0220] Acquire first description information inputted for a facial image to be generated in a target round, and a set of historical information corresponding to the target round;
[0221] According to the first description information and the historical information set, the target description information and target editing strength corresponding to the target round are generated; wherein the target editing strength corresponding to the target round represents the degree of influence of the target description information corresponding to the target round on the continuous type of facial attributes when generating the target facial parameters corresponding to the target round.
[0222] Optionally, the historical information set includes historical information corresponding to at least one historical round before the target round, and the historical information includes description information input for the facial image to be generated in the historical round, target description information corresponding to the historical round, and target editing strength corresponding to the historical round.
[0223] Optionally, the first generating unit 401 is specifically configured to:
[0224] Generate an initial facial image corresponding to the target round according to the target description information corresponding to the target round;
[0225] Initial facial parameters corresponding to the initial facial image are determined.
[0226] Optionally, the second generating unit 403 is specifically configured to:
[0227] Determining target facial parameters corresponding to the target round according to the facial attribute adjustment information, the initial facial parameters corresponding to the target round, and the target facial parameters corresponding to the previous round;
[0228] According to the target facial parameters corresponding to the target round, a target facial image corresponding to the target round is generated.
[0229] Optionally, the second generating unit 403 is specifically configured to:
[0230] Determine, according to the facial attribute adjustment information, a first facial attribute whose adjustment state is a first state among the facial attributes; wherein the first state indicates that the first facial attribute is a facial attribute to be adjusted in the target round;
[0231] Determine a third sub-parameter corresponding to the first facial attribute in the target round according to the first sub-parameter corresponding to the first facial attribute in the initial facial parameters corresponding to the target round and the second sub-parameter corresponding to the first facial attribute in the target facial parameters corresponding to the previous round;
[0232] The second sub-parameter in the target facial parameters corresponding to the previous round is replaced by the third sub-parameter to obtain the target facial parameters corresponding to the target round.
[0233] Optionally, the second generating unit 403 is specifically configured to:
[0234] Based on the parameter type of the first facial attribute, determine a first weight corresponding to a first sub-parameter corresponding to the first facial attribute in the initial facial parameters corresponding to the target round, and a second weight corresponding to a second sub-parameter corresponding to the first facial attribute in the target facial parameters corresponding to the previous round; wherein the sum of the first weight and the second weight is 1;
[0235] The first sub-parameter and the second sub-parameter are weighted according to the first weight and the second weight to obtain a third sub-parameter corresponding to the first facial attribute in the target round.
[0236] Optionally, the second generating unit 403 is specifically configured to:
[0237] If the parameter type of the first facial attribute is a discrete type, it is determined that a first weight corresponding to a first sub-parameter corresponding to the first facial attribute in the initial facial parameters corresponding to the target round is 1.
[0238] Optionally, the second generating unit 403 is specifically configured to:
[0239] If the parameter type of the first facial attribute is a continuous type, determining normalized position data corresponding to the first sub-parameter;
[0240] A first weight corresponding to the first sub-parameter is determined according to the normalized position data.
[0241] Optionally, the second generating unit 403 is specifically configured to:
[0242] Determine the position reference data corresponding to the first sub-parameter according to the data relationship between the first sub-parameter and the second sub-parameter;
[0243] The normalized position data corresponding to the first sub-parameter is determined according to the position reference data, the first sub-parameter and the second sub-parameter.
[0244] Optionally, the second generating unit 403 is specifically configured to:
[0245] determining a first difference between the position reference data and the second sub-parameter;
[0246] determining a second difference between the first sub-parameter and the second sub-parameter;
[0247] An absolute value corresponding to the ratio of the first difference to the second difference is determined as normalized position data corresponding to the first sub-parameter.
[0248] Optionally, the second generating unit 403 is specifically configured to:
[0249] If the first sub-parameter is greater than the second sub-parameter, the maximum value of the first facial attribute in each target facial parameter corresponding to each round is determined as the position reference data corresponding to the first sub-parameter.
[0250] Optionally, the second generating unit 403 is specifically configured to:
[0251] If the first sub-parameter is smaller than the second sub-parameter, the minimum value of the first facial attribute among the target facial parameters corresponding to each round is determined as the position reference data corresponding to the first sub-parameter.
[0252] Optionally, the second generating unit 403 is specifically configured to:
[0253] The product of the normalized position data and the target editing strength is determined as the first weight corresponding to the first sub-parameter.
[0254] Optionally, the second generating unit 403 is specifically configured to:
[0255] The first description information and the historical information set are input into a pre-trained language model to output target description information and target editing strength corresponding to the target round through the pre-trained language model.
[0256] Optionally, the data processing device 400 further includes a training unit, which is used to obtain a pre-trained language model by training in the following manner:
[0257] Acquire at least one sample corpus; wherein the sample corpus includes sample current description information, sample history information set and sample target information, and the sample target information includes sample target description information and sample target editing strength;
[0258] Inputting the sample corpus into the language model to be trained, so that the language model outputs corresponding prediction target information according to the current description information of the sample and the historical description information of the sample; wherein the prediction target information includes the prediction target description information and the prediction target editing strength;
[0259] The model parameters of the language model are adjusted so that the difference between the predicted target information and the sample target information is less than a first threshold, thereby obtaining a pre-trained language model.
[0260] Optionally, the second generating unit 403 is specifically configured to:
[0261] Input the target description information into the pre-trained image diffusion model to obtain the initial noisy image and the noise value corresponding to each time node through the image diffusion model;
[0262] The noise value corresponding to each time node is removed in turn from the initial noisy image to obtain the initial facial image corresponding to the target round.
[0263] Optionally, the training unit is used to train a pre-trained image diffusion model in the following manner:
[0264] Constructing at least one sample data pair; wherein the sample data pair includes a sample facial image and annotation information corresponding to the sample facial image;
[0265] For each time node, a preset noise value is sampled from a sample facial image to obtain a sample noisy image;
[0266] Inputting the annotation information corresponding to the sample noisy image into the image diffusion model to be trained, so as to output a predicted noise value through the image diffusion model;
[0267] The model parameters of the image diffusion model are adjusted so that the difference between the predicted noise value and the preset noise value belonging to the same time node is smaller than the second threshold value, thereby obtaining a pre-trained image diffusion model.
[0268] Optionally, the second generating unit 403 is specifically configured to:
[0269] The initial facial image is input into the pre-trained parameter conversion model to output corresponding initial facial parameters through the pre-trained parameter conversion model.
[0270] Optionally, the training unit is used to train a pre-trained parameter conversion model in the following way:
[0271] Obtain at least one sample facial parameter, and render the sample facial parameter into a sample facial image;
[0272] Inputting sample facial parameters into a parameter conversion model to be trained, so as to output a predicted facial image through the parameter conversion model;
[0273] The parameters of the parameter conversion model are adjusted so that the difference between the predicted facial image and the sample facial image is within a preset range, thereby obtaining a pre-trained parameter conversion model.
[0274] Optionally, the training unit is specifically used for:
[0275] obtaining at least one initial sample facial parameter;
[0276] Random perturbations are applied to the facial attributes of the initial sample facial parameters to obtain sample facial parameters.
[0277] Optionally, the second generating unit 403 is specifically configured to:
[0278] The target description information is input into a pre-trained facial attribute classification model to output facial attribute adjustment information corresponding to the target round through the pre-trained facial attribute classification model.
[0279] Optionally, the training unit is used to train a pre-trained facial attribute classification model in the following manner:
[0280] Acquire at least one sample description information; wherein the sample description information is annotated with sample facial attribute adjustment information;
[0281] Inputting the sample description information into the facial attribute classification model to be trained, so as to output predicted facial attribute adjustment information through the facial attribute classification model;
[0282] The model parameters of the facial attribute classification model are adjusted so that the difference between the predicted facial attribute adjustment information and the sample facial attribute adjustment information is smaller than a third threshold, thereby obtaining a pre-trained facial attribute classification model.
[0283] Corresponding to the data processing method provided in the first embodiment of the present application, the third embodiment of the present application also provides an electronic device for data processing.
[0284] like Figure 5 , which is a structural block diagram of an example of an electronic device for data processing provided in an embodiment of the present application.
[0285] In this embodiment, an optional hardware structure of the electronic device 500 can be as follows: Figure 5 As shown, it includes: at least one processor 501, at least one memory 502 and at least one communication bus 505; the memory 502 contains a program 503 and data 504.
[0286] The bus 505 may be a communication device for transmitting data between components inside the electronic device 500, such as an internal bus (e.g., a CPU-memory bus, where the processor is a central processing unit, CPU for short), an external bus (e.g., a universal serial bus port, a peripheral component interconnect express port), etc.
[0287] In addition, the electronic device also includes: at least one network interface 506, at least one peripheral interface 507. The network interface 506 provides wired or wireless communication related to an external network 508 (for example, the Internet, an intranet, a local area network, a mobile communication network, etc.); in some embodiments, the network interface 506 may include any number of network interface controllers (NICs), radio frequency (RF) modules, repeaters, transceivers, modems, routers, gateways, any combination of wired network adapters, wireless network adapters, Bluetooth adapters, infrared adapters, near field communication (NFC) adapters, cellular network chips, etc.
[0288] The peripheral interface 507 is used to connect to the peripheral device, and the peripheral device can be the peripheral device 1 ( Figure 5 509), peripheral 2 ( Figure 5 510) and peripheral 3 ( Figure 5 511 in the figure). Peripherals are peripheral devices, which may include but are not limited to cursor control devices (such as mice, touch pads or touch screens), keyboards, displays (such as cathode ray tube displays, liquid crystal displays), displays or light emitting diode displays, video input devices (such as cameras or communication interfaces coupled to video files), etc.
[0289] The processor 501 may be a CPU, or an application specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present application.
[0290] The memory 502 may include a high-speed RAM (full name: Random Access Memory) memory, and may also include a non-volatile memory, such as at least one disk memory.
[0291] The processor 501 calls the program and data stored in the memory 502 to perform the following steps:
[0292] Generating initial facial parameters corresponding to the target round according to target description information corresponding to the target round; wherein the target description information is used to describe the adjustment to be performed on the facial image in the target round;
[0293] Determine facial attribute adjustment information corresponding to the target round according to the target description information; wherein the facial attribute adjustment information is used to indicate the adjustment status of each facial attribute in the target round;
[0294] A target facial image corresponding to the target round is generated according to the facial attribute adjustment information, the initial facial parameters corresponding to the target round, and the target facial parameters corresponding to the previous round of the target round.
[0295] Corresponding to the data processing method provided in the first embodiment of the present application, the fourth embodiment of the present application provides a computer-readable storage medium storing a program of the data processing method, which is executed by a processor to perform the following steps:
[0296] Generating initial facial parameters corresponding to the target round according to target description information corresponding to the target round; wherein the target description information is used to describe the adjustment to be performed on the facial image in the target round;
[0297] Determine facial attribute adjustment information corresponding to the target round according to the target description information; wherein the facial attribute adjustment information is used to indicate the adjustment status of each facial attribute in the target round;
[0298] A target facial image corresponding to the target round is generated according to the facial attribute adjustment information, the initial facial parameters corresponding to the target round, and the target facial parameters corresponding to the previous round of the target round.
[0299] It should be noted that for the detailed description of the devices, electronic devices and computer-readable storage media provided in the second, third and fourth embodiments of the present application, reference can be made to the relevant description of the first embodiment of the present application, and no further details will be given here.
[0300] Although the present application is disclosed as above in the form of a preferred embodiment, it is not intended to limit the present application. Any technical personnel in this field may make possible changes and modifications without departing from the spirit and scope of the present application. Therefore, the scope of protection of the present application shall be based on the scope defined by the claims of the present application.
[0301] In a typical configuration, a node device in a blockchain includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0302] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0303] 1. Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), random access memory (RAM) of other attributes, read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage media or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined in this article, computer-readable media does not include non-transitory media such as modulated data signals and carrier waves.
[0304] 2. Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0305] Although the present application is disclosed as above in the form of a preferred embodiment, it is not intended to limit the present application. Any technical personnel in this field may make possible changes and modifications without departing from the spirit and scope of the present application. Therefore, the scope of protection of the present application shall be based on the scope defined by the claims of the present application.
Claims
1. A data processing method, characterized in that: The method comprises: Generating initial facial parameters corresponding to the target round according to target description information corresponding to the target round; wherein the target description information is used to describe the adjustment to be performed on the facial image in the target round; Determine facial attribute adjustment information corresponding to the target round according to the target description information; wherein the facial attribute adjustment information is used to indicate the adjustment status of each facial attribute in the target round; A target facial image corresponding to the target round is generated according to the facial attribute adjustment information, the initial facial parameters corresponding to the target round, and the target facial parameters corresponding to the previous round of the target round.
2. The method according to claim 1, characterized in that: The step of generating initial facial parameters corresponding to the target round according to the target description information corresponding to the target round includes: Generating an initial facial image corresponding to the target round according to the target description information corresponding to the target round; Determine initial facial parameters corresponding to the initial facial image.
3. The method according to claim 1, characterized in that The step of generating a target facial image corresponding to the target round according to the facial attribute adjustment information, the initial facial parameters corresponding to the target round, and the target facial parameters corresponding to the previous round of the target round includes: Determining target facial parameters corresponding to the target round according to the facial attribute adjustment information, the initial facial parameters corresponding to the target round, and the target facial parameters corresponding to the previous round; A target facial image corresponding to the target round is generated according to the target facial parameters corresponding to the target round.
4. The method according to claim 3, characterized in that The step of determining the target facial parameters corresponding to the target round according to the facial attribute adjustment information, the initial facial parameters corresponding to the target round, and the target facial parameters corresponding to the previous round includes: Determine, according to the facial attribute adjustment information, a first facial attribute whose adjustment state is a first state among the facial attributes; wherein the first state indicates that the first facial attribute is a facial attribute to be adjusted in the target round; Determine a third sub-parameter corresponding to the first facial attribute in the target round according to the first sub-parameter corresponding to the first facial attribute in the initial facial parameters corresponding to the target round and the second sub-parameter corresponding to the first facial attribute in the target facial parameters corresponding to the previous round; The second sub-parameter in the target facial parameters corresponding to the previous round is replaced by the third sub-parameter to obtain the target facial parameters corresponding to the target round.
5. The method according to claim 4, characterized in that The determining, according to the first sub-parameter corresponding to the first facial attribute in the initial facial parameters corresponding to the target round and the second sub-parameter corresponding to the first facial attribute in the target facial parameters corresponding to the previous round, a third sub-parameter corresponding to the first facial attribute in the target round includes: Based on the parameter type of the first facial attribute, determining a first weight corresponding to a first sub-parameter corresponding to the first facial attribute in the initial facial parameters corresponding to the target round, and a second weight corresponding to a second sub-parameter corresponding to the first facial attribute in the target facial parameters corresponding to the previous round; wherein the sum of the first weight and the second weight is 1; The first sub-parameter and the second sub-parameter are weighted according to the first weight and the second weight to obtain a third sub-parameter corresponding to the first facial attribute in the target round.
6. The method according to claim 5, characterized in that The determining, based on the parameter type of the first facial attribute, a first weight corresponding to a first sub-parameter corresponding to the first facial attribute in the initial facial parameter corresponding to the target round includes: If the parameter type of the first facial attribute is a discrete type, it is determined that a first weight corresponding to a first sub-parameter corresponding to the first facial attribute in the initial facial parameters corresponding to the target round is 1.
7. The method according to claim 5, characterized in that The determining, based on the parameter type of the first facial attribute, a first weight corresponding to a first sub-parameter corresponding to the first facial attribute in the initial facial parameter corresponding to the target round includes: If the parameter type of the first facial attribute is a continuous type, determining normalized position data corresponding to the first sub-parameter; A first weight corresponding to the first sub-parameter is determined according to the normalized position data.
8. The method according to claim 7, characterized in that The determining the normalized position data corresponding to the first sub-parameter includes: Determining position reference data corresponding to the first sub-parameter according to a data relationship between the first sub-parameter and the second sub-parameter; The normalized position data corresponding to the first sub-parameter is determined according to the position reference data, the first sub-parameter and the second sub-parameter.
9. The method according to claim 8, characterized in that The determining, according to the position reference data, the first sub-parameter and the second sub-parameter, the normalized position data corresponding to the first sub-parameter includes: determining a first difference between the position reference data and the second sub-parameter; determining a second difference between the first sub-parameter and the second sub-parameter; An absolute value corresponding to the ratio of the first difference to the second difference is determined as normalized position data corresponding to the first sub-parameter.
10. The method according to claim 8, characterized in that The determining, according to the data relationship between the first sub-parameter and the second sub-parameter, the position reference data corresponding to the first sub-parameter includes: If the first sub-parameter is greater than the second sub-parameter, the maximum value of the first facial attribute in each target facial parameter corresponding to each round is determined as the position reference data corresponding to the first sub-parameter.
11. The method according to claim 8, characterized in that The determining, according to the data relationship between the first sub-parameter and the second sub-parameter, the position reference data corresponding to the first sub-parameter includes: If the first sub-parameter is smaller than the second sub-parameter, the minimum value of the first facial attribute among the target facial parameters corresponding to each round is determined as the position reference data corresponding to the first sub-parameter.
12. The method according to claim 1, characterized in that Before generating the initial facial parameters corresponding to the target round according to the target description information corresponding to the target round, the method further includes: Acquire first description information inputted for a facial image to be generated in a target round, and a set of historical information corresponding to the target round; According to the first description information and the historical information set, the target description information and target editing strength corresponding to the target round are generated; wherein the target editing strength corresponding to the target round represents the degree of influence of the target description information corresponding to the target round on the continuous type of facial attributes when generating the target facial parameters corresponding to the target round.
13. The method according to claim 12, characterized in that The historical information set includes historical information corresponding to at least one historical round before the target round, and the historical information includes description information input for the facial image to be generated in the historical round, target description information corresponding to the historical round, and target editing strength corresponding to the historical round.
14. The method according to claim 7, characterized in that The determining, according to the normalized position data, a first weight corresponding to the first sub-parameter includes: The product of the normalized position data and the target editing strength corresponding to the target round is determined as the first weight corresponding to the first sub-parameter.
15. The method according to claim 12, characterized in that The step of generating target description information and target editing strength corresponding to the target round according to the first description information and the historical information set includes: The first description information and the historical information set are input into a pre-trained language model, so as to output target description information and target editing strength corresponding to the target round through the pre-trained language model.
16. The method according to claim 15, characterized in that The pre-trained language model is trained in the following way: Acquire at least one sample corpus; wherein the sample corpus includes sample current description information, a sample history information set, and sample target information, and the sample target information includes sample target description information and sample target editing strength; Inputting the sample corpus into a language model to be trained, so that the language model outputs corresponding prediction target information according to the current description information of the sample and the historical description information of the sample; wherein the prediction target information includes prediction target description information and prediction target editing strength; The model parameters of the language model are adjusted so that the difference between the predicted target information and the sample target information is smaller than a first threshold, thereby obtaining a pre-trained language model.
17. The method according to claim 2, characterized in that Generating the initial facial image corresponding to the target round according to the target description information includes: Inputting the target description information into a pre-trained image diffusion model to obtain an initial noisy image and a noise value corresponding to each time node through the image diffusion model; The noise value corresponding to each time node is removed in sequence from the initial noisy image to obtain an initial facial image corresponding to the target round.
18. The method according to claim 17, characterized in that The pre-trained image diffusion model is trained in the following way: Constructing at least one sample data pair; wherein the sample data pair includes a sample facial image and annotation information corresponding to the sample facial image; For each time node, a preset noise value is sampled by the sample facial image to obtain a sample noisy image; Inputting the annotation information corresponding to the sample noisy image into the image diffusion model to be trained, so as to output a predicted noise value through the image diffusion model; The model parameters of the image diffusion model are adjusted so that the difference between the predicted noise value and the preset noise value belonging to the same time node is less than a second threshold, thereby obtaining a pre-trained image diffusion model.
19. The method according to claim 2, characterized in that The determining of initial facial parameters corresponding to the initial facial image comprises: The initial facial image is input into a pre-trained parameter conversion model to output corresponding initial facial parameters through the pre-trained parameter conversion model.
20. The method according to claim 19, characterized in that The pre-trained parameter conversion model is trained in the following way: Acquire at least one sample facial parameter, and render the sample facial parameter into a sample facial image; Inputting the sample facial parameters into a parameter conversion model to be trained, so as to output a predicted facial image through the parameter conversion model; The parameters of the parameter conversion model are adjusted so that the difference between the predicted facial image and the sample facial image is within a preset range, thereby obtaining a pre-trained parameter conversion model.
21. The method according to claim 20, characterized in that The obtaining of at least one sample facial parameter comprises: obtaining at least one initial sample facial parameter; Random perturbations are applied to the facial attributes of the initial sample facial parameters to obtain sample facial parameters.
22. The method according to claim 1, characterized in that The step of determining facial attribute adjustment information corresponding to the target round according to the target description information includes: The target description information is input into a pre-trained facial attribute classification model, so as to output facial attribute adjustment information corresponding to the target round through the pre-trained facial attribute classification model.
23. The method according to claim 22, characterized in that The pre-trained facial attribute classification model is trained in the following way: Acquire at least one sample description information; wherein the sample description information is annotated with sample facial attribute adjustment information; Inputting the sample description information into a facial attribute classification model to be trained, so as to output predicted facial attribute adjustment information through the facial attribute classification model; The model parameters of the facial attribute classification model are adjusted so that the difference between the predicted facial attribute adjustment information and the sample facial attribute adjustment information is smaller than a third threshold, thereby obtaining a pre-trained facial attribute classification model.
24. A data processing device, characterized in that: The device comprises: A first generating unit, configured to generate initial facial parameters corresponding to the target round according to target description information corresponding to the target round; wherein the target description information is used to describe the adjustment to be performed on the facial image in the target round; A determination unit, configured to determine facial attribute adjustment information corresponding to the target round according to the target description information; wherein the facial attribute adjustment information is used to indicate an adjustment state of each facial attribute in the target round; The second generating unit is used to generate a target facial image corresponding to the target round according to the facial attribute adjustment information, the initial facial parameters corresponding to the target round and the target facial parameters corresponding to the previous round of the target round.
25. An electronic device, characterized in that: include: processor; as well as The memory is used to store a data processing program. After the electronic device is powered on and the program is run by the processor, the method as described in any one of claims 1 to 23 is executed.
26. A computer-readable storage medium, characterized in that: A data processing program is stored, and the program is run by a processor to execute the method according to any one of claims 1 to 23.
Citation Information
Patent Citations
Face recognition method
CN104200194A
Speaking face video generation method and device based on multi-modal information control
CN117456587A
Image processing and model distillation training method and device, equipment and storage medium
CN117636136A
Virtual face shaping engine based on natural language interaction
CN118710809A
Virtual character efficient generation system based on multi-modal large model
CN118987612A