Digital human generation method and device, equipment and medium
By generating grid digital people suitable for client network parameters, the problem of large network overhead in digital people live broadcast is solved, and data volume reduction and rendering efficiency improvement are achieved.
Patent Information
- Application Number
- CN202411998801.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-07-29
AI Technical Summary
In digital live broadcast, the server transmits a large amount of data to the client, resulting in high network overhead and heavy rendering burden on the client.
The server generates grid parameters suitable for the client's network parameters, generates grid digital people based on this parameter and point cloud parameters, and sends them to the client to render. The amount of grid digital people data is smaller than that of point cloud digital people. The client can adjust grid parameters according to network parameters to adapt to changes in bandwidth and clarity.
It reduces the amount of data transmission between the server and the client, reduces network overhead, balances the performance overhead of the client and the server, and improves rendering efficiency and effect.
Smart Images

Figure CN120390118A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technologies, and in particular, to a digital human generation method, apparatus, device, and medium. Background Art
[0002] In recent years, with the continuous progress of computer technologies, digital human live streaming has become a new live streaming method. Digital human live streaming can reduce the labor costs of live streaming and the like, and is widely applied to various industries such as education and health.
[0003] In related technologies, when performing digital human live streaming, the server usually needs to send the rendered live video stream to the client, and the client plays the rendered live video stream. The live video stream rendered by the server is usually generated based on a point cloud digital human, and the data volume of this live video stream is usually large, which results in a large data volume transmitted between the server and the client and a large network overhead. Summary of the Invention
[0004] This application provides a digital human generation method, apparatus, device, and medium, which are used to reduce network overhead and perform digital human live streaming quickly and simply.
[0005] In a first aspect, this application provides a digital human generation method. The method is applied to a server, and the method includes:
[0006] Receiving first network parameters of a client, and generating first grid parameters based on the first network parameters;
[0007] Generating a first grid digital human based on the first grid parameters and first point cloud parameters;
[0008] Sending the first grid digital human to the client so that the client renders the first grid digital human to generate a target digital human; wherein, the first grid digital human includes a plurality of grid units in a three-dimensional space, the first grid parameters are used to represent the number of grid units of the first grid digital human; the first point cloud parameters are used to represent the three-dimensional space coordinates of a plurality of points on a point cloud digital human corresponding to the first grid digital human, and the data volume of the point cloud digital human is greater than the data volume of the first grid digital human.
[0009] In the above manner, since, compared with the rendered live video stream generated by the server based on the point cloud digital human, the amount of data of the first mesh digital human sent by the server to the client in the embodiments of the present application is smaller than the amount of data of the point cloud digital human and smaller than the amount of data of the rendered live video stream generated based on the point cloud digital human, therefore, the embodiments of the present application can reduce the amount of data transmitted between the server and the client, can reduce network overhead, and can perform digital human live broadcast quickly and simply. In addition, in the embodiments of the present application, the client can render the first mesh digital human to generate a target digital human for live broadcast, and the client can help the server share the task of rendering and generating the target digital human, thereby balancing the performance overhead between the client and the server to a certain extent.
[0010] In a possible implementation manner, the first network parameter includes the bandwidth parameter of the client and / or the playback clarity of the digital human by the client, and generating the first mesh parameter based on the first network parameter includes:
[0011] Inputting the first network parameter into a mesh parameter generation model to generate the first mesh parameter.
[0012] In the above manner, based on the mesh parameter generation model, the mesh parameter suitable for the network parameter of the client can be generated quickly and accurately.
[0013] In a possible implementation manner, generating the first mesh digital human based on the first mesh parameter and the first point cloud parameter includes:
[0014] Obtaining first audio, and inputting the first audio into an audio-visual model to generate the first point cloud digital human;
[0015] Extracting the first point cloud parameter of the first point cloud digital human, and generating the first mesh digital human based on the first mesh parameter and the first point cloud parameter.
[0016] In the above manner, based on the audio-visual model, the first audio can be converted into the corresponding first point cloud digital human. By extracting the first point cloud parameter of the first point cloud digital human, based on the first mesh parameter and the first point cloud parameter, the point cloud digital human with a large amount of data can be quickly converted into a mesh digital human with a network parameter suitable for the client and a small amount of data.
[0017] In a possible implementation manner, generating the first mesh digital human based on the first mesh parameter and the first point cloud parameter includes:
[0018] Encoding the first mesh parameter and the first point cloud parameter respectively to generate a first feature vector of the first point cloud parameter and a second feature vector of the first mesh parameter;
[0019] Input the first feature vector and the second feature vector into the grid digital human generation model to generate a first grid digital human.
[0020] In the above manner, by encoding the first point cloud parameters, a first feature vector of the first point cloud parameters can be obtained, and by encoding the first grid parameters, a second feature vector of the first grid parameters can be obtained. Inputting the first feature vector and the second feature vector into the grid digital human generation model, a grid digital human suitable for the client with a small amount of data can be quickly generated based on the grid digital human generation model.
[0021] In a possible implementation manner, the method further includes:
[0022] Send the first grid digital human and the texture map corresponding to the first grid digital human to the client, so that the client renders the first grid digital human based on the texture map to generate the target digital human; wherein, the resolution of the texture map corresponds to the number of grid cells of the first grid digital human.
[0023] In the above manner, a texture map suitable for the grid cells of the first grid digital human can be determined, and the first grid digital human and the texture map Figure 1 are sent to the client, which can improve the rendering efficiency and rendering effect of the client for the first grid digital human.
[0024] In a possible implementation manner, the method further includes:
[0025] When the first network parameters of the client change to second network parameters, receive the second network parameters of the client and generate second grid parameters based on the second network parameters;
[0026] Generate a second grid digital human based on the second grid parameters and the first point cloud parameters;
[0027] Send the second grid digital human to the client, so that the client renders the second grid digital human to generate the target digital human; wherein the number of grid cells of the second grid digital human is different from the number of grid cells of the first grid digital human.
[0028] In the above manner, when there are fluctuations in the network bandwidth of the client or the user changes the playback clarity selection of the digital human on the client, the changed second network parameters can be sent to the server. The server can determine the second grid parameters and the second grid digital human suitable for the second network parameters, and send the second grid digital human suitable for the second network parameters of the client to the client. The target digital human for live broadcast generated by the client based on the second grid digital human can better conform to the second network parameters of the client, can maximize the live broadcast effect, and improve the user experience.
[0029] In a possible implementation manner, the change of the first network parameters of the client to the second network parameters includes: the bandwidth in the second network parameters is greater than the bandwidth in the first network parameters, or the playback clarity in the second network parameters is greater than the playback clarity in the first network parameters; or the bandwidth in the second network parameters is less than the bandwidth in the first network parameters, or the playback clarity in the second network parameters is less than the playback clarity in the first network parameters;
[0030] In the case where the bandwidth in the second network parameters is greater than the bandwidth in the first network parameters, or the playback clarity in the second network parameters is greater than the playback clarity in the first network parameters, the number of grid units of the second grid digital human is greater than the number of grid units of the first grid digital human, and the data volume of the second grid digital human is greater than the data volume of the first grid digital human;
[0031] In the case where the bandwidth in the second network parameters is less than the bandwidth in the first network parameters, or the playback clarity in the second network parameters is less than the playback clarity in the first network parameters, the number of grid units of the second grid digital human is less than the number of grid units of the first grid digital human, and the data volume of the second grid digital human is less than the data volume of the first grid digital human.
[0032] In the above manner, according to the change of the network parameters of the client, the grid digital human suitable for the network parameters of the client can be obtained.
[0033] In a second aspect, the present application provides a method for training a grid parameter generation model, the method including:
[0034] Obtain sample network parameters in a sample set, where the sample network parameters correspond to sample grid parameters;
[0035] Input the sample network parameters into a grid parameter generation model to be trained to generate predicted grid parameters;
[0036] Train the grid parameter generation model to be trained according to the sample grid parameters and the predicted grid parameters.
[0037] In a third aspect, the present application provides a method for training a grid digital human generation model, the method comprising:
[0038] Obtain a first sample feature vector of sample point cloud parameters in a sample set, and a second sample feature vector of sample grid parameters corresponding to the first sample feature vector; wherein, the sample grid parameters are grid parameters of a sample grid digital human;
[0039] Input the first sample feature vector and the second sample feature vector into the grid digital human generation model to be trained, generate a predicted grid digital human, extract the predicted grid parameters of the predicted grid digital human, and encode the predicted grid parameters to generate a predicted feature vector of the predicted grid parameters;
[0040] Train the grid digital human generation model to be trained according to the predicted feature vector and the second sample feature vector.
[0041] In a fourth aspect, the present application provides a digital human display method, which is applied to a client, and the method comprises:
[0042] Receive a first grid digital human sent by a server;
[0043] Render the first grid digital human to generate a target digital human;
[0044] Display the target digital human;
[0045] Wherein, the first grid digital human is generated based on first grid parameters and first point cloud parameters; the first grid parameters are used to characterize the number of grid cells of the first grid digital human; the first point cloud parameters are used to characterize the three-dimensional spatial coordinates of multiple points on a first point cloud digital human corresponding to the first grid digital human, and the data volume of the point cloud digital human is greater than the data volume of the first grid digital human.
[0046] In a possible implementation manner, the first grid parameters are generated based on first network parameters of the client, and the first network parameters include the bandwidth parameter of the client and / or the playback clarity of the digital human by the client;
[0047] In the case that the first network parameters of the client change to second network parameters, receive a second grid digital human sent by the server;
[0048] Render the second grid digital human to generate a target digital human;
[0049] Display the target digital human; wherein, the number of grid cells of the second grid digital human is different from the number of grid cells of the first grid digital human.
[0050] In a possible implementation manner, the change of the first network parameter to the second network parameter includes: the bandwidth in the second network parameter is greater than the bandwidth in the first network parameter, or the playback clarity in the second network parameter is greater than the playback clarity in the first network parameter; or the bandwidth in the second network parameter is less than the bandwidth in the first network parameter, or the playback clarity in the second network parameter is less than the playback clarity in the first network parameter;
[0051] When the bandwidth in the second network parameter is greater than the bandwidth in the first network parameter, or the playback clarity in the second network parameter is greater than the playback clarity in the first network parameter, the number of grid cells of the second grid digital human is greater than the number of grid cells of the first grid digital human, and the data volume of the second grid digital human is greater than the data volume of the first grid digital human;
[0052] When the bandwidth in the second network parameter is less than the bandwidth in the first network parameter, or the playback clarity in the second network parameter is less than the playback clarity in the first network parameter, the number of grid cells of the second grid digital human is less than the number of grid cells of the first grid digital human, and the data volume of the second grid digital human is less than the data volume of the first grid digital human.
[0053] In a possible implementation manner, the rendering of the first grid digital human to generate a target digital human includes:
[0054] Receiving the texture map corresponding to the first grid digital human;
[0055] Rendering the first grid digital human based on the texture map to generate the target digital human; wherein, the resolution of the texture map corresponds to the number of grid cells of the first grid digital human.
[0056] In a possible implementation manner, the target digital human includes a plurality of sub-target digital humans, and each sub-target digital human corresponds to each audio frame of the first audio;
[0057] The display of the target digital human includes:
[0058] When playing each audio frame of the first audio, synchronously display the sub-target digital human corresponding to each audio frame.
[0059] In a fifth aspect, an embodiment of the present application provides a digital human generation device, including units or modules for executing the method described in any one of the first aspects. The digital human generation device includes:
[0060] A first generation module, configured to receive first network parameters of a client, and generate first grid parameters based on the first network parameters;
[0061] A second generation module, configured to generate a first grid digital human based on the first grid parameters and first point cloud parameters;
[0062] A sending module, configured to send the first grid digital human to the client, so that the client renders the first grid digital human to generate a target digital human; wherein, the first grid digital human includes a plurality of grid units in a three-dimensional space, the first grid parameters are used to characterize the number of grid units of the first grid digital human; the first point cloud parameters are used to characterize the three-dimensional space coordinates of a plurality of points on a point cloud digital human corresponding to the first grid digital human, and the data volume of the point cloud digital human is greater than the data volume of the first grid digital human.
[0063] In a possible implementation manner, the first generation module is specifically configured to:
[0064] The first network parameters include the bandwidth parameter of the client and / or the playback clarity of the digital human by the client. Input the first network parameters into a grid parameter generation model to generate the first grid parameters.
[0065] In a possible implementation manner, the second generation module is specifically configured to:
[0066] Obtain first audio, input the first audio into an audio-visual model to generate the first point cloud digital human;
[0067] Extract the first point cloud parameters of the first point cloud digital human, and generate a first grid digital human based on the first grid parameters and the first point cloud parameters.
[0068] In a possible implementation manner, the second generation module is specifically configured to:
[0069] Encode the first grid parameters and the first point cloud parameters respectively to generate a first feature vector of the first point cloud parameters and a second feature vector of the first grid parameters;
[0070] Input the first feature vector and the second feature vector into a grid digital human generation model to generate a first grid digital human.
[0071] In a possible implementation manner, the sending module is further configured to:
[0072] Send the first mesh digital human and the texture map corresponding to the first mesh digital human to the client, so that the client renders the first mesh digital human based on the texture map to generate the target digital human; wherein, the resolution of the texture map corresponds to the number of mesh units of the first mesh digital human.
[0073] In a possible implementation manner, the first generation module is further configured to:
[0074] When the first network parameter of the client changes to a second network parameter, receive the second network parameter of the client, and generate a second mesh parameter based on the second network parameter;
[0075] The second generation module is further configured to:
[0076] Generate a second mesh digital human based on the second mesh parameter and the first point cloud parameter;
[0077] The sending module is further configured to:
[0078] Send the second mesh digital human to the client, so that the client renders the second mesh digital human to generate a target digital human.
[0079] In a possible implementation manner, the change of the first network parameter of the client to the second network parameter includes: the bandwidth in the second network parameter is greater than the bandwidth in the first network parameter, or the playback clarity in the second network parameter is greater than the playback clarity in the first network parameter; or the bandwidth in the second network parameter is less than the bandwidth in the first network parameter, or the playback clarity in the second network parameter is less than the playback clarity in the first network parameter;
[0080] When the bandwidth in the second network parameter is greater than the bandwidth in the first network parameter, or the playback clarity in the second network parameter is greater than the playback clarity in the first network parameter, the number of mesh units of the second mesh digital human is greater than the number of mesh units of the first mesh digital human, and the data volume of the second mesh digital human is greater than the data volume of the first mesh digital human;
[0081] When the bandwidth in the second network parameter is less than the bandwidth in the first network parameter, or the playback clarity in the second network parameter is less than the playback clarity in the first network parameter, the number of mesh units of the second mesh digital human is less than the number of mesh units of the first mesh digital human, and the data volume of the second mesh digital human is less than the data volume of the first mesh digital human.
[0082] In a sixth aspect, an embodiment of the present application provides a model training device, which includes units or modules for executing the method in the second aspect, or includes units or modules for executing the method in the third aspect.
[0083] When the model training device executes the method in the second aspect, the model training device includes:
[0084] A first acquisition module, configured to acquire sample network parameters in a sample set, where the sample network parameters correspond to sample grid parameters;
[0085] A first prediction module, configured to input the sample network parameters into a grid parameter generation model to be trained to generate predicted grid parameters;
[0086] A first training module, configured to train the grid parameter generation model to be trained according to the sample grid parameters and the predicted grid parameters.
[0087] When the model training device executes the method in the third aspect, the model training device includes:
[0088] A second acquisition module, configured to acquire a first sample feature vector of sample point cloud parameters in a sample set, and a second sample feature vector of sample grid parameters corresponding to the first sample feature vector; wherein, the sample grid parameters are grid parameters of a sample grid digital human;
[0089] A second prediction module, configured to input the first sample feature vector and the second sample feature vector into a grid digital human generation model to be trained to generate a predicted grid digital human, extract predicted grid parameters of the predicted grid digital human, and encode the predicted grid parameters to generate a predicted feature vector of the predicted grid parameters;
[0090] A second training module, configured to train the grid digital human generation model to be trained according to the predicted feature vector and the second sample feature vector.
[0091] In a seventh aspect, the present application provides a digital human display device, which includes units or modules for executing any of the methods in the fourth aspect. The digital human display device includes:
[0092] A receiving module, configured to receive a first grid digital human sent by a server;
[0093] A rendering module, configured to render the first grid digital human to generate a target digital human;
[0094] A display module, configured to display the target digital human;
[0095] Among them, the first grid digital human is generated based on the first grid parameters and the first point cloud parameters; the first grid parameters are used to characterize the number of grid cells of the first grid digital human; the first point cloud parameters are used to characterize the three-dimensional spatial coordinates of multiple points on the first point cloud digital human corresponding to the first grid digital human, and the data volume of the point cloud digital human is greater than the data volume of the first grid digital human.
[0096] In a possible implementation manner, the first grid parameters are generated based on the first network parameters of the client, and the first network parameters include the bandwidth parameter of the client and / or the playback clarity of the digital human by the client;
[0097] The receiving module is further configured to receive the second grid digital human sent by the server when the first network parameters of the client change to the second network parameters;
[0098] The rendering module is further configured to render the second grid digital human to generate a target digital human;
[0099] The display module is further configured to display the target digital human; among them, the number of grid cells of the second grid digital human is different from the number of grid cells of the first grid digital human.
[0100] In a possible implementation manner, the bandwidth in the second network parameters is greater than the bandwidth in the first network parameters, or the playback clarity in the second network parameters is greater than the playback clarity in the first network parameters; or the bandwidth in the second network parameters is less than the bandwidth in the first network parameters, or the playback clarity in the second network parameters is less than the playback clarity in the first network parameters;
[0101] When the bandwidth in the second network parameters is greater than the bandwidth in the first network parameters, or the playback clarity in the second network parameters is greater than the playback clarity in the first network parameters, the number of grid cells of the second grid digital human is greater than the number of grid cells of the first grid digital human, and the data volume of the second grid digital human is greater than the data volume of the first grid digital human;
[0102] When the bandwidth in the second network parameters is less than the bandwidth in the first network parameters, or the playback clarity in the second network parameters is less than the playback clarity in the first network parameters, the number of grid cells of the second grid digital human is less than the number of grid cells of the first grid digital human, and the data volume of the second grid digital human is less than the data volume of the first grid digital human.
[0103] In a possible implementation manner, the rendering module is specifically configured to:
[0104] Receive the texture map corresponding to the first mesh digital human;
[0105] Render the first mesh digital human based on the texture map to generate the target digital human; wherein, the resolution of the texture map corresponds to the number of mesh units of the first mesh digital human.
[0106] In a possible implementation manner, the target digital human includes a plurality of sub-target digital humans, and each sub-target digital human corresponds to each audio frame of the first audio;
[0107] The display module is specifically configured to:
[0108] When playing each audio frame of the first audio, synchronously display the sub-target digital human corresponding to each audio frame.
[0109] In an eighth aspect, the present application provides a computing device, which includes a processor and a memory. The processor is used to call computer program instructions stored in the memory to execute the method described in any item of the first aspect, or execute the method described in the second aspect, or execute the method described in the third aspect, or execute the method described in the fourth aspect.
[0110] Among them, the computing device can be a server or a terminal device.
[0111] In a ninth aspect, the present application provides a digital human live broadcast system, which includes: a server that executes the method described in any item of the first aspect, and a client described in the fourth aspect.
[0112] In a tenth aspect, the present application provides a computer-readable storage medium, in which computer programs or instructions are stored. When the computer programs or instructions are executed by a communication device, the method described in any item of the first aspect is implemented, or the method described in the second aspect is implemented, or the method described in the third aspect is executed, or the method described in the fourth aspect is executed.
[0113] In an eleventh aspect, the present application provides a computer program product, which, when called by a computer, causes the computer to execute the method described in any item of the first aspect, or execute the method described in the second aspect, or execute the method described in the third aspect, or execute the method described in the fourth aspect.
[0114] In a twelfth aspect, the present application provides a computing device, which includes a processor and a memory. The processor is configured to call computer program instructions stored in the memory to execute the method described in any one of the first aspect, or execute the method described in the second aspect, or execute the method described in the third aspect, or execute the method described in the fourth aspect. Description of the Drawings
[0115] To more clearly illustrate the embodiments of the present application or the implementation manners in the related art, the following will briefly introduce the drawings required for use in the description of the embodiments or the related art. Obviously, the drawings in the following description are some embodiments of the present application, and those of ordinary skill in the art can also obtain other drawings based on these drawings.
[0116] Figure 1 Schematic diagram of a digital human generation process provided by an embodiment of the present application;
[0117] Figure 2A Schematic diagram of a grid digital human provided by an embodiment of the present application;
[0118] Figure 2B Another schematic diagram of a grid digital human provided by an embodiment of the present application;
[0119] Figure 2C Another schematic diagram of a grid digital human provided by an embodiment of the present application;
[0120] Figure 2D Schematic diagram of a target digital human provided by an embodiment of the present application;
[0121] Figure 2E Another schematic diagram of a target digital human provided by an embodiment of the present application;
[0122] Figure 3 Another schematic diagram of a digital human generation process provided by an embodiment of the present application;
[0123] Figure 4 Another schematic diagram of a digital human generation process provided by an embodiment of the present application;
[0124] Figure 5 Schematic diagram of the structure of a digital human generation device provided by an embodiment of the present application;
[0125] Figure 6 Schematic diagram of the structure of a model training device provided by an embodiment of the present application;
[0126] Figure 7 Another schematic diagram of the structure of a model training device provided by an embodiment of the present application;
[0127] Figure 8Schematic structural diagram of a digital human display device provided by an embodiment of the present application;
[0128] Figure 9 Schematic structural diagram of an electronic device provided by an embodiment of the present application;
[0129] Figure 10 Schematic diagram of a digital human generation system provided by an embodiment of the present application. Detailed implementation manners
[0130] Embodiments of the present application provide a method, apparatus, device and medium for live streaming of a digital human. Among them, a digital human refers to a digital human image created by using digital technology and similar to the human image. A digital human is the product of the integration of information science and life science, and can use the methods of information science to perform virtual simulation on the morphology and functions of the human body at different levels.
[0131] To make the purpose and implementation manners of the present application clearer, the following will clearly and completely describe the exemplary implementation manners of the present application with reference to the accompanying drawings in the exemplary embodiments of the present application. Obviously, the described exemplary embodiments are only a part of the embodiments of the present application, rather than all of the embodiments.
[0132] It should be noted that the brief description of the terms in the present application is only for the convenience of understanding the subsequent described implementation manners, rather than intending to limit the implementation manners of the present application. Unless otherwise specified, these terms should be understood in their ordinary and common meanings.
[0133] The terms "first", "second", "third", etc. in the specification, claims and the above drawings of the present application are used to distinguish similar or like objects or entities, and do not necessarily mean to limit a specific order or sequence, unless otherwise noted. It should be understood that such terms can be interchanged under appropriate circumstances.
[0134] The terms "comprising" and "having" and any variations thereof are intended to cover but not exclude inclusion. For example, a product or device comprising a series of components does not necessarily have to be limited to all the clearly listed components, but may include other components not clearly listed or inherent to these products or devices.
[0135] The term "module" refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic or a combination of hardware or / and software code that can perform functions related to the element.
[0136] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the various embodiments of the present application.
[0137] For ease of explanation, the above description has been presented in the context of specific embodiments. However, the above exemplary discussion is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed. Numerous modifications and variations are possible in light of the above teachings. The selection and description of the above embodiments are for the purpose of better explaining the principles and the practical application, so that those skilled in the art can better use the embodiments and various different variations of the embodiments suitable for specific use considerations.
[0138] Embodiment 1:
[0139] Figure 1 A schematic diagram of a digital human generation process provided by an embodiment of the present application, the process including:
[0140] S101: The server receives the first network parameter sent by the client, and generates the first grid parameter based on the first network parameter.
[0141] In a possible implementation manner, considering that the bandwidth of the client may fluctuate, and the user may also have different requirements for the playback clarity of the digital human during the live broadcast, etc., in order to determine a digital human suitable for the bandwidth of the client and the playback clarity requirements, the server may receive the first network parameter of the client, and may generate the first grid parameter suitable for the first network parameter based on the first network parameter. Optionally, the first network parameter may include at least one of the current bandwidth parameter of the client and the playback clarity of the digital human for the client. Exemplarily, the first network parameter may only include the current bandwidth parameter of the client, may only include the playback clarity of the digital human for the client, or may include both the current bandwidth parameter of the client and the playback clarity of the digital human for the client.
[0142] In a possible implementation, in order to quickly determine the first grid parameter suitable for the first network parameter, the corresponding relationship between the network parameter and the grid parameter can be configured based on methods such as user questionnaire surveys and digital human performance tests. After that, after receiving the first network parameter of the client, the grid parameter corresponding to the first network parameter can be determined based on the pre-configured corresponding relationship between the network parameter and the grid parameter, and this grid parameter can be determined as the first grid parameter suitable for the first network parameter. Among them, the present application does not specifically limit the corresponding relationship between the first grid parameter and the first network parameter, and it can be flexibly set according to requirements. Exemplarily, please refer to Figure 2A , Figure 2A which is a schematic diagram of a grid digital human provided by an embodiment of the present application. The grid digital human can be constructed by multiple grid units with different shapes (such as triangles, parallelograms, irregular shapes, etc.). Among them, the grid parameter can be used to characterize the number of grid units of the grid digital human. The grid parameter can include the number of grid units that make up the grid digital human, or further include the three-dimensional coordinates of each grid unit. For example, please refer to Figure 2B , Figure 2B which is another schematic diagram of a grid digital human provided by an embodiment of the present application. The more the number of grid units included in the grid parameter, the more the number of sub-grid faces of the corresponding grid digital human, and the higher the clarity of the digital human rendered based on this grid parameter. This grid parameter is more suitable for playback modes with larger bandwidths and higher requirements for clarity. On the contrary, please refer to Figure 2C , Figure 2C which is still another schematic diagram of a grid digital human provided by an embodiment of the present application. The fewer the number of grid units included in the grid parameter, the fewer the number of sub-grid faces of the corresponding grid digital human, and the lower the clarity of the digital human rendered based on this grid parameter. This grid parameter is more suitable for playback modes with smaller bandwidths and lower requirements for clarity. That is to say, the grid parameter of the grid digital human corresponds to the network parameter of the client. When the network bandwidth (bandwidth parameter) of the client is low and the playback clarity is low, the number of grid units included in the grid parameter of the corresponding grid digital human is small; while when the network bandwidth of the client is high and the playback clarity is high, the number of grid units included in the grid parameter of the corresponding grid digital human is large.
[0143] In a possible implementation, when generating the first grid parameter suitable for the first network parameter based on the first network parameter, the first network parameter can also be input into a pre-trained grid parameter generation model, and based on the output result of the grid parameter generation model, the first grid parameter suitable for the first network parameter is determined (generated). Among them, the grid parameter generation model can also be referred to as the grid parameter generation model. The training process of the grid parameter generation model can be as follows:
[0144] The sample set for training the grid parameter generation model contains multiple sample network parameters, and each sample network parameter can respectively correspond to a sample grid parameter. For each sample network parameter, the sample network parameter can include a bandwidth parameter and / or a playback clarity.
[0145] When training the grid parameter generation model, any sample network parameter in the sample set can be obtained, and the sample network parameter is input into the grid parameter generation model to be trained, and the grid parameter corresponding to the sample network parameter output by the grid parameter generation model (for ease of description, called the predicted grid parameter) is obtained. Based on the configured error loss function, the error between the predicted grid parameter and the sample grid parameter can be measured. The larger the value of the error loss function, the greater the error between the predicted grid parameter and the sample grid parameter, and the less accurate the recognition result (prediction result) of the model. The gradient descent algorithm can be used to perform backpropagation on the gradient of the model parameters and adjust the model parameters, thereby training the model and improving the prediction performance of the model by minimizing the value of the loss function.
[0146] In specific implementation, the above operations can be performed on each sample network parameter in the sample set. When the preset convergence condition is met, it is determined that the training of the grid parameter generation model is completed. Among them, meeting the preset convergence condition can be that the number of sample network parameters correctly recognized by the model in the sample set is greater than the set number, or the number of iterations for training the model reaches the set maximum number of iterations, etc. It can be flexibly set in specific implementation and will not be specifically limited here.
[0147] In a possible implementation manner, the client can send a first network parameter to the server under a preset first condition. For example, when the user starts watching the digital human live broadcast using the client by clicking the play button, etc., the client can receive the play live instruction triggered by the user. At this time, the client can obtain the current bandwidth parameter of the client. For example, the current bandwidth parameter of the client can be obtained through navigator.connection.downlink (an API interface of the browser), which will not be elaborated here. In addition, the user can also select a play mode with different requirements for the playback clarity of the digital human, and the client can determine the playback clarity corresponding to the play mode selected by the user according to the corresponding relationship between the pre-saved play mode and the playback clarity. The client can send the first network parameter including the bandwidth parameter and / or the playback clarity to the server, which will not be elaborated here.
[0148] S102: The server generates a first grid digital human based on the first grid parameter and the first point cloud parameter.
[0149] In a possible implementation, when generating a digital human for live streaming, the server can obtain the audio during the live streaming of the digital human (referred to as the first audio for convenience of description) from the client, input the first audio into a pre-trained audio-visual model, and the audio-visual model can convert the first audio into a corresponding first point cloud digital human. The first point cloud digital human is an animated digital human. The first audio includes multiple frames of audio, and each frame of audio has a corresponding sub-point cloud digital human. The sub-point cloud digital humans that are continuous in time form the above-mentioned first point cloud digital human. Each sub-point cloud digital human has a corresponding expression, posture, movement, and mouth shape. The audio-visual model is an audio-visual model that can transform sound memes into digital humans with specific forms. The first point cloud digital human is a high-precision digital human and includes point cloud data. Point cloud data refers to a set of vectors in a three-dimensional coordinate system. The scanned data is recorded in the form of points, and each point contains three-dimensional coordinates. Some may also carry other information such as the color and reflection intensity of the attributes of the point. The main characteristics of point cloud data are high-precision, high-resolution, and high-dimensional geometric information, which can intuitively represent the shape, surface, and texture of objects in space. It can be understood that the first point cloud digital human is constructed through point cloud data. The first point cloud digital human includes a large number of point sets distributed in space, and each point includes three-dimensional coordinate information, or further includes additional attributes such as color, reflection intensity, and normal vector.
[0150] Among them, the audio-visual model can also be called a sound-visual model, and the audio-visual model can transform sound memes into digital humans with specific forms (point cloud digital humans). The following is a brief introduction to the training process of the audio-visual model:
[0151] Before training the audio-visual model, multiple videos (massive videos) used for training the audio-visual model can be pre-processed first. For example, the videos can be classified first, and the videos can be classified into videos suitable for training digital humans with different character images such as young women, young men, elderly women, and elderly men. For another example, the noise and background noise of the sound generated by the audio data in the videos can also be reduced. The resolution, aspect ratio, and other parameters of the video frames in different videos can also be configured to the same unified parameter values.
[0152] For each video, the audio in the video can be extracted, as well as the features of the expressions, actions, lip shapes, etc. of the people in the video frames corresponding to each frame of the audio. The extracted audio for each frame and the corresponding video frames are input into a sound and video encoder (neural network model), and the model is trained to obtain a trained sound and video model. Among them, the training process of the sound and video model can use the Neural Radiance Field (Nerf) method in the three-dimensional scene reconstruction method, which will not be elaborated here.
[0153] In a possible implementation, considering that the first point cloud digital human contains a large number of point sets distributed in space that construct the human figure of the digital human, the data volume of the first point cloud digital human is usually relatively large. For example, the size of the first point cloud digital human can reach the gigabyte (GB) level. In order to not only reflect the human figure of the digital human contained in the first point cloud digital human in terms of appearance but also reduce network overhead, after obtaining the first point cloud digital human, the point cloud parameters (the first point cloud parameters) of the first point cloud digital human can be extracted, where the first point cloud parameters can represent the three-dimensional coordinate information of multiple points on the first point cloud digital human. After that, based on the first point cloud parameters and the first mesh parameters, a first mesh digital human can be generated. Among them, the data volume of the first mesh digital human is smaller than that of the first point cloud digital human, and the first mesh digital human includes multiple mesh units in three-dimensional space. Please refer to again Figure 2A or Figure 2B or Figure 2C , the mesh digital human can contain mesh units in the shapes of a set number of triangles, parallelograms, irregular shapes, etc. The mesh digital human constructed based on these mesh units can not only reflect the human figure of the point cloud digital human in terms of appearance but also has a smaller data volume than the point cloud digital human. For example, usually, the data volume of the point cloud digital human can reach the gigabyte (GB) level, while compared with the point cloud digital human, the data volume of the mesh digital human can be reduced to the megabyte (MB) level.
[0154] In a possible implementation, the process of generating the first mesh digital human can be as follows:
[0155] Encode the first cloud point parameters and the first mesh parameters respectively to generate the first feature vector of the first point cloud parameters and the second feature vector of the first mesh parameters; after that, the first feature vector and the second feature vector can be input into a pre-trained mesh digital human generation model to generate the first mesh digital human.
[0156] The training process of the mesh digital human generation model is introduced below. Among them, the point cloud digital human can be converted into a mesh digital human based on the mesh digital human generation model, and the mesh digital human generation model can be an autoregressive neural network model such as Transformer.
[0157] In a possible implementation, based on the sample point cloud digital humans corresponding to different sample audios, the point cloud parameters (sample point cloud parameters) of each sample point cloud digital human can be extracted, and the sample point cloud parameters can be encoded to generate a first sample feature vector corresponding to the sample point cloud parameters. Among them, the sample point cloud parameters include the three-dimensional spatial coordinates of each point on the digital human, as well as the color and reflection intensity information corresponding to each point.
[0158] A grid can be constructed for the same sample point cloud digital human to obtain multiple different sample grid digital humans corresponding thereto. Among them, each sample grid digital human corresponds to different grid parameters, and the grid parameters include the number of grid cells of the digital human, or further include the three-dimensional spatial coordinates of each grid cell. After the construction of the sample grid digital human is completed, for the same sample point cloud digital human, the sample grid parameters of its corresponding multiple different sample grid digital humans can be extracted, and the sample grid parameters can be encoded to generate a second sample feature vector corresponding to the sample grid parameters.
[0159] When training the grid digital human generation model, for each sample point cloud digital human, there are corresponding different sample grid digital humans, and each sample grid digital human has different grid parameters. The grid parameters of the sample grid digital human correspond to the network parameters of the client, which will not be elaborated here. Please refer to again Figure 2B and 2C , since one sample point cloud digital human can correspond to multiple different sample grid digital humans, therefore, the first sample feature vector of each sample point cloud parameter corresponds to the second sample feature vectors of multiple different sample grid parameters. Among them, the sample grid parameters correspond to the network parameters of the client, and the sample grid parameters can include the number of grid cells of the digital human, or further include the three-dimensional coordinates of each grid cell.
[0160] Exemplarily, during the training process of the grid digital human generation model, the first sample feature vector and one of the second sample feature vectors corresponding to the first sample feature vector can be input into the grid digital human generation model to be trained, and the grid digital human output by the grid digital human generation model (for ease of description, called the predicted grid digital human) can be obtained. The grid parameters of the predicted grid digital human can be extracted, and the grid parameters can be encoded to generate the feature vector corresponding to the grid parameters (for ease of description, called the predicted feature vector). Based on the configured error loss function, the error between the predicted feature vector and the corresponding second sample feature vector can be measured. The larger the value of the error loss function, the greater the error between the predicted feature vector and the corresponding second sample feature vector, and the more inaccurate the recognition result (prediction result) of the model. The gradient descent algorithm can be used to perform backpropagation on the gradient of the model parameters and adjust the model parameters, thereby training the model. By minimizing the value of the loss function, the prediction performance of the model can be improved.
[0161] In addition, the first sample feature vector and other second sample feature vectors corresponding to the first sample feature vector can be separately input into the grid digital human generation model to be trained, and the above process can be repeated to train the model. In a possible implementation manner, a linear projection layer can be added to the grid digital human generation model to learn the mapping relationship between the first feature vector of the sample point cloud parameters, the second feature vector of different sample grid parameters, and the sample grid digital humans of different sample grid parameters based on the linear projection layer, so as to improve the prediction performance of the model.
[0162] In a possible implementation manner, when the preset model convergence condition is satisfied, it can be determined that the training of the grid digital human generation model is completed. Among them, satisfying the preset model convergence condition can be that the number of iterations for training the grid digital human generation model reaches the set maximum number of iterations, or the difference between the corresponding predicted feature vector and the second sample feature vector of the grid digital human generation model is less than the preset difference threshold, etc. The model convergence condition can be flexibly set according to requirements, and the present application does not make specific limitations on this.
[0163] In a possible implementation, considering that there may be defects in the token sequence generated by the grid digital human generation model during the actual generation of the predicted grid digital human, in order to ensure the accuracy of the grid digital human output by the grid digital human generation model to the greatest extent when there are defects in the token sequence generated by the grid digital human generation model, during the training of the grid digital human generation model, Gaussian noise can be added to the second sample feature vector used for training the grid digital human generation model. These Gaussian noises can be regarded as the defects in the token sequence generated by the grid digital human generation model during the generation of the predicted grid digital human. Since the defects in the token sequence generated by the grid digital human generation model during actual use are simulated by adding Gaussian noise to the second sample feature vector during the training process of the grid digital human generation model, the training process of the grid digital human generation model can be closest to the actual use process to the greatest extent, thereby improving the accuracy of actually generating grid digital humans based on the trained grid digital human generation model.
[0164] Specifically, the process of adding Gaussian noise to the second sample feature vector can be as follows: First, encode the sample grid parameters to obtain the original feature vector of the sample grid parameters. Then, Gaussian noise can be added to this original feature vector to obtain the second sample feature vector with added Gaussian noise. Among them, Gaussian noise can be a type of noise whose probability density function follows a Gaussian distribution (i.e., a normal distribution). Gaussian noise can include fluctuation noise, cosmic noise, thermal noise, shot noise, etc., and this application does not make specific limitations in this regard. Among them, the process of training the grid digital human generation model based on the second sample feature vector with added Gaussian noise is the same as the process of training the grid digital human generation model based on the second sample feature vector introduced in the above embodiment, and will not be elaborated here.
[0165] In a possible implementation, when the server obtains the network parameters of the client, it can first convert the network parameters of the client into corresponding grid parameters based on the pre-trained network parameter and grid parameter conversion model (grid parameter generation model). Then, the first feature vector of the point cloud parameters and the second feature vector of the grid parameters can be input into the trained grid digital human generation model together, and the grid digital human generation model can output a grid digital human adapted to the network parameters of the client. That is to say, based on the trained grid digital human generation model, when the network parameters of the client change, a grid digital human corresponding to the network parameters can be output. For example, when the bandwidth in the network parameters is large, a digital human with more grid cells can be output, and at this time, the clarity of the digital human seen by the user is higher; while when the bandwidth in the network parameters is small, a digital human with fewer grid cells can be output, and at this time, the clarity of the digital human seen by the user is lower.
[0166] S103: The server sends the first mesh digital human to the client.
[0167] S104: The client renders the first mesh digital human to generate a target digital human, where the mesh parameters of the target digital human correspond to the network parameters of the client.
[0168] In a possible implementation manner, after the server determines the first mesh digital human, it can send the first mesh digital human to the client. After receiving the first mesh digital human sent by the server, the client can render the first mesh digital human. For example, it can add colors, hues, lighting, and apply texture maps with corresponding precision to the first mesh digital human, so as to generate a target digital human, which can be used for live streaming.
[0169] Exemplarily, the texture map may include images representing the figure of the digital human, the materials, textures, etc. of the digital human required for rendering the digital human. Among them, the material may include information such as the color and brightness of the figure representing the digital human. The texture can be used to wrap the image representing the figure onto the surface of the mesh model, etc. Among them, the process of rendering and generating the target digital human based on the mesh digital human and the texture map can adopt related technologies, which will not be elaborated here. Please refer to Figure 2D and Figure 2E , Figure 2D which is a schematic diagram of a target digital human provided by an embodiment of the present application, Figure 2E which is another schematic diagram of a target digital human provided by an embodiment of the present application. The target digital human viewed by the user is a digital human with certain colors, expressions, postures, movements, etc. after rendering, which will not be elaborated here.
[0170] Optionally, to facilitate the client to render the first mesh digital human, the server can further send the first mesh digital human and the texture map with the corresponding precision of the first mesh digital human Figure 1 to the client, so that the client can quickly render the first mesh digital human based on the texture map and quickly generate the target digital human. Specifically, the resolution of the texture map corresponding to the first mesh digital human corresponds to the number of mesh units of the first mesh digital human. For example, the more the number of mesh units (the number of sub-mesh faces) of the first mesh digital human, the larger the resolution of the texture map corresponding to the first mesh digital human, and the higher the clarity of the digital human rendered based on the texture map. On the contrary, the fewer the number of mesh units (the number of sub-mesh faces) of the first mesh digital human, the smaller the resolution of the texture map corresponding to the first mesh digital human, and the lower the clarity of the digital human rendered based on the texture map.
[0171] In a possible implementation, the server can pre-save the correspondence between mesh digital humans with different numbers of grid cells and textures with different resolutions, which facilitates timely sending of textures suitable for the mesh digital humans to the client and improves efficiency. Additionally, to save storage space, the server can also periodically clean up some textures that have been used less than a set number of times within a set time period to save storage space. If a certain texture that is not currently saved in the server needs to be used later, the texture can be generated in a timely manner. The specific process of generating the texture is not specifically limited in this application.
[0172] In a possible implementation, the first audio includes multiple audio frames, and the rendered target digital human is an animated digital human. The target digital human is generated based on the multiple audio frames in the first audio. Each audio frame has a corresponding sub-target digital human, and multiple temporally consecutive sub-target digital humans form the above-mentioned target digital human. That is to say, the target digital human includes sub-target digital humans corresponding to each audio frame of the first audio, and each sub-target digital human has a corresponding expression, posture, movement, and mouth shape. During the live broadcast based on the target digital human, the client can synchronize each audio frame of the first audio with the sub-target digital human corresponding to each audio frame while playing the first audio. Since the target digital human is generated based on the first audio, during the live broadcast, the expressions, movements, and mouth shapes shown by the target digital human correspond to the voice, intonation, speech rate, and content of the first audio.
[0173] In a possible implementation, considering that during the live broadcast by the client, the bandwidth of the client may fluctuate, and the user's requirements for the playback clarity of the digital human during the live broadcast may also change. To maximize the determination of the mesh digital human that suits the current network parameters of the client and maximize the guarantee of the live broadcast effect, when the client recognizes a bandwidth fluctuation (increase or decrease), for example, when During the duration the fluctuation amplitude of the bandwidth parameter exceeds the set threshold, it can be considered that the network parameters of the client have changed from the first network parameters to the second network parameters. The latest bandwidth parameter can be used as the changed second network parameter and sent to the server. The client can also send the latest bandwidth parameter and the playback clarity corresponding to the current playback mode as the second network parameter to the server. Among them, the bandwidth in the second network parameter can be greater than the bandwidth in the first network parameter or less than the bandwidth in the first network parameter. The specific duration of the first duration and the specific value of the set threshold corresponding to the fluctuation amplitude of the bandwidth parameter can be flexibly set according to requirements, and this application does not specifically limit this.
[0174] Optionally, the client can also identify the playback clarity corresponding to the changed playback mode when the user changes the playback mode and the client receives a playback clarity change instruction. The playback clarity can be a clarity that is larger than before or a clarity that is smaller than before. It can be considered that the network parameters of the client have changed, from the first network parameters to the second network parameters. The latest playback clarity can be used as the changed second network parameters, or the latest playback clarity and the current bandwidth parameters can be used together as the second network parameters, and the second network parameters are sent to the server. Among them, the playback clarity in the second network parameters can be greater than the playback clarity in the first network parameters or less than the playback clarity in the first network parameters.
[0175] After receiving the second network parameters of the client, the server can generate second grid parameters suitable for the second network parameters based on the second network parameters. Among them, the process of generating the second grid parameters based on the second network parameters is the same as the process of generating the first grid parameters based on the first network parameters introduced in the above embodiments, and will not be elaborated here.
[0176] After generating the second grid parameters, the server can generate a second grid digital human based on the second grid parameters and the first point cloud parameters. Among them, the number of grid cells of the second grid digital human is different from the number of grid cells of the first grid digital human. The server can send the second grid digital human to the client, and the client can render the second grid digital human to generate a target digital human. Among them, the process of generating the second grid digital human is the same as the process of generating the first grid digital human, and the process of rendering the second grid digital human to generate a target digital human is the same as the process of rendering the first grid digital human to generate a target digital human, and will not be elaborated here.
[0177] Optionally, when the bandwidth in the second network parameters is less than the bandwidth in the first network parameters, or the playback clarity in the second network parameters is less than the playback clarity in the first network parameters, the number of grid cells included in the second grid parameters is less than the number of grid cells included in the first grid parameters. Correspondingly, the number of sub-grid faces (grid cells) of the second grid digital human is less than the number of sub-grid faces (grid cells) of the first grid digital human, and the data volume of the second grid digital human is less than the data volume of the first grid digital human. The clarity of the target digital human rendered based on the second grid digital human is less than the clarity of the target digital human rendered based on the first grid digital human.
[0178] Conversely, when the bandwidth in the second network parameter is greater than the bandwidth in the first network parameter, or the playback clarity in the second network parameter is greater than the playback clarity in the first network parameter, the number of grid cells included in the second grid parameter is greater than the number of grid cells included in the first grid parameter. Correspondingly, the number of sub-grid faces (grid cells) of the second grid digital human is greater than the number of sub-grid faces (grid cells) of the first grid digital human, the data volume of the second grid digital human is greater than the data volume of the first grid digital human, and the clarity of the target digital human rendered based on the second grid digital human is greater than the clarity of the target digital human rendered based on the first grid digital human.
[0179] In a possible implementation manner, considering that the second grid digital human and the first grid digital human are grid digital humans with different numbers of sub-grid faces (grid cells), when sending the second grid digital human to the client, the server can also determine the texture map corresponding to the second grid digital human from the correspondence between the grid digital humans with different numbers of grid cells and the texture maps with different resolutions pre-stored, and can send the second grid digital human and the corresponding Figure 1 texture map to the client. After receiving the second grid digital human and the corresponding texture map, the client can render the second grid digital human based on the texture map to obtain a target digital human, and conduct a live broadcast based on the rendered target digital human and the first audio. During the live broadcast of the target digital human, the expressions, actions, and lip shapes shown by it correspond to the voice, intonation, speech rate, and content of the first audio.
[0180] In a possible implementation, the grid digital human is generated based on multiple audio frames in the first audio. Each audio frame corresponds to a sub-grid digital human, and multiple temporally consecutive sub-grid digital humans form the above-mentioned grid digital human. That is to say, the grid digital human includes sub-grid digital humans corresponding to each audio frame of the first audio, and each sub-grid digital human has a corresponding expression, posture, action, and mouth shape. Considering that during the live explanation process, the digital human may only involve changes in the positions of parts such as the mouth, face, and eyes, and the positions of parts such as the forehead and hands may not change. To minimize network overhead, when the server sends the grid digital human to the client, for each sub-grid digital human that is temporally consecutive in the grid digital human, the server can compare the sub-grid digital human (for ease of description, called the first sub-grid digital human) with the previous sub-grid digital human (for ease of description, called the second sub-grid digital human) located before this sub-grid digital human. For example, for each sub-grid (grid unit) in this sub-grid digital human, it is determined whether the position of this sub-grid in the first sub-grid digital human relative to its position in the second sub-grid digital human has changed. If it has changed, the position change information of which position (such as the (x1, y1, z1) coordinate position) in the second sub-grid digital human this sub-grid has changed to which position (such as the (x2, y2, z2) coordinate position) in the first sub-grid digital human can be further determined. The sub-grid digital human may include the position change information of the sub-grid whose position has changed in the first sub-grid digital human relative to the second sub-grid digital human, or further include the change information of additional attributes such as the color, reflection intensity, and normal of these sub-grids. When the client subsequently receives the grid digital human, it can adjust the positions and additional attributes of the sub-grids in the sub-grid digital human according to the position change information and the change information of the additional attributes of the sub-grids in each sub-grid digital human, and obtain the rendered sub-target digital human corresponding to each sub-grid digital human.
[0181] In a possible implementation, considering that when the digital human makes different sounds during the live explanation process, although the positions of the sub-grids in positions such as the face may change slightly, it may be difficult for the naked eye to detect such changes. If the position change information of these sub-grids is sent to the client, it may increase the bandwidth overhead. To save the bandwidth overhead to the greatest extent, for any sub-grid in the sub-grid digital human, after identifying that the position of the sub-grid has changed between two adjacent sub-grid digital humans, it can be further determined whether the position change of the sub-grid exceeds the set position change threshold. For example, assuming that the position coordinates of a certain sub-grid in the second sub-grid digital human are (x1, y1, z1) and the position coordinates in the first sub-grid digital human are (x2, y2, z2), the distance between the two coordinates (x1, y1, z1) and (x2, y2, z2) can be obtained, and it can be determined whether this distance exceeds the set position change threshold. Among them, the position change threshold can be a distance threshold. If this distance exceeds the set position change threshold, the position change information of the sub-grid can be included in the corresponding sub-grid digital human and sent to the client. Conversely, if this distance does not exceed the set position change threshold, the position change information of the sub-grid may not be included in the sub-grid digital human.
[0182] In the embodiment of the present application, the server can generate the first grid parameter based on the first network parameter of the client, and can generate the first grid digital human based on the first grid parameter and the first point cloud parameter characterizing the three-dimensional spatial coordinates of multiple points on the point cloud digital human, and send the first grid digital human to the client. The client renders the first grid digital human to obtain the target digital human, where the data volume of the first grid digital human is smaller than that of the point cloud digital human. Since compared with the rendered live video stream generated by the server based on the point cloud digital human in the related art and sending the live video stream to the client, in the embodiment of the present application, the data volume of the first grid digital human sent by the server to the client is smaller than the data volume of the point cloud digital human and smaller than the data volume of the rendered live video stream generated based on the point cloud digital human. Therefore, the embodiment of the present application can reduce the size of the data transmitted between the server and the client, can reduce the network overhead, and can quickly and simply perform digital human live broadcast. In addition, in the embodiment of the present application, the client can render the first grid digital human to generate the target digital human, and the client can share the task of rendering and generating the target digital human for the server, so as to balance the performance overhead between the client and the server to a certain extent.
[0183] In a possible implementation, since the client can render the first grid digital human to generate the target digital human for live broadcast, if the network connection between the client and the server is disconnected at this time, the client can play based on the currently rendered target digital human and the preset prompt mode. For example, the target digital human that has been rendered can still be displayed on the display interface of the client, and at the same time, a friendly prompt message such as "Network disconnected, please wait" can also be displayed, and play is carried out based on the rendered target digital human and the preset prompt mode. Compared with the related art, where the client receives the live video stream rendered by the server and cannot maintain the playback scene or prompt the user if the network connection between the client and the server is disconnected, the present application can maintain the playback scene and prompt the user, thereby improving the user experience.
[0184] In a possible implementation, when the network connection between the client and the server is restored, the client can output a prompt message "The network has been restored. Do you want to watch the live broadcast again?", and if it is recognized that the "Continue to watch" button is triggered, it can be considered that a live broadcast viewing instruction triggered by the user is received. The client can obtain the current bandwidth parameter and the playback clarity corresponding to the playback mode selected by the user, and send the bandwidth parameter and the playback clarity as network parameters to the server. The server can perform the steps of determining the grid parameters corresponding to the network parameters and generating the grid digital human as described above, and send the grid digital human to the client. After receiving the grid digital human, the client can render the grid digital human to obtain the target digital human and restore the live broadcast scene, which will not be elaborated here.
[0185] In a possible implementation, when the network is restored and the client recognizes that the "Continue to watch" button is triggered, it can then output a selection message "Play from the current time point or from the time point before the network disconnection". When it is recognized that playing from the current time point is selected, the client can send the network parameters and the information of playing from the current time point to the server together. The server can determine the audio corresponding to the current time point and generate the grid digital human corresponding to the audio, which will not be elaborated here.
[0186] When it is recognized that playing from the time point before the network disconnection is selected, the client can send the network parameters and the information of starting the live broadcast from the time point when the network was disconnected to the server together. The server can obtain the audio corresponding to the time period when the client's network was disconnected from the saved live broadcast data and generate the grid digital human corresponding to the audio, which will not be elaborated here.
[0187] For ease of understanding, the following uses a specific embodiment to explain the digital human generation process provided by the present application. Refer to Figure 3 , Figure 3Another schematic diagram of the digital human generation process provided by the embodiments of this application. This process includes the following steps:
[0188] S301: The server receives the first network parameters of the client and generates first grid parameters based on the first network parameters. Among them, the first network parameters include the bandwidth parameter of the client and / or the playback clarity of the digital human on the client.
[0189] Specifically, the first network parameters can be input into a grid parameter generation model (linear model) to generate the first grid parameters.
[0190] S302: The server generates a first point cloud digital human based on the first audio, generates a first grid digital human based on the first feature vector of the point cloud parameters of the first point cloud digital human and the second feature vector of the first grid parameters, and sends the first grid digital human to the client.
[0191] Specifically, the server can input the first audio into a sound and picture model to generate a first point cloud digital human, extract the point cloud parameters (the first point cloud parameters) of the first point cloud digital human, and encode the first point cloud parameters to obtain the feature vector corresponding to the first point cloud parameters; the server can also encode the first grid parameters to obtain the second feature vector corresponding to the first grid parameters; where the first grid digital human includes multiple grid units.
[0192] In one implementation, the server can determine the texture map corresponding to the first grid digital human. Among them, the resolution of the texture map corresponds to the number of grid units of the first grid digital human. The server can send the first grid digital human and the texture map corresponding to the first grid digital human to the client.
[0193] S303: The client renders the first grid digital human to obtain the target digital human.
[0194] For ease of understanding, the digital human generation process provided by this application will be further explained through a specific embodiment below. Refer to Figure 4 , Figure 4 Another schematic diagram of the digital human generation process provided by the embodiments of this application. This process includes the following steps:
[0195] S401: The server receives the first network parameters of the client, inputs the first network parameters into a grid parameter generation model, and generates first grid parameters. Among them, the first network parameters include the bandwidth parameter of the client and / or the playback clarity of the digital human on the client.
[0196] S402: The server generates a first point cloud digital human based on the first audio, generates a first mesh digital human based on the first feature vector of the point cloud parameters and the second feature vector of the first mesh parameters of the first point cloud digital human, and sends the first mesh digital human to the client.
[0197] S403: The client renders the first mesh digital human to obtain a target digital human.
[0198] S404: When the first network parameters of the client change, the server obtains the changed second network parameters of the client.
[0199] Specifically, when the client recognizes that the fluctuation range of the network bandwidth of the client exceeds the set threshold within the first duration, or the user changes the playback clarity selection of the digital human of the client, the changed second network parameters can be sent to the server (the server side).
[0200] S405: The server generates a second mesh digital human based on the first feature vector of the point cloud parameters and the second feature vector of the second mesh parameters of the first point cloud digital human, and sends the second mesh digital human to the client.
[0201] Among them, when the network bandwidth becomes larger, the number of mesh units of the second mesh digital human determined in S405 is greater than the number of mesh units of the first mesh digital human in S402, and vice versa; or when the user increases the playback clarity, the number of mesh units of the second mesh digital human determined in S405 is greater than the number of mesh units of the first mesh digital human in S402, and vice versa.
[0202] S406: The client receives the second mesh digital human sent by the server and renders the second mesh digital human to obtain a target digital human.
[0203] Embodiment 2:
[0204] Based on the same inventive concept as the method embodiment, the embodiment of the present application further provides a digital human generation device, which is used to execute the method executed by the server in the above method embodiment. For related features, reference can be made to the above method embodiment and will not be elaborated here. As Figure 5 shown, Figure 5 is a schematic structural diagram of a digital human generation device provided by the embodiment of the present application. The digital human generation device 500 includes:
[0205] A first generation module 501, configured to receive the first network parameters of the client and generate first mesh parameters based on the first network parameters;
[0206] A second generation module 502, configured to generate a first mesh digital human based on the first mesh parameters and the first point cloud parameters;
[0207] A sending module 503, configured to send the first mesh digital human to the client, so that the client renders the first mesh digital human to generate a target digital human; wherein, the first mesh digital human includes a plurality of mesh units in a three-dimensional space, and the first mesh parameter is used to characterize the number of mesh units of the first mesh digital human; the first point cloud parameter is used to characterize the three-dimensional space coordinates of a plurality of points on the point cloud digital human corresponding to the first mesh digital human, and the data volume of the point cloud digital human is greater than the data volume of the first mesh digital human.
[0208] In a possible implementation manner, the first generation module 501 is specifically configured to:
[0209] The first network parameter includes the bandwidth parameter of the client and / or the playback clarity of the digital human by the client, and inputs the first network parameter into a mesh parameter generation model to generate the first mesh parameter.
[0210] In a possible implementation manner, the second generation module 502 is specifically configured to:
[0211] Obtain a first audio, input the first audio into a sound and picture model to generate the first point cloud digital human;
[0212] Extract the first point cloud parameter of the first point cloud digital human, and generate a first mesh digital human based on the first mesh parameter and the first point cloud parameter.
[0213] In a possible implementation manner, the second generation module 502 is specifically configured to:
[0214] Encode the first mesh parameter and the first point cloud parameter respectively to generate a first feature vector of the first point cloud parameter and a second feature vector of the first mesh parameter;
[0215] Input the first feature vector and the second feature vector into a mesh digital human generation model to generate a first mesh digital human.
[0216] In a possible implementation manner, the sending module 503 is further configured to:
[0217] Send the first mesh digital human and the texture map corresponding to the first mesh digital human to the client, so that the client renders the first mesh digital human based on the texture map to generate the target digital human; wherein, the resolution of the texture map corresponds to the number of mesh units of the first mesh digital human.
[0218] In a possible implementation manner, the first generation module 501 is further configured to:
[0219] When the first network parameter of the client changes to a second network parameter, receive the second network parameter of the client, and generate a second grid parameter based on the second network parameter;
[0220] The second generation module 502 is further configured to:
[0221] Generate a second grid digital human based on the second grid parameter and the first point cloud parameter;
[0222] The sending module 503 is further configured to:
[0223] Send the second grid digital human to the client, so that the client renders the second grid digital human to generate a target digital human.
[0224] In a possible implementation manner, the change of the first network parameter of the client to the second network parameter includes: the bandwidth in the second network parameter is greater than the bandwidth in the first network parameter, or the playback clarity in the second network parameter is greater than the playback clarity in the first network parameter; or the bandwidth in the second network parameter is less than the bandwidth in the first network parameter, or the playback clarity in the second network parameter is less than the playback clarity in the first network parameter;
[0225] When the bandwidth in the second network parameter is greater than the bandwidth in the first network parameter, or the playback clarity in the second network parameter is greater than the playback clarity in the first network parameter, the number of grid cells of the second grid digital human is greater than the number of grid cells of the first grid digital human, and the data volume of the second grid digital human is greater than the data volume of the first grid digital human;
[0226] When the bandwidth in the second network parameter is less than the bandwidth in the first network parameter, or the playback clarity in the second network parameter is less than the playback clarity in the first network parameter, the number of grid cells of the second grid digital human is less than the number of grid cells of the first grid digital human, and the data volume of the second grid digital human is less than the data volume of the first grid digital human.
[0227] Based on the same inventive concept as the method embodiment, the embodiment of the present application further provides a model training device, which includes units or modules for executing the grid parameter generation model training method, or includes units or modules for executing the grid digital human generation model training method. For related features, reference can be made to the above method embodiment, and details are not described herein again.
[0228] As Figure 6 shownFigure 6 A schematic structural diagram of a model training device 600 provided by an embodiment of the present application. When the model training device executes the grid parameter generation model training method, the model training device 600 includes:
[0229] A first acquisition module 601, configured to acquire sample network parameters in a sample set, where the sample network parameters correspond to sample grid parameters;
[0230] A first prediction module 602, configured to input the sample network parameters into a grid parameter generation model to be trained to generate predicted grid parameters;
[0231] A first training module 603, configured to train the grid parameter generation model to be trained according to the sample grid parameters and the predicted grid parameters.
[0232] As Figure 7 shown, Figure 7 A schematic structural diagram of another model training device 700 provided by an embodiment of the present application. When the model training device executes the grid digital human generation model training method, the model training device 700 includes:
[0233] A second acquisition module 701, configured to acquire a first sample feature vector of sample point cloud parameters in a sample set, and a second sample feature vector of sample grid parameters corresponding to the first sample feature vector; wherein, the sample grid parameters are grid parameters of a sample grid digital human;
[0234] A second prediction module 702, configured to input the first sample feature vector and the second sample feature vector into a grid digital human generation model to be trained to generate a predicted grid digital human, extract predicted grid parameters of the predicted grid digital human, and encode the predicted grid parameters to generate a predicted feature vector of the predicted grid parameters;
[0235] A second training module 703, configured to train the grid digital human generation model to be trained according to the predicted feature vector and the second sample feature vector.
[0236] Based on the same inventive concept as the method embodiment, an embodiment of the present application further provides a digital human display device, which is configured to execute the method executed by the client in the above method embodiment. For related features, reference can be made to the above method embodiment, and details are not described herein again. As Figure 8 shown, Figure 8 A schematic structural diagram of a digital human display device provided by an embodiment of the present application. The digital human display device 800 includes:
[0237] A receiving module 801, configured to receive a first grid digital human sent by a server;
[0238] A rendering module 802, configured to render the first mesh digital human to generate a target digital human;
[0239] A display module 803, configured to display the target digital human;
[0240] Wherein, the first mesh digital human is generated based on first mesh parameters and first point cloud parameters; the first mesh parameters are used to characterize the number of mesh cells of the first mesh digital human; the first point cloud parameters are used to characterize the three-dimensional spatial coordinates of multiple points on the first point cloud digital human corresponding to the first mesh digital human, and the data volume of the point cloud digital human is greater than the data volume of the first mesh digital human.
[0241] In a possible implementation manner, the first mesh parameters are generated based on first network parameters of the client, and the first network parameters include the bandwidth parameter of the client and / or the playback clarity of the digital human by the client;
[0242] The receiving module 801 is further configured to receive a second mesh digital human sent by the server when the first network parameters of the client change to second network parameters;
[0243] The rendering module 802 is further configured to render the second mesh digital human to generate a target digital human;
[0244] The display module 803 is further configured to display the target digital human; wherein, the number of mesh cells of the second mesh digital human is different from the number of mesh cells of the first mesh digital human.
[0245] In a possible implementation manner, the bandwidth in the second network parameters is greater than the bandwidth in the first network parameters, or the playback clarity in the second network parameters is greater than the playback clarity in the first network parameters; or the bandwidth in the second network parameters is less than the bandwidth in the first network parameters, or the playback clarity in the second network parameters is less than the playback clarity in the first network parameters;
[0246] When the bandwidth in the second network parameters is greater than the bandwidth in the first network parameters, or the playback clarity in the second network parameters is greater than the playback clarity in the first network parameters, the number of mesh cells of the second mesh digital human is greater than the number of mesh cells of the first mesh digital human, and the data volume of the second mesh digital human is greater than the data volume of the first mesh digital human;
[0247] When the bandwidth in the second network parameter is less than the bandwidth in the first network parameter, or the playback clarity in the second network parameter is less than the playback clarity in the first network parameter, the number of grid units of the second grid digital human is less than the number of grid units of the first grid digital human, and the data volume of the second grid digital human is less than the data volume of the first grid digital human.
[0248] In a possible implementation manner, the rendering module 802 is specifically configured to:
[0249] Receive the texture map corresponding to the first grid digital human;
[0250] Render the first grid digital human based on the texture map to generate the target digital human; wherein, the resolution of the texture map corresponds to the number of grid units of the first grid digital human.
[0251] In a possible implementation manner, the target digital human includes a plurality of sub-target digital humans, and each sub-target digital human corresponds to each audio frame of the first audio;
[0252] The display module 803 is specifically configured to:
[0253] When playing each audio frame of the first audio, synchronously display the sub-target digital human corresponding to each audio frame.
[0254] Based on the same technical concept, the present application further provides an electronic device, Figure 9 which is a schematic structural diagram of an electronic device provided by an embodiment of the present application, as Figure 9 shown, including: a processor 901, a communication interface 902, a memory 903, and a communication bus 904, wherein the processor 901, the communication interface 902, and the memory 903 communicate with each other through the communication bus 904;
[0255] The memory 903 stores a computer program, and when the program is executed by the processor 901, the processor 901 is caused to execute the steps of any of the above method embodiments, which will not be elaborated herein.
[0256] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a thick line is used in the figure to represent it, but it does not mean that there is only one bus or one type of bus.
[0257] The communication interface 902 is used for communication between the above-mentioned electronic device and other devices.
[0258] The memory may include a Random Access Memory (RAM), or may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.
[0259] The above-mentioned processor may be a general-purpose processor, including a central processing unit, a Network Processor (NP), etc.; it may also be a Digital Signal Processing (DSP), an application-specific integrated circuit, a field-programmable gate array, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0260] Based on the same technical concept, this application also provides a digital human live broadcast system. Refer to Figure 10 , Figure 10 which is a schematic diagram of a digital human generation system provided by an embodiment of this application. The system includes a server 1001 in any of the above method embodiments and a client 1002 in any of the above method embodiments. Details will not be repeated for the repeated parts.
[0261] Based on the same technical concept, an embodiment of this application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program executable by an electronic device. When the program runs on the electronic device, it enables the electronic device to implement any of the above method embodiments when executed.
[0262] The above computer-readable storage medium may be any available medium or data storage device accessible by the processor in the electronic device, including but not limited to magnetic memories such as floppy disks, hard disks, magnetic tapes, magneto-optical discs (MO), etc., optical memories such as CDs, DVDs, BDs, HVDs, etc., and semiconductor memories such as ROMs, EPROMs, EEPROMs, non-volatile memories (NANDFLASH), solid-state drives (SSD), etc.
[0263] Based on the same technical concept, an embodiment of this application also provides a computer program product. The computer program product includes: computer program code. When the computer program code runs on a computer, it enables the computer to implement any of the above method embodiments. Since the principle of solving problems by the above computer program product is similar to that of the digital human generation method, the implementation of the above computer program product can refer to the implementation of the method. Details will not be repeated for the repeated parts.
[0264] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0265] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0266] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0267] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0268] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these changes and modifications.
Claims
1. A digital life generation method, characterized in that: The method includes: Receiving first network parameters of a client, and generating first grid parameters based on the first network parameters; Generating a first grid digital human based on the first grid parameters and first point cloud parameters; Sending the first grid digital human to the client so that the client renders the first grid digital human to generate a target digital human; wherein, the first grid digital human includes a plurality of grid units in three-dimensional space, and the first grid parameters are used to characterize the number of grid units of the first grid digital human; the first point cloud parameters are used to characterize the three-dimensional spatial coordinates of a plurality of points on a first point cloud digital human corresponding to the first grid digital human, and the data volume of the point cloud digital human is greater than the data volume of the first grid digital human.
2. The method according to claim 1, wherein The first network parameters include the bandwidth parameter of the client and / or the playback clarity of the digital human by the client. Generating the first grid parameters based on the first network parameters includes: Inputting the first network parameters into a grid parameter generation model to generate the first grid parameters.
3. The method according to claim 1 or 2, characterized in that: Generating the first grid digital human based on the first grid parameters and first point cloud parameters includes: Obtaining first audio, and inputting the first audio into an audio-visual model to generate the first point cloud digital human; Extracting the first point cloud parameters of the first point cloud digital human, and generating a first grid digital human based on the first grid parameters and the first point cloud parameters.
4. The method according to claim 3, wherein: Generating the first grid digital human based on the first grid parameters and the first point cloud parameters includes: Encoding the first grid parameters and the first point cloud parameters respectively to generate a first feature vector of the first point cloud parameters and a second feature vector of the first grid parameters; Inputting the first feature vector and the second feature vector into a grid digital human generation model to generate a first grid digital human.
5. The method according to any one of claims 1 to 4, characterized in that: The method further includes: Sending the first grid digital human and a texture map corresponding to the first grid digital human to the client so that the client renders the first grid digital human based on the texture map to generate the target digital human; wherein, the resolution of the texture map corresponds to the number of grid units of the first grid digital human.
6. The method according to any one of claims 1-5, characterized in that: The method further includes: When the first network parameters of the client change to second network parameters, receiving the second network parameters of the client, and generating second grid parameters based on the second network parameters; Generating a second grid digital human based on the second grid parameters and first point cloud parameters; Sending the second grid digital human to the client so that the client renders the second grid digital human to generate a target digital human; wherein the number of grid units of the second grid digital human is different from the number of grid units of the first grid digital human.
7. The method according to claim 6, characterized in that, The first network parameters of the client are changed to second network parameters, including: the bandwidth in the second network parameters is greater than the bandwidth in the first network parameters, or the playback clarity in the second network parameters is greater than the playback clarity in the first network parameters; or the bandwidth in the second network parameters is less than the bandwidth in the first network parameters, or the playback clarity in the second network parameters is less than the playback clarity in the first network parameters; In the case where the bandwidth in the second network parameters is greater than the bandwidth in the first network parameters, or the playback clarity in the second network parameters is greater than the playback clarity in the first network parameters, the number of grid units of the second grid digital human is greater than the number of grid units of the first grid digital human, and the data volume of the second grid digital human is greater than the data volume of the first grid digital human; In the case where the bandwidth in the second network parameters is less than the bandwidth in the first network parameters, or the playback clarity in the second network parameters is less than the playback clarity in the first network parameters, the number of grid units of the second grid digital human is less than the number of grid units of the first grid digital human, and the data volume of the second grid digital human is less than the data volume of the first grid digital human.
8. A method for training a grid parameter generation model, characterized in that, The method includes: Obtain sample network parameters in a sample set, where the sample network parameters correspond to sample grid parameters; Input the sample network parameters into a grid parameter generation model to be trained to generate predicted grid parameters; Train the grid parameter generation model to be trained according to the sample grid parameters and the predicted grid parameters.
9. A method for training a grid digital life generation model, characterized in that, The method includes: Obtain a first sample feature vector of sample point cloud parameters in a sample set, and a second sample feature vector of sample grid parameters corresponding to the first sample feature vector; wherein, the sample grid parameters are grid parameters of a sample grid digital human; Input the first sample feature vector and the second sample feature vector into a grid digital human generation model to be trained to generate a predicted grid digital human, extract the predicted grid parameters of the predicted grid digital human, and encode the predicted grid parameters to generate a predicted feature vector of the predicted grid parameters; Train the grid digital human generation model to be trained according to the predicted feature vector and the second sample feature vector.
10. A digital human display method, characterized in that: The method includes: Receive a first grid digital human sent by a server; Render the first grid digital human to generate a target digital human; Display the target digital human; Wherein, the first grid digital human is generated based on first grid parameters and first point cloud parameters; the first grid parameters are used to represent the number of grid units of the first grid digital human; the first point cloud parameters are used to represent the three-dimensional spatial coordinates of multiple points on a first point cloud digital human corresponding to the first grid digital human, and the data volume of the point cloud digital human is greater than the data volume of the first grid digital human.
11. The method according to claim 10, characterized in that, The first mesh parameter is generated based on the first network parameter of the client, and the first network parameter includes the bandwidth parameter of the client and / or the playback clarity of the digital human by the client; When the first network parameter of the client changes to a second network parameter, receive the second mesh digital human sent by the server; Render the second mesh digital human to generate a target digital human; Display the target digital human; wherein, the number of mesh units of the second mesh digital human is different from the number of mesh units of the first mesh digital human.
12. The method according to claim 11, wherein, The change of the first network parameter to the second network parameter includes: the bandwidth in the second network parameter is greater than the bandwidth in the first network parameter, or the playback clarity in the second network parameter is greater than the playback clarity in the first network parameter; or the bandwidth in the second network parameter is less than the bandwidth in the first network parameter, or the playback clarity in the second network parameter is less than the playback clarity in the first network parameter; When the bandwidth in the second network parameter is greater than the bandwidth in the first network parameter, or the playback clarity in the second network parameter is greater than the playback clarity in the first network parameter, the number of mesh units of the second mesh digital human is greater than the number of mesh units of the first mesh digital human, and the data volume of the second mesh digital human is greater than the data volume of the first mesh digital human; When the bandwidth in the second network parameter is less than the bandwidth in the first network parameter, or the playback clarity in the second network parameter is less than the playback clarity in the first network parameter, the number of mesh units of the second mesh digital human is less than the number of mesh units of the first mesh digital human, and the data volume of the second mesh digital human is less than the data volume of the first mesh digital human.
13. The method according to claim 10, wherein The rendering of the first mesh digital human to generate a target digital human includes: Receive the texture map corresponding to the first mesh digital human; Render the first mesh digital human based on the texture map to generate the target digital human; wherein, the resolution of the texture map corresponds to the number of mesh units of the first mesh digital human.
14. The method according to claim 10 or 13, characterized in that The target digital human includes a plurality of sub-target digital humans, and each sub-target digital human corresponds to each audio frame of the first audio; The display of the target digital human includes: When playing each audio frame of the first audio, synchronously display the sub-target digital human corresponding to each audio frame.
15. A computing device, characterized in that, The computing device includes a processor and a memory, and the processor is used to call the computer program instructions stored in the memory to execute the method according to any one of claims 1-7, the method for training the mesh parameter generation model according to claim 8, the method for training the mesh digital human generation model according to claim 9, or the method according to any one of claims 10-14.