Digital human video generation method and device, equipment, medium and product
By generating digital human videos based on user interaction instructions and using personalized rendering parameters to process response instructions, the problem of excessive computer resource usage when generating digital human videos is solved, and the user experience is improved.
Patent Information
- Application Number
- CN202511254383.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-03
- Publication Date
- 2025-10-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, when generating digital human videos, the uniform resolution and frame rate result in excessive computer resource usage, affecting user experience.
According to the interactive instructions input by the user, the digital human's response voice and response instructions are generated, the response instructions are rendered based on the preset rendering parameters, a video stream is generated, and the response voice and video stream are fused, and different rendering parameters are used to generate the digital human video.
Through personalized rendering parameter processing, computer resource usage is reduced and user experience is improved.
Smart Images

Figure CN120751201A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a method, apparatus, device, medium, and product for generating digital human videos. Background Art
[0002] In related technologies, when generating digital human videos, the digital human needs to be rendered. Usually, the entire digital human is rendered, but the uniform resolution and frame rate result in excessive computer resource usage. Summary of the Invention
[0003] This application mainly provides a method, device, equipment, medium and product for generating digital human video.
[0004] The technical solution of this application is achieved as follows: A method for generating a digital human video, the method comprising: Get the interactive instructions input by the user; Generate a response voice and at least one response instruction of the digital human based on the interaction instruction; the response instruction is used to instruct the rendering of the scene where the digital human is located or the body part of the digital human; Rendering at least one of the response instructions based on preset rendering parameters to generate at least one video stream of the digital human; The response voice is merged with at least one of the video streams to obtain the digital human video.
[0005] In the above solution, the response instruction includes one or more of the following: Facial lip shape response command; Respond to commands with body movements; Scene switching response command.
[0006] In the above solution, the step of rendering at least one of the response instructions based on preset rendering parameters to generate at least one video stream of the digital human includes: Determining a first rendering parameter corresponding to the facial lip shape response instruction from the preset rendering parameters; At a preset timestamp sequence, respond to the facial lip shape instruction based on the first rendering parameter Rendering is performed to generate a first video stream of the digital human.
[0007] In the above solution, the step of rendering at least one of the response instructions based on preset rendering parameters to generate at least one video stream of the digital human includes: Determining a second rendering parameter corresponding to the body movement response instruction from the preset rendering parameters; The body movement response instruction is rendered based on the second rendering parameter on a preset timestamp sequence to generate a second video stream of the digital human.
[0008] In the above solution, the step of rendering at least one of the response instructions based on preset rendering parameters to generate at least one video stream of the digital human includes: Determining a third rendering parameter corresponding to the scene switching response instruction from the preset rendering parameters; The scene switching response instruction is rendered based on the third rendering parameter on a preset timestamp sequence to generate a third video stream of the digital human.
[0009] In the above solution, the step of fusing the response voice with at least one of the video streams to obtain the digital human video includes: Generate a video vector of the digital human based on the preset rendering parameters and at least one of the response instructions on a preset timestamp sequence; fusing at least one of the video streams based on the video vector to obtain a fused video stream; The digital human video is generated based on the fused video stream and the response voice.
[0010] In the above solution, the fusing of at least one video stream based on the video vector to obtain a fused video stream includes: Comparing the resolution and frame rate of each of the video streams at a current timestamp based on the video vector to obtain a comparison result; Based on the comparison result, each of the video streams is aligned on the time axis to obtain the fused video stream.
[0011] A device for generating a digital human video, the device comprising: An acquisition unit, used to acquire an interaction instruction input by a user; A processing unit, configured to generate a response voice of the digital human and at least one response instruction based on the interaction instruction; the response instruction is used to instruct the rendering of the scene where the digital human is located or a body part of the digital human; The processing unit is configured to render at least one of the response instructions based on preset rendering parameters to generate at least one video stream of the digital human; The processing unit is used to fuse the response voice with at least one of the video streams to obtain the digital human video.
[0012] An electronic device comprising: a processor and a memory for storing a computer program capable of running on the processor, Wherein, the processor is used to execute the steps of any of the above methods when running the computer program.
[0013] A storage medium stores a computer program thereon, wherein the computer program implements the steps of any of the above methods when executed by a processor.
[0014] A computer product comprises a computer program, wherein when the computer program is executed by a processor, the steps of any one of the above methods are implemented.
[0015] The present invention provides a method, apparatus, device, medium, and product for generating a digital human video. The method comprises the following steps: obtaining an interactive instruction input by a user; generating a digital human's response voice and at least one response instruction based on the interactive instruction; the response instruction being used to instruct the rendering of the scene in which the digital human is located or a body part of the digital human; rendering at least one of the response instructions based on preset rendering parameters to generate at least one video stream of the digital human; and fusing the response voice with at least one of the video streams to obtain the digital human video. In other words, the present application generates a digital human's response voice and at least one response instruction based on the interactive instruction input by the user, renders at least one response instruction based on preset rendering parameters to generate at least one video stream of the digital human, and fuses the response voice with the at least one video stream to obtain the digital human video. This enables the generation of different response instructions based on the interactive instruction, the use of different rendering parameters when rendering different response instructions, and the generation of a digital human video by fusing at least one video stream. This solves the problem in the related art of using a uniform resolution and frame rate when rendering digital humans, which results in excessive computer resource usage. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 A schematic diagram of a flow chart of a method for generating a digital human video provided in an embodiment of the present application; Figure 2 A schematic flow chart of another method for generating a digital human video provided in an embodiment of the present application; Figure 3 A schematic diagram of a flow chart of a third method for generating a digital human video provided in an embodiment of the present application; Figure 4 A schematic flow chart of a fourth method for generating a digital human video provided in an embodiment of the present application; Figure 5 A schematic diagram of the structure of a device for generating a digital human video provided in an embodiment of the present application; Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0017] In order to make the purpose, technical solutions and advantages of this application clearer, the technical solutions of this application are further elaborated in detail below with reference to the accompanying drawings and embodiments. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0018] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0019] The terms "first / second / third" involved in this application are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It is understandable that "first / second / third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0021] In related technologies, when generating digital human videos, the digital human itself is mostly rendered independently. However, in most application scenarios, the digital human's facial lip shape and limbs require high-resolution, high-frame-rate, and fine-grained rendering, while background scenes and other scenes do not need to be rendered at a high frame rate. If a unified standard is used to generate virtual digital human videos, it will bring higher hardware configuration requirements and higher network overhead, which will seriously affect the user experience.
[0022] The embodiment of the present application provides a method for generating a digital human video, referring to Figure 1 As shown, the method includes the following steps: Step S101: Obtaining an interaction instruction input by a user.
[0023] It can be understood that the executing entity of this embodiment is a server. The server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0024] In practical applications, interactive instructions may include voice instructions and / or text instructions, as well as action instructions and scene switching instructions. Voice instructions and / or text instructions can be understood as instructions related to the user's intentions. The user can enter voice instructions by greeting or asking questions, such as "Hello!" or "What's the weather like today?" Action instructions can be understood as instructions specified by the user through the interactive interface requiring the virtual digital human to perform simulated actions, such as spreading hands, waving, turning around, smiling, etc. Scene switching instructions can be understood as instructions specified by the user through the interactive interface for the virtual digital human to switch scenes. Scene switching instructions include but are not limited to changes in the user's observation perspective, switching scene types, and changes in embedded virtual objects in the scene.
[0025] Step S102: Generate a digital human's response voice and at least one response instruction based on the interaction instruction.
[0026] The response instruction is used to instruct the rendering of the scene where the digital human is located or the body part of the digital human.
[0027] It is understandable that the server may include a voice processing module and a controller module. The voice processing module understands the user's intention based on voice instructions and / or text instructions and generates a response voice. For example, the voice processing module may use a pre-trained mapping model to map voice tag parameters to response parameters and generate a response voice based on the response parameters. The voice processing module may also call an artificial intelligence question-answering interface to generate a response voice based on voice instructions and / or text instructions. , generate corresponding facial lip shape response instructions according to the response voice F, F =[ T , ], Can be understood as the current timestamp T The corresponding facial and lip shape response command. The body parts of the digital human can include but are not limited to limbs, lips, and face.
[0028] In actual applications, in addition to the body movement response instructions and scene switching response instructions input by the user, voice instructions and text instructions can also generate body movement response instructions M, M =[ T , ], Can be understood as the current timestamp T Corresponding body movement response instructions can also generate scene switching response instructions S, S =[ T , ], Can be understood as the current timestamp TThe speech processing module can send body movement response commands, scene change response commands, and facial lip shape response commands to the controller module for processing. Response commands can be used to trigger the virtual digital human's body movements, lip shapes, facial expressions, and scene changes.
[0029] Step S103: Render at least one response instruction based on preset rendering parameters to generate at least one video stream of the digital human.
[0030] In practical applications, reference Figure 2 As shown, the controller module categorizes received commands into the following categories: the first category is user-entered action commands and scene switching commands; the second category is body movement response commands, scene switching response commands, and facial lip shape response commands generated based on voice commands and text commands. Preset rendering parameters can include resolution and frame rate corresponding to different types of commands.
[0031] After these instructions are classified by the controller module, a video vector (resolution-frame) corresponding to the timestamp is generated based on the preset resolution and frame rate requirements of different types of instructions. =[T,[ , , , , , ]], together with the response command, is sent to the module that executes these commands for parallel processing. Response command for scene switching at the current timestamp T The corresponding resolution, The current timestamp for the body movement response command T The corresponding resolution, The lip gesture response command at the current timestamp T The corresponding resolution, The current timestamp for the body movement response command T The corresponding frame rate, The current timestamp for the body movement response command T The corresponding frame rate, The lip gesture response command at the current timestamp T The corresponding frame rate.
[0032] The server may further include an action instruction module, a facial lip shape driving module, and a scene switching module, each of which receives a response instruction sent by the controller module and renders the response instruction according to rendering parameters corresponding to the response instruction to obtain a corresponding video stream. The preset rendering parameters include rendering parameters corresponding to different response instructions.
[0033] Step S104: Fusing the response voice with at least one video stream to obtain a digital human video.
[0034] In practical applications, the multiple video streams generated can be improved to the highest preset resolution and frame rate based on the video vector using dynamic interpolation and interpolation or super-resolution, and then fused and combined with response instructions to generate digital human videos.
[0035] As can be seen from the above content, the embodiment of the present application generates a digital human's response voice and at least one response instruction based on the interactive instructions input by the user, renders the at least one response instruction according to preset rendering parameters, generates at least one video stream of the digital human, and fuses the response voice with the at least one video stream to obtain a digital human video, so as to realize the generation of different response instructions according to the interactive instructions, use different rendering parameters when rendering different response instructions, and fuse at least one video stream to generate a digital human video, thereby solving the problem of excessive computer resource occupation caused by the use of a unified resolution and frame rate when rendering digital humans in related technologies.
[0036] In some embodiments of the present application, the response instruction includes one or more of the following: Facial lip shape response command; Respond to commands with body movements; Scene switching response command.
[0037] In actual applications, in addition to the body movement response instructions and scene switching response instructions input by the user, voice instructions and text instructions can also generate body movement response instructions M, M =[ T , ] and scene switching response command S, S =[ T , ], the voice processing module can send body movement response instructions, scene switching response instructions, and facial lip shape response instructions to the controller module for processing. The response instructions can be used to trigger the virtual digital human's body movements, lip shapes, facial expressions, and scene switching.
[0038] In some embodiments of the present application, rendering at least one response instruction based on preset rendering parameters to generate at least one video stream of the digital human includes: Determining a first rendering parameter corresponding to the facial lip shape response instruction from preset rendering parameters; At a preset timestamp sequence, respond to the facial lip shape instruction based on the first rendering parameter Rendering is performed to generate the first video stream of the digital human.
[0039] In practical applications, the first rendering parameter can be understood as the resolution and frame rate corresponding to the facial lip shape response instruction. Figure 3 As shown, the Facial Lip Shaping Driver Module receives the Facial Lip Shaping Response Command F from the Controller Module. Based on the timestamp T, it searches for the corresponding facial lip shape image in the preset animation and renders it into a partial animation in chronological order. The Facial Lip Shaping Response Command can be rendered at the highest preset frame rate and resolution.
[0040] In some embodiments of the present application, rendering at least one response instruction based on preset rendering parameters to generate at least one video stream of the digital human includes: Determining a second rendering parameter corresponding to the body movement response instruction from the preset rendering parameters; On a preset timestamp sequence, the body movement response instruction is rendered based on the second rendering parameter to generate a second video stream of the digital human.
[0041] In practical applications, the second rendering parameter can be understood as the resolution and frame rate corresponding to the body movement response instruction. Figure 3 As shown, the action instruction module receives the body movement response command M issued by the controller module. It is understandable that action switching does not occur continuously during virtual digital human application interaction. Therefore, the preset frame rate and resolution of the action instructions can be lower than those corresponding to facial expressions and lip movements to reduce the software and hardware overhead required for rendering. After processing the action M corresponding to timestamp T, a copy of the frame rendered at that moment can be used to fill the time until the timestamp of the next action M', and then the new action M' can be processed.
[0042] In some embodiments of the present application, rendering at least one response instruction based on preset rendering parameters to generate at least one video stream of the digital human includes: Determining a third rendering parameter corresponding to the scene switching response instruction from the preset rendering parameters; On a preset timestamp sequence, the scene switching response instruction is rendered based on a third rendering parameter to generate a third video stream of the digital human.
[0043] In practical applications, the third rendering parameter can be understood as the resolution and frame rate corresponding to the scene switching response instruction. Figure 3 As shown, the scene switching module receives the scene switching response instruction S distributed by the controller module. In the virtual digital human application, scene switching may include user perspective switching. Switching between virtual objects in the scene . The user perspective switching can be pulling in, back, moving up, moving down, rotating around, etc.; the switching of virtual objects can be the movement of objects in the scene, such as raising and lowering curtains, switching background images, etc. Similar to the execution method of the action instruction module, scene switching does not occur continuously in the application interaction of virtual digital people. Therefore, the preset frame rate and resolution of scene switching can be lower than the corresponding values of facial lip movements to reduce the software and hardware overhead required for rendering. After processing the instruction S corresponding to the timestamp T, the copy frame of the frame rendered at that moment can be used to fill in until the timestamp time point of the next instruction S', and then the new instruction S' can be processed.
[0044] refer to Figure 3 As shown in the figure, the controller module and each action processing module are encapsulated into independent container images and run in a cluster environment built by a container orchestration system (such as Kubernetes). This can achieve clustering capabilities of multiple copies on a single host and multiple copies on multiple hosts, thereby improving the clustered control output effect of the digital human.
[0045] In some embodiments of the present application, the response voice is merged with at least one video stream to obtain a digital human video, including: Generate a video vector of the digital human based on preset rendering parameters and at least one response instruction on a preset timestamp sequence; fusing at least one video stream based on the video vector to obtain a fused video stream; Digital human video is generated based on the fused video stream and response voice.
[0046] In actual applications, each action processing module outputs multiple video streams, which can be combined with video vectors to fuse multiple videos according to preset coordinates. During fusion, dynamic interpolation and interpolation super-resolution can be used to normalize video streams with different resolutions and frame rates to the highest resolution and frame rate video stream for playback.
[0047] In some embodiments of the present application, fusing at least one video stream based on a video vector to obtain a fused video stream includes: Comparing the resolution and frame rate of each video stream at the current timestamp based on the video vector to obtain a comparison result; Based on the comparison results, each video stream is aligned on the time axis to obtain a fused video stream.
[0048] In practical applications, after receiving the vector Then, the parameters are applied to the three video streams respectively, according to the resolution r, frame rate f and the highest preset resolution of each video stream at the current timestamp T , frame rate In comparison, dynamic frame insertion and interpolation or super-resolution are used to increase the video stream to the highest preset resolution and frame rate. The three-way video stream can be expressed as follows: (Formula 1) is the video stream of facial lip shape, The highest preset resolution for the current timestamp of the video stream of the face lip shape, The highest preset resolution frame rate of the current timestamp of the video stream of the facial lip shape; For the video stream of body movements, The highest preset resolution for the current timestamp of the video stream of the body movement, The highest preset resolution frame rate of the current timestamp of the video stream of the body movement; Video stream for scene switching, The highest preset resolution of the current timestamp of the video stream for scene switching. The highest preset resolution frame rate of the video stream at the current timestamp for scene switching.
[0049] The final video output is obtained by performing frame shift calculation on the corresponding reference initial position of each video stream frame, as shown in Formula 2: (Formula 2) is the frame displacement corresponding to the video stream of the facial lip shape, Frame displacement corresponding to the video stream of body movements and the frame displacement corresponding to the video stream of scene switching When all video streams are aligned on the timeline based on the same benchmark, multiple video streams are merged and the fused video stream is output. .
[0050] In a realistic scenario, refer to Figure 4 As shown, the method for generating a digital human video in the embodiment of the present application can be implemented in the following manner: Step S401: receiving user interaction input.
[0051] Step S402: The interactive input is processed by artificial intelligence voice to generate response language, triggering response actions, scene switching, and facial lip movements.
[0052] Step S403: The controller distributes different task instructions to process scene changes, body movements, and facial lip movements respectively.
[0053] Step S404: The client receives multiple video streams and performs overlay synthesis. This step can be understood as the server performing rendering processing through different executors and outputting video streams with different resolutions and frame rates, which are finally fused into a complete video on the client.
[0054] Based on the same inventive concept as above, Figure 5 This is a schematic diagram of the structure of a device for generating a digital human video provided by an embodiment of the present invention. The device 500 includes: An acquisition unit 501 is configured to acquire an interaction instruction input by a user; Processing unit 502, configured to generate a response voice of the digital human and at least one response instruction based on the interaction instruction; the response instruction is used to instruct the rendering of the scene where the digital human is located or the body part of the digital human; The processing unit 502 is configured to render at least one response instruction based on preset rendering parameters to generate at least one video stream of the digital human; The processing unit 502 is configured to fuse the response voice with at least one video stream to obtain a digital human video.
[0055] In some embodiments of the present application, the response instruction includes one or more of the following: Facial lip shape response command; Respond to commands with body movements; Scene switching response command.
[0056] In some embodiments of the present application, the processing unit 502 is configured to determine a first rendering parameter corresponding to the facial lip shape response instruction from preset rendering parameters; The facial lip shape response instruction is rendered based on the first rendering parameter on a preset timestamp sequence to generate a first video stream of the digital human.
[0057] In some embodiments of the present application, the processing unit 502 is configured to determine a second rendering parameter corresponding to the body movement response instruction from the preset rendering parameters; On a preset timestamp sequence, the body movement response instruction is rendered based on the second rendering parameter to generate a second video stream of the digital human.
[0058] In some embodiments of the present application, the processing unit 502 is configured to determine a third rendering parameter corresponding to the scene switching response instruction from preset rendering parameters; On a preset timestamp sequence, the scene switching response instruction is rendered based on a third rendering parameter to generate a third video stream of the digital human.
[0059] In some embodiments of the present application, the processing unit 502 is configured to generate a video vector of the digital human based on preset rendering parameters and at least one response instruction on a preset timestamp sequence; fusing at least one video stream based on the video vector to obtain a fused video stream; Digital human video is generated based on the fused video stream and response voice.
[0060] In some embodiments of the present application, the processing unit 502 is configured to compare the resolution and frame rate of each video stream at a current timestamp based on the video vector to obtain a comparison result; Based on the comparison results, each video stream is aligned on the time axis to obtain a fused video stream.
[0061] Based on the above embodiments, an embodiment of the present application provides an electronic device, Figure 6 This is a hardware structure diagram of an electronic device according to an embodiment of the present invention. The electronic device 600 includes: at least one processor 601, a memory 602, and optionally, the electronic device 600 may further include at least one communication interface 603. The various components in the electronic device 600 are coupled together through a bus system 604. It can be understood that the bus system 604 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 604 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, Figure 6 Various buses are labeled as bus system 604 .
[0062] Based on the hardware implementation of the above program modules, the communication interface 603 can exchange information with other communication devices; The processor 601 is connected to the communication interface 603 to implement information exchange with other communication devices, and is used to execute the methods provided by one or more of the above technical solutions when running a computer program; The computer program is stored in the memory 602 .
[0063] Specifically, the processor 601 is configured to obtain an interaction instruction input by a user; Generate the digital human's response voice and at least one response instruction based on the interactive instruction; the response instruction is used to instruct the rendering of the scene where the digital human is located or the body part of the digital human; Rendering at least one response instruction based on preset rendering parameters to generate at least one video stream of the digital human; The answering voice is fused with at least one video stream to obtain a digital human video.
[0064] In some embodiments of the present application, the response instruction includes one or more of the following: Facial lip shape response command; Respond to commands with body movements; Scene switching response command.
[0065] In some embodiments of the present application, the processor 601 is configured to determine a first rendering parameter corresponding to the facial lip shape response instruction from preset rendering parameters; At a preset timestamp sequence, respond to the facial lip shape instruction based on the first rendering parameter Rendering is performed to generate the first video stream of the digital human.
[0066] In some embodiments of the present application, the processor 601 is configured to determine a second rendering parameter corresponding to the body movement response instruction from preset rendering parameters; On a preset timestamp sequence, the body movement response instruction is rendered based on the second rendering parameter to generate a second video stream of the digital human.
[0067] In some embodiments of the present application, the processor 601 is configured to determine a third rendering parameter corresponding to the scene switching response instruction from preset rendering parameters; On a preset timestamp sequence, the scene switching response instruction is rendered based on a third rendering parameter to generate a third video stream of the digital human.
[0068] In some embodiments of the present application, the processor 601 is configured to generate a video vector of the digital human based on preset rendering parameters and at least one response instruction on a preset timestamp sequence; fusing at least one video stream based on the video vector to obtain a fused video stream; Digital human video is generated based on the fused video stream and response voice.
[0069] In some embodiments of the present application, the processor 601 is configured to compare the resolution and frame rate of each video stream at a current timestamp based on the video vector to obtain a comparison result; Based on the comparison results, each video stream is aligned on the time axis to obtain a fused video stream.
[0070] It is understood that memory 602 can be volatile memory or non-volatile memory, or can include both volatile and non-volatile memory. Non-volatile memory can include read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic random access memory (FRAM), flash memory, magnetic surface storage, optical disk, or compact disc read-only memory (CD-ROM); magnetic surface storage can include magnetic disk storage or magnetic tape storage. Volatile memory can include random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Synchronous Static Random Access Memory (SSRAM), Dynamic Random Access Memory (DRAM), Synchronous Dynamic Random Access Memory (SDRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDRSDRAM), Enhanced Synchronous Dynamic Random Access Memory (ESDRAM), Sync Link Dynamic Random Access Memory (SLDRAM), and Direct Rambus Random Access Memory (DRRAM).The memory 602 described in the embodiments of the present invention is intended to include, but is not limited to, these and any other suitable types of memory.
[0071] The memory 602 in the embodiment of the present invention is used to store various types of data to support the operation of the electronic device 600. Examples of such data include any computer program for operating on the electronic device 600. The program for implementing the method of the embodiment of the present invention may be included in the memory 602.
[0072] The methods disclosed in the above embodiments of the present invention can be applied to or implemented by processor 601. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits in the processor or by software instructions. The above processor may be a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of the present invention can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module may be located in a storage medium located in a memory. The processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method.
[0073] In an exemplary embodiment, the electronic device 600 may be implemented by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers (MCUs), microprocessors, or other electronic components to perform the above method.
[0074] Based on the above embodiments, the embodiments of the present application further provide a computer product, including a computer program, which is implemented when the computer program is executed by a processor. Figure 1 The corresponding embodiment provides steps in the method for generating a digital human video.
[0075] Based on the above embodiments, the embodiments of the present application further provide a storage medium, wherein the storage medium stores computer executable instructions, and the computer executable instructions are configured to execute Figure 1 The corresponding embodiment provides a method for generating a digital human video.
[0076] It should be noted that the above-mentioned computer storage medium can be a memory such as ROM, PROM, EPROM, EEPROM, FRAM, FlashMemory, magnetic surface storage, optical disk, or CD-ROM; it can also be various electronic devices that include one or any combination of the above-mentioned memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.
[0077] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0078] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0079] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a more preferred embodiment. Based on this understanding, the technical solution of this application, or the part that contributes to the existing technology, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of this application.
[0080] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0081] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0082] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0083] The above are only preferred embodiments of the present application and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A method for generating a digital human video, characterized in that: The method comprises: Get the interactive instructions input by the user; Generate a response voice and at least one response instruction of the digital human based on the interaction instruction; the response instruction is used to instruct the rendering of the scene where the digital human is located or the body part of the digital human; Rendering at least one of the response instructions based on preset rendering parameters to generate at least one video stream of the digital human; The response voice is merged with at least one of the video streams to obtain the digital human video.
2. The method according to claim 1, characterized in that The response instruction includes one or more of the following: Facial lip shape response command; Respond to commands with body movements; Scene switching response command.
3. The method according to claim 2, characterized in that The step of rendering at least one of the response instructions based on preset rendering parameters to generate at least one video stream of the digital human comprises: Determining a first rendering parameter corresponding to the facial lip shape response instruction from the preset rendering parameters; At a preset timestamp sequence, respond to the facial lip shape instruction based on the first rendering parameter Rendering is performed to generate a first video stream of the digital human.
4. The method according to claim 2, characterized in that The step of rendering at least one of the response instructions based on preset rendering parameters to generate at least one video stream of the digital human comprises: Determining a second rendering parameter corresponding to the body movement response instruction from the preset rendering parameters; The body movement response instruction is rendered based on the second rendering parameter on a preset timestamp sequence to generate a second video stream of the digital human.
5. The method according to claim 2, characterized in that The step of rendering at least one of the response instructions based on preset rendering parameters to generate at least one video stream of the digital human comprises: Determining a third rendering parameter corresponding to the scene switching response instruction from the preset rendering parameters; The scene switching response instruction is rendered based on the third rendering parameter on a preset timestamp sequence to generate a third video stream of the digital human.
6. The method according to claim 1, characterized in that The step of fusing the response voice with at least one of the video streams to obtain the digital human video includes: Generate a video vector of the digital human based on the preset rendering parameters and at least one of the response instructions on a preset timestamp sequence; fusing at least one of the video streams based on the video vector to obtain a fused video stream; The digital human video is generated based on the fused video stream and the response voice.
7. The method according to claim 6, characterized in that The fusing at least one of the video streams based on the video vector to obtain a fused video stream includes: Comparing the resolution and frame rate of each of the video streams at a current timestamp based on the video vector to obtain a comparison result; Based on the comparison result, each of the video streams is aligned on the time axis to obtain the fused video stream.
8. A device for generating a digital human video, characterized in that: The device comprises: An acquisition unit, used to acquire an interaction instruction input by a user; A processing unit, configured to generate a response voice of the digital human and at least one response instruction based on the interaction instruction; the response instruction is used to instruct the rendering of the scene where the digital human is located or a body part of the digital human; The processing unit is configured to render at least one of the response instructions based on preset rendering parameters to generate at least one video stream of the digital human; The processing unit is used to fuse the response voice with at least one of the video streams to obtain the digital human video.
9. An electronic device, characterized in that: include: a processor and a memory for storing a computer program capable of being executed on the processor, Wherein, when the processor is used to run the computer program, it executes the steps of the method according to any one of claims 1 to 7.
10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
11. A computer product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Scene rendering method and device, storage medium and electronic equipment
CN113850898A
Human-computer interaction method and device based on digital human, electronic equipment and storage medium
CN113901190A
Digital human real-time interaction system and digital human real-time interaction method
CN119440254A
System and method for an interactive digitally rendered avatar of a subject person
US20220150287A1
Cited By
Virtual digital human rendering method, device and system and electronic equipment
CN121309936A
Digital human generation and interaction method, equipment and medium
CN121547664A
A digital human generation and interaction method, device and medium
CN121547664B