Computer system, local device and cloud device for digital human rendering

Through computer systems that support local and cloud rendering modes, the problems of time-consuming and labor-intensive digital rendering and single business needs are solved, and flexible digital rendering and high-quality rendering effects are achieved.

CN120378649APending Publication Date: 2025-07-25MOORE THREADS TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510696155.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The prior art is time-consuming and labor-intensive in digital human rendering and is difficult to meet a variety of business needs, especially in terms of low latency and high precision.

Method used

It provides a computer system that supports local rendering mode and cloud rendering mode, and realizes digital human rendering through the collaborative work of local devices and cloud devices.

Benefits of technology

Meet business needs in different scenarios, improve the flexibility and interactivity of digital human rendering, reduce latency and improve rendering quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120378649A_ABST
    Figure CN120378649A_ABST
Patent Text Reader

Abstract

The invention discloses a computer system for digital human rendering, local equipment and cloud equipment, and belongs to the field of digital human. The computer system comprises local equipment and cloud equipment, and the computer system supports a local rendering mode and a cloud rendering mode. The local equipment is used for sending input information to the cloud equipment; the cloud equipment is used for generating rendering related information according to the input information; in the local rendering mode, the cloud device is used for sending rendering related information to the local device, and the local device is used for obtaining video stream data of the digital human through rendering according to the rendering related information; and in the cloud rendering mode, the cloud device is used for rendering according to the rendering related information to obtain the video stream data of the digital person, and the local device is used for receiving the video stream data of the digital person. The computer system supports two rendering modes, rendering is performed by the local device in the local rendering mode, and rendering is performed by the cloud device in the cloud rendering mode, so that service requirements in different scenes are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of digital humans, and particularly to a computer system, a local device, a cloud device, a digital human rendering method, a medium, and a program product for digital human rendering. Background Art

[0002] With the development of artificial intelligence technology, the creation and application of digital humans have become increasingly common. A digital human is a virtual human image created using digital technology.

[0003] In related technologies, a model of a digital human is obtained using 3D scanning technology, and then the facial expressions and body movements of the digital human are optimized through a large number of manual operations.

[0004] However, this technology is not only time-consuming and laborious, but also usually can only meet a single business requirement and cannot meet multiple business requirements in the case of different business requirements such as digital humans that require low latency or high precision. Summary of the Invention

[0005] The present application provides a computer system, a local device, a cloud device, a digital human rendering method, a medium, and a program product for digital human rendering, and the technical solution at least includes at least one of the following aspects.

[0006] According to one aspect of the embodiments of the present application, there is provided a computer system for digital human rendering, the computer system includes a local device and a cloud device, and the computer system supports a local rendering mode and a cloud rendering mode;

[0007] The local device is configured to send input information to the cloud device, and the input information is used for the cloud device to generate rendering-related information associated with the digital human through a digital human generation service;

[0008] The cloud device is configured to generate rendering-related information according to the input information, and the rendering-related information is used to provide information required for rendering the digital human;

[0009] In the local rendering mode, the cloud device is configured to send the rendering-related information to the local device, and the local device is configured to render the video stream data of the digital human according to the rendering-related information;

[0010] In the cloud rendering mode, the cloud device is configured to render the video stream data of the digital human according to the rendering-related information, and the local device is configured to receive the video stream data of the digital human.

[0011] According to another aspect of the embodiments of the present application, a local device for digital human rendering is provided. The local device supports a local rendering mode and a cloud rendering mode. The local device includes a client module, or includes a client module and a rendering engine module;

[0012] The client module is configured to send input information to a cloud device. The input information is used for the cloud device to generate rendering-related information associated with the digital human through a digital human generation service module;

[0013] In the local rendering mode, the client module is configured to receive the rendering-related information sent by the cloud device. The rendering-related information is used to provide the information required for rendering the digital human;

[0014] The rendering engine module is configured to render the video stream data of the digital human according to the rendering-related information, and send the video stream data of the digital human to the client module;

[0015] In the cloud rendering mode, the client module is configured to receive the video stream data of the digital human. The video stream data of the digital human is rendered by the cloud device according to the rendering-related information.

[0016] According to another aspect of the embodiments of the present application, a cloud device for digital human rendering is provided. The cloud device supports a local rendering mode and a cloud rendering mode. The cloud device includes an access gateway module and a digital human generation service module, or includes an access gateway module, a digital human generation service module and a rendering engine module;

[0017] The digital human generation service module is configured to receive the input information sent by the local device. The input information is used for the cloud device to generate rendering-related information associated with the digital human through the digital human generation service module;

[0018] The digital human generation service module is configured to generate rendering-related information according to the input information. The rendering-related information is used to provide the information required for rendering the digital human;

[0019] In the local rendering mode, the access gateway module is configured to send the rendering-related information to the local device. The rendering-related information is used for the local device to render the video stream data of the digital human;

[0020] In the cloud rendering mode, the rendering engine module is configured to render the video stream data of the digital human according to the rendering-related information, and send the video stream data of the digital human to the local device.

[0021] According to one aspect of the embodiments of the present application, a digital human rendering method is provided. The method is executed by a local device. The local device supports a local rendering mode and a cloud rendering mode. The method includes:

[0022] Send input information to a cloud device, where the input information is used for the cloud device to generate rendering-related information associated with a digital human through a digital human generation service;

[0023] In the local rendering mode, receive the rendering-related information sent by the cloud device, where the rendering-related information is used to provide the information required for rendering the digital human; Render the video stream data of the digital human according to the rendering-related information;

[0024] In the cloud rendering mode, receive the video stream data of the digital human, where the video stream data of the digital human is rendered by the cloud device according to the rendering-related information.

[0025] According to another aspect of the embodiments of the present application, there is provided a digital human rendering method, which is executed by a cloud device. The cloud device supports the local rendering mode and the cloud rendering mode. The method includes:

[0026] Receive the input information sent by the local device, where the input information is used for the cloud device to generate rendering-related information associated with the digital human through a digital human generation service;

[0027] Generate rendering-related information according to the input information, where the rendering-related information is used to provide the information required for rendering the digital human;

[0028] In the local rendering mode, send the rendering-related information to the local device, where the rendering-related information is used for the local device to render the video stream data of the digital human;

[0029] In the cloud rendering mode, render the video stream data of the digital human according to the rendering-related information, and send the video stream data of the digital human to the local device.

[0030] According to another aspect of the embodiments of the present application, there is provided a computer-readable storage medium, in which at least one segment of program is stored, and the at least one segment of program is loaded and executed by a processor to implement the above digital human rendering method.

[0031] According to another aspect of the embodiments of the present application, there is provided a computer program product or a computer program, which includes computer instructions. The computer instructions are stored in a computer-readable storage medium, and the processor obtains the computer instructions from the computer-readable storage medium and executes the computer instructions to implement the above digital human rendering method.

[0032] The technical solutions provided by the embodiments of the present application may include the following beneficial effects:

[0033] The computer system supports two rendering modes, namely the local rendering mode and the cloud rendering mode. In the local rendering mode, the local device performs rendering, and in the cloud rendering mode, the cloud device performs rendering, thereby meeting the business requirements in different scenarios and solving the problem that related technologies usually can only meet single business requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] To more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the following described drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0035] Figure 1 FIG. shows a schematic diagram of a computer system for digital human rendering provided by an exemplary embodiment of the present application;

[0036] Figure 2 FIG. shows a schematic diagram of digital human rendering provided by an exemplary embodiment of the present application;

[0037] Figure 3 FIG. shows a flowchart of a computer system for digital human rendering provided by an exemplary embodiment of the present application;

[0038] Figure 4 FIG. shows a flowchart of generating rendering-related information provided by an exemplary embodiment of the present application;

[0039] Figure 5 FIG. shows a flowchart of generating rendering-related information provided by an exemplary embodiment of the present application;

[0040] Figure 6 FIG. shows a flowchart of generating rendering-related information provided by an exemplary embodiment of the present application;

[0041] Figure 7 FIG. shows a flowchart of generating rendering-related information provided by an exemplary embodiment of the present application;

[0042] Figure 8 FIG. shows a schematic diagram of a digital human generation service provided by an exemplary embodiment of the present application;

[0043] Figure 9 FIG. shows a flowchart of a computer system for digital human rendering provided by an exemplary embodiment of the present application;

[0044] Figure 10 FIG. shows a schematic diagram of switching from the cloud rendering mode to the local rendering mode provided by an exemplary embodiment of the present application;

[0045] Figure 11 Shows a schematic diagram of switching from the local rendering mode to the cloud rendering mode provided by an exemplary embodiment of the present application;

[0046] Figure 12 Shows a schematic diagram of a local device for digital human rendering provided by an exemplary embodiment of the present application;

[0047] Figure 13 Shows a schematic diagram of a cloud device for digital human rendering provided by an exemplary embodiment of the present application;

[0048] Figure 14 Shows a flowchart of a digital human rendering method provided by an exemplary embodiment of the present application;

[0049] Figure 15 Shows a flowchart of a digital human rendering method provided by an exemplary embodiment of the present application;

[0050] Figure 16 Shows a flowchart of a digital human rendering method provided by an exemplary embodiment of the present application;

[0051] Figure 17 Shows a flowchart of a digital human rendering method provided by an exemplary embodiment of the present application;

[0052] Figure 18 Shows a flowchart of a digital human rendering method provided by an exemplary embodiment of the present application;

[0053] Figure 19 Shows a structural block diagram of a computer device provided by an exemplary embodiment of the present application. Detailed implementation manners

[0054] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.

[0055] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0056] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. The singular forms "a", "the", and "said" used in this application and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0057] It should be noted that the object information (including but not limited to object device information, object personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the object or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0058] It should be understood that although the terms first, second, etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of this application, the first parameter may also be referred to as the second parameter, and similarly, the second parameter may also be referred to as the first parameter. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".

[0059] First, some of the terms involved in this application are explained as follows:

[0060] Digital Human: It refers to a virtual human image created by using digital technology that is close to the human image. A digital human appears as a two-dimensional image or a three-dimensional model and can have functions such as speech recognition and natural language processing.

[0061] Full-functional Graphics Processing Unit (GPU): It refers to a GPU with complete functions. Such a GPU can not only provide graphics rendering capabilities, video encoding and decoding capabilities, and ultra-high-definition display capabilities, but also execute general computing tasks, and is mainly applied in fields such as artificial intelligence, graphics rendering, multimedia, deep learning, and high-performance computing.

[0062] With the development of artificial intelligence technology, the creation and application of digital humans have become more and more common. A digital human refers to a virtual human image created by using digital technology. In related technologies, a model of a digital human is obtained by using three-dimensional scanning technology, and then the facial expressions and limb movements of the digital human are optimized through a large number of manual operations. However, this technology is not only time-consuming and laborious, but also due to the related technologies coming from different companies, there are problems such as high latency and unstable services for digital humans, and it cannot meet the business requirements.

[0063] To solve the above problems, the present application proposes a computer system for digital human rendering. Figure 1 FIG. shows a schematic diagram of a computer system for digital human rendering provided by an exemplary embodiment of the present application. The computer system 10 includes: a local device 100 and a cloud device 200. The computer system supports a local rendering mode and a cloud rendering mode. Figure 1 FIG. (a) shows the computer system 10 in the cloud rendering mode. Figure 1 FIG. (b) shows the computer system 10 in the local rendering mode.

[0064] The local device 100 may be a computer device such as a mobile phone, a tablet computer, an in-vehicle terminal (car computer), a wearable device, a personal computer, etc. The cloud device 200 may be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers. It may also be a cloud server providing services such as cloud database, cloud computing, cloud function, cloud storage, network service, cloud communication, middleware service, domain name service, security service, content delivery network, and big data and artificial palm image recognition platform.

[0065] The local device 100 and the cloud device 200 can communicate with each other through a network, such as a wired network or a wireless network.

[0066] For Figure 1 FIGS. (a) and (b), the local device 100 runs a client 110, and the cloud device 200 runs an access gateway 210 and a digital human generation service 220. The difference is that in Figure 1 FIG. (a), the cloud device 200 also runs a rendering engine 230; in Figure 1 FIG. (b), the local device 100 also runs a rendering engine 230.

[0067] Among them, the client 110 is used to send input information to the access gateway 210 and receive the video stream data of the digital human sent by the rendering engine 230. If the business requirement is to display conversation information on the client 110, the client 110 will also receive the conversation information sent by the access gateway 210.

[0068] The access gateway 210 is used to receive the input information sent by the client 110, call the digital human generation service 220 to generate rendering-related information, and then send the rendering-related information to the rendering engine 230.

[0069] The digital human generation service 220 includes: a keyword spotting (KWS) service 221, an automatic speech recognition (ASR) service 222, a large language model (LLM) 223, a text-to-speech (TTS) service 224, and a voice-driven service 225. Information can be directly transmitted between each service, or forwarded through the access gateway 210. In this embodiment, the example of forwarding information through the access gateway 210 is used for illustration. For example, after converting the voice information into text information, the ASR service 222 sends it to the access gateway 210, and the access gateway 210 then forwards it to the LLM 223.

[0070] The rendering engine 230 is used to render rendering-related information to obtain the video stream data of the digital human, and send the video stream data of the digital human to the client 110.

[0071] In some embodiments, the local device 100 is used to send input information to the cloud device 200, and the input information is used for the cloud device 200 to generate rendering-related information associated with the digital human through the digital human generation service 220.

[0072] The input information is a question or request for the digital human input by the user through the local device 100, including but not limited to text information and voice information. For example, the user makes a voice input "Little X, wake up" through the local device 100, or the user makes a text input "What's the weather like today?" through the local device 100.

[0073] Optionally, the local device 100 runs the client 110, and the cloud device 200 runs the access gateway 210 and the digital human generation service 220; the client 110 is used to send input information to the access gateway 210.

[0074] The access gateway 210 is used to perform information interaction with the client 110. After receiving the input information, the access gateway 210 can send the input information to the digital human generation service 220.

[0075] In some embodiments, the cloud device 200 is used to generate rendering-related information according to the input information, and the rendering-related information is used to provide the information required for rendering the digital human.

[0076] The input information includes the user's question or request for the digital human. In order for the digital human to respond to the input information more flexibly, it is necessary to use the rendering-related information to render the digital human to improve the interactivity and flexibility of the digital human.

[0077] Optionally, the access gateway 210 is configured to call the digital human generation service 220 according to the input information to generate rendering-related information; the access gateway 210 is configured to send the rendering-related information to the rendering engine 230.

[0078] The digital human generation service 220 includes at least one service. The access gateway 210 calls the service corresponding to the input information according to the input information, and finally generates rendering-related information.

[0079] By way of example and not limitation, the digital human generation service 220 includes at least one of the following services: KWS service 221, ASR service 222, LLM 223, TTS service 224, voice-driven service 225.

[0080] Among them, the cloud device 200 is used to wake up the digital human through the KWS service 221. For example, if the input information includes "Xiaox wake up", the KWS service 221 will wake up the digital human corresponding to "Xiaox".

[0081] The cloud device 200 is used to receive the input information through the ASR service 222 and convert the voice information in the input information into text information. For example, if the user does not input text information but directly gives a voice instruction "Provide a recipe for breakfast", after the ASR service 222 receives the voice information, it converts it into the corresponding text information.

[0082] In some embodiments, the input information includes at least one of the first input information and the second input information. The first input information is used to indicate the problem content or request content that the digital human is to respond to, and the second input information is used to indicate at least one of the conversation style and virtual image type of the digital human;

[0083] The cloud device 200 is configured to generate rendering-related information according to at least one of the first input information and the second input information.

[0084] Figure 2 Shows a schematic diagram of digital human rendering provided by an exemplary embodiment of the present application.

[0085] The access gateway 210 receives the input information sent by the client 110. The input information includes at least one of the first input information and the second input information. After interacting with the LLM 213, TTS service 214, and voice-driven service 215, the access gateway 210 sends the generated rendering-related information to the rendering engine 230. The rendering-related information includes digital human dynamic information, and the digital human dynamic information is used to indicate the movement information of at least one body part of the digital human.

[0086] The access gateway 210 sends the input information to the LLM 213, and the LLM generates conversation information according to the input information and sends it to the access gateway 210. The conversation information is used to respond to the question content or request content.

[0087] The access gateway 210 sends the conversation information to the TTS service 214, and the TTS service 214 converts the conversation information into voice information and sends it to the access gateway 210.

[0088] The access gateway 210 sends the voice information to the voice drive service 215, and the voice drive service 215 converts the voice information into digital human dynamic information. The digital human dynamic information includes at least one of the digital human's limb movement information, facial expression information, and lip movement information. The embodiments of the present application do not limit this.

[0089] In some embodiments, the input information includes first input information.

[0090] The cloud device is used to generate conversation information according to the first input information through the LLM 213; convert the conversation information into first voice information through the TTS service 214; and convert the first voice information into digital human dynamic information through the voice drive service 215.

[0091] The first input information (User Prompt, also known as user prompt, user input information) is a request or question that the user poses to the digital human in the input information, providing a specific task for the LLM 223 to solve. If the problem scope in the first input information is not clear, the reply generated by the large language model may be more general and lack pertinence.

[0092] For example, the first input information 1 is "How to write a bubble sort algorithm in Python?", and the first input information 2 is "How to write a bubble sort algorithm?". The problem scope of the first input information 1 is clearer than that of the first input information 2. The reply for the first input information 1 includes the steps of writing a bubble sort algorithm in Python, while the reply for the first input information 2 may include the steps of writing a bubble sort algorithm in C++ or JAVA, but may not include how to write a bubble sort algorithm in Python.

[0093] In some embodiments, the input information includes second input information.

[0094] A cloud device is used to generate conversation information through LLM213 according to a preset text and the conversation style of a digital human; convert the conversation information into second voice information corresponding to a first voice tone through TTS service 214 according to the first voice tone, where the first voice tone matches the virtual image type of the digital human; convert the second voice information into digital human dynamic information corresponding to a first virtual model through voice driving service 215 according to the first virtual model, where the first virtual model matches the virtual image type of the digital human.

[0095] The second input information (System Prompt, also known as system prompt, system input information) is used to indicate the conversation style, tone, behavior pattern, virtual image type, etc. of the digital human. The second input information can indicate that the digital human is a "child", a "programming assistant", or an "English teacher". For example, if the user needs to have a conversation with a digital human of the programming assistance type, the second input information can be "You are a programming assistant who helps users answer programming-related questions".

[0096] The second input information can also be input by selecting options. For example, if the user selects the digital human "Little X" on the client 110, then the second input information is used to indicate that the conversation style of the digital human is the conversation style of Little X, and the virtual image type of the digital human is the virtual image of Little X.

[0097] Exemplarily, the preset text is "Today is also a wonderful day". The user selects the digital human "Little X" on the local device. The conversation style of Little X is to add "Little X personally believes" before the conversation. The first virtual model corresponding to Little X is a humanoid model. LLM213 generates conversation information "Little X personally believes that today is also a wonderful day" according to "Today is also a wonderful day" and the conversation style of Little X. The first voice tone that matches the virtual image of Little X is a humanoid voice tone. TTS service 214 converts the conversation information into second voice information corresponding to the humanoid voice tone according to the humanoid voice tone. Voice driving service 215 converts the second voice information into lip movement information corresponding to the humanoid model according to the humanoid model corresponding to Little X.

[0098] In some embodiments, the input information includes first input information and second input information;

[0099] A cloud device is used to generate conversation information through LLM223 according to the first input information and the conversation style of the digital human; convert the conversation information into third voice information corresponding to a second voice tone through TTS service 224 according to the second voice tone, where the second voice tone matches the virtual image type of the digital human.

[0100] For example, if the virtual image type of the digital human is the virtual image of Little X, then the second voice is the simulated human voice that matches the virtual image of Little X. The TTS service 224 converts the conversation information into the third voice information corresponding to the simulated human voice according to the simulated human voice.

[0101] In some embodiments, the cloud device 200 is used to convert the third voice information into the digital human dynamic information corresponding to the second virtual model according to the second virtual model through the voice driving service 225, and the second virtual model matches the virtual image type of the digital human.

[0102] For example, if the virtual image type of the digital human is the virtual image of Little X, then the second virtual model is the simulated human model that matches the virtual image of Little X. The voice driving service 225 converts the third voice information into the lip movement information corresponding to the simulated human model according to the simulated human model.

[0103] Exemplarily, the first input information is "What's the weather like today?", the user selects the digital human "Little X" on the client 110, the conversation style of Little X is to add "Little X personally thinks" before the conversation, and the virtual model corresponding to Little X is the simulated human model. The LLM 223 generates the conversation information "Little X personally thinks it will be a sunny day today" according to "What's the weather like today?" and the conversation style of Little X. The voice that matches the virtual image of Little X is the simulated human voice, and the TTS service 224 converts the conversation information into the third voice information corresponding to the simulated human voice according to the simulated human voice. The voice driving service 225 converts the third voice information into the lip movement information corresponding to the simulated human model according to the simulated human model corresponding to Little X.

[0104] In some embodiments, the cloud device 200 is used to send the first input information and / or the conversation information to the client 110 through the access gateway 210.

[0105] In the case where the client 110 needs to display the first input information and the conversation information, the access gateway 210 sends the first input information and the conversation information to the client 110. In the case where the client 110 does not need to display the first input information and the conversation information, the access gateway 210 does not send the first input information and the conversation information.

[0106] In some embodiments, the rendering engine 230 is used to render the rendering-related information to obtain the video stream data of the digital human, and send the video stream data of the digital human to the client 110.

[0107] Taking the rendering of relevant information including the lip information corresponding to the simulated human model as an example, the rendering engine 230 is used to render the lip information corresponding to the simulated human model to obtain the video stream data of Little X. After the video stream data is sent to the client 110, the client 110 can decode it to obtain a video. The content of the video includes a simulated human saying "Little X personally believes that today will be a sunny day", and the lip movement matches the dialogue information.

[0108] Optionally, the rendering engine 230 sends the video stream data of the digital human and the voice information of the digital human to the client 110.

[0109] In the embodiments of the present application, the rendering engine 230 is usually taken as an example of the Unreal Engine for illustration, and the specific implementation of the rendering engine 230 is not limited.

[0110] In some embodiments, in the cloud rendering mode, the cloud device 200 is used to render the video stream data of the digital human according to the rendering relevant information, and the local device 100 is used to receive the video stream data of the digital human.

[0111] Optionally, in the cloud rendering mode, the rendering engine 230 runs in the cloud device 200. As Figure 1 shown in (a) of, the rendering engine 230 runs in the cloud device 200. The rendering engine 230 renders the video stream data of the digital human according to the rendering relevant information and sends the video stream data of the digital human to the client 110.

[0112] In some embodiments, in the local rendering mode, the cloud device 200 is used to send the rendering relevant information to the local device 100, and the local device 100 is used to render the video stream data of the digital human according to the rendering relevant information.

[0113] Optionally, in the local rendering mode, the rendering engine 230 runs in the local device 100. As Figure 1 shown in (b) of, the rendering engine 230 runs in the local device 100. The rendering engine 230 renders the video stream data of the digital human according to the rendering relevant information and sends the video stream data of the digital human to the client 110.

[0114] In summary, the computer system provided in this embodiment supports two rendering modes, namely the local rendering mode and the cloud rendering mode. In the local rendering mode, the local device performs the rendering, and in the cloud rendering mode, the cloud device performs the rendering, thereby meeting the business requirements in different scenarios and solving the problem that related technologies usually can only meet single business requirements.

[0115] The cloud device provided in this embodiment is used to generate rendering-related information according to the first input information and the second input information, so that the rendering-related information is associated with the first input information and the second input information. Ultimately, the rendered digital human also has a stronger association with the first input information and the second input information, enhancing the interaction experience between the user and the digital human.

[0116] In the computer system provided in this embodiment, at least one of the local device and the cloud device also runs a rendering engine. In the local rendering mode, the rendering engine runs in the local device; in the cloud rendering mode, the rendering engine runs in the cloud device. In different modes, the rendering engine runs in different devices to adapt to the business requirements in different scenarios. For example, in the local rendering mode, the business requirement is mainly low latency, and running the rendering engine in the local device can render the digital human more quickly; in the cloud rendering mode, the business requirement is mainly high quality, and running the rendering engine in the cloud device can provide a digital human with higher precision.

[0117] Figure 3 The flowchart of the computer system for digital human rendering provided by an exemplary embodiment of the present application is shown. The computer system includes a local device 100 and a cloud device 200, and the computer system supports the local rendering mode and the cloud rendering mode.

[0118] In some embodiments, the local device 100 is used to send input information to the cloud device 200, and the input information is used for the cloud device 200 to generate rendering-related information associated with the digital human through the digital human generation service.

[0119] The local device 100 can be a computer device such as a mobile phone, a tablet computer, a vehicle-mounted terminal (in-vehicle computer), a wearable device, a personal computer, etc. The cloud device 200 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides services such as cloud database, cloud computing, cloud function, cloud storage, network service, cloud communication, middleware service, domain name service, security service, content delivery network, and big data and artificial palm image recognition platform.

[0120] The local device 100 and the cloud device 200 can communicate through a network, such as a wired network or a wireless network.

[0121] A digital human refers to a virtual human image created by using digital technology and similar to the human image. The digital human can be represented as a two-dimensional image or a three-dimensional model and can have functions such as speech recognition and natural language processing. In the embodiments of the present application, the digital human is described by taking the three-dimensional model as an example.

[0122] The input information is a question or request for the digital human input by the user through the local device 100, including but not limited to text information and voice information. For example, the user makes a voice input of "Little X, wake up" through the local device 100, or the user makes a text input of "What's the weather like today?" through the local device 100.

[0123] In some embodiments, the cloud device 200 is used to generate rendering-related information according to the input information, and the rendering-related information is used to provide the information required for rendering the digital human.

[0124] The rendering-related information includes digital human dynamic information, and the digital human dynamic information is used to indicate the motion information of at least one body part of the digital human. The digital human dynamic information includes at least one of the limb movement information, facial expression information, and lip movement information of the digital human. The embodiments of the present application do not limit this.

[0125] The input information includes a question or request from the user for the digital human. In order for the digital human to respond to the input information more flexibly, it is necessary to use the rendering-related information to render the digital human, improving the interactivity and flexibility of the digital human.

[0126] In some embodiments, in the local rendering mode, the cloud device 200 is used to send the rendering-related information to the local device 100, and the local device 100 is used to render the video stream data of the digital human according to the rendering-related information.

[0127] As Figure 1 shown in (b) of, the cloud device 200 sends the rendering-related information to the local device 100, and the local device 100 renders the video stream data of the digital human according to the rendering-related information. Since the rendering operation is performed in the local device 100, the rendering load of the cloud device 200 is reduced, and the response speed of the digital human is improved.

[0128] In some embodiments, in the cloud rendering mode, the cloud device 200 is used to render the video stream data of the digital human according to the rendering-related information, and the local device 100 is used to receive the video stream data of the digital human.

[0129] As Figure 1 shown in (a) of, the cloud device 200 renders the video stream data of the digital human according to the rendering-related information, and sends the video stream data of the digital human to the local device 100. Since the rendering operation is performed in the cloud device 200, the graphics card used by the cloud device 200 to render the digital human is better than the graphics card used by the local device 100, so the rendering quality of the digital human is higher.

[0130] Optionally, the cloud device 200 adopts a distributed architecture. By adopting a distributed architecture, in the cloud rendering mode, the rendering tasks related to the digital human are dispersed to different servers, improving the rendering efficiency and enabling more flexible management of the rendering tasks.

[0131] In some embodiments, by adopting cloud rendering technology, the computer system supports both the local rendering mode and the cloud rendering mode, improving the remote access performance of the computer system.

[0132] Build a cloud rendering cluster based on a distributed architecture and disperse the rendering tasks of the digital human to different nodes (cloud devices). By adopting load balancing technology, the rendering tasks are dynamically allocated to different nodes to ensure efficient and stable rendering. Adopting a distributed architecture can not only improve the rendering efficiency but also ensure the redundancy of task processing. Even if a certain node is overloaded or fails, the computer system can quickly switch to other nodes to avoid task jams and maintain a smooth operation experience.

[0133] In summary, the computer system provided in this embodiment supports two rendering modes, namely the local rendering mode and the cloud rendering mode. In the local rendering mode, the local device performs the rendering, and in the cloud rendering mode, the cloud device performs the rendering, thereby meeting the business requirements in different scenarios and solving the problem that related technologies usually can only meet single business requirements.

[0134] Figure 4 The flowchart of generating rendering-related information provided by an exemplary embodiment of the present application is shown. The computer system includes a local device 100 and a cloud device 200, and the computer system supports the local rendering mode and the cloud rendering mode; the input information includes at least one of the first input information and the second input information. The first input information is used to indicate the problem content or request content that the digital human is to respond to, and the second input information is used to indicate at least one of the conversation style and virtual image type of the digital human.

[0135] In some embodiments, the cloud device 200 is configured to generate rendering-related information according to at least one of the first input information and the second input information.

[0136] The first input information (User Prompt, also known as user prompt, user input information) is a request or question that the user poses to the digital human in the input information, providing the specific task that the cloud device 200 needs to solve. If the problem scope in the first input information is not clear, the reply generated by the large language model may be more general and lack pertinence.

[0137] For example, the first input message 1 is "How to write a bubble sort algorithm in Python?", and the first input message 2 is "How to write a bubble sort algorithm?". The question scope of the first input message 1 is more specific than that of the first input message 2. The reply to the first input message 1 includes the steps of writing a bubble sort algorithm in Python, while the reply to the first input message 2 may include the steps of writing a bubble sort algorithm in C++ or JAVA, but may not include how to write a bubble sort algorithm in Python.

[0138] The second input information (System Prompt, also known as system prompt or system input information) is used to indicate the conversation style, tone, behavior pattern, virtual image type, etc. of the digital human. The second input information can indicate that the digital human is a "child", a "programming assistant", or an "English teacher". For example, if the user needs to have a conversation with a digital human of the programming assistance type, the second input information can be "You are a programming assistant who helps users answer programming-related questions".

[0139] The second input information can also be input by selecting options. For example, if the user selects the digital human "Little X" on the client 110, then the second input information is used to indicate that the conversation style of the digital human is the conversation style of Little X, and the virtual image type of the digital human is the virtual image of Little X.

[0140] The cloud device 200 can generate rendering-related information according to the first input information, can also generate rendering-related information according to the second input information, and can also generate rendering-related information according to the first input information and the second input information. The specific implementation details are as follows.

[0141] (1) Generating rendering-related information according to the first input information:

[0142] Figure 5 The flowchart of generating rendering-related information provided by an exemplary embodiment of the present application is shown.

[0143] In some embodiments, the rendering-related information includes digital human dynamic information, and the digital human dynamic information is used to indicate the motion information of at least one body part of the digital human. The cloud device 200 includes an LLM, a TTS service, and a voice driving service, and the input information includes the first input information.

[0144] The cloud device 200 is used to generate conversation information according to the first input information through the LLM, and the conversation information is used to respond to the question content or request content; convert the conversation information into the first voice information through the TTS service; and convert the first voice information into digital human dynamic information through the voice driving service.

[0145] The dynamic information of the digital human includes at least one of the limb movement information, facial expression information, and lip movement information of the digital human, and the embodiments of the present application do not limit this.

[0146] Exemplarily, the first input information is "What's the weather like today?", and the LLM generates the dialogue information "It will be a sunny day today" according to "What's the weather like today?". This dialogue information is only used to respond to the question content and does not use any dialogue style. The TTS service converts the dialogue information into the first voice information with the default voice. The virtual model corresponding to the digital human is the default virtual model, and the voice driving service converts the first voice information into the lip movement information corresponding to the default virtual model.

[0147] (2) Generate rendering-related information according to the second input information:

[0148] Figure 6 The flowchart of generating rendering-related information provided by an exemplary embodiment of the present application is shown.

[0149] In some embodiments, the rendering-related information includes the digital human dynamic information, which is used to indicate the movement information of at least one body part of the digital human. The cloud device 200 includes an LLM, a TTS service, and a voice driving service, and the input information includes the second input information;

[0150] The cloud device 200 is used to generate dialogue information through the LLM according to the preset text and the dialogue style of the digital human, and the dialogue information is used to respond to the question content or request content;

[0151] Through the TTS service, convert the dialogue information into the second voice information corresponding to the first voice according to the first voice, and the first voice matches the virtual image type of the digital human;

[0152] Through the voice driving service, convert the second voice information into the digital human dynamic information corresponding to the first virtual model according to the first virtual model, and the first virtual model matches the virtual image type of the digital human.

[0153] Exemplarily, the preset text is "Today is also a wonderful day", the user selects the digital human "Little X" on the local device 100, the dialogue style of Little X is to add "Little X personally thinks" before the dialogue, and the virtual model corresponding to Little X is a realistic human model. The LLM generates the dialogue information "Little X personally thinks, today is also a wonderful day" according to "Today is also a wonderful day" and the dialogue style of Little X. The voice that matches the virtual image of Little X is the realistic human voice, and the TTS service converts the dialogue information into the second voice information corresponding to the realistic human voice according to the realistic human voice. The voice driving service converts the second voice information into the lip movement information corresponding to the realistic human model according to the realistic human model corresponding to Little X.

[0154] By way of example and not limitation, the types of virtual avatars include: children, working women, domineering CEOs, small animals, realistic humans, etc. The embodiments of the present application do not limit the specific types of virtual avatars.

[0155] For example, when the role of the digital human is a child, the generated dialogue information includes at least one of the following characteristics: simple sentence structure, few words, direct emotional expression, single view, simple narrative method.

[0156] Among them, the simple sentence structure is because children usually use relatively simple sentence structures and have not fully mastered complex clauses and long sentences; few words are because children have relatively fewer words than adults and only use some common words when expressing; direct emotional expression is because children tend to express emotions directly and usually do not have too many euphemistic or complex expressions; single view is because children's thinking has not been fully developed and they will not think from multiple perspectives to obtain a comprehensive view; simple narrative method is because children lack complex plots and details when narrating and usually focus on directly narrating basic facts.

[0157] Exemplarily, when the role of the digital human is a child, the generated dialogue information is as follows:

[0158] Question 1: "Do you have any particularly favorite toys?"

[0159] Reply 1: "Hi, I'm Musa! I really like toys! My favorite toys are building blocks. They allow me to freely build various shapes and structures, which is very interesting! I also like playing video games, especially those puzzle games that require thinking, which can challenge my thinking ability. In addition, I like watching cartoons and reading storybooks. They can let me feel different worlds and wonderful imaginations. In short, my toys are very diverse, and each one can make me happy and excited!"

[0160] Question 2: "What was your dream when you were a child?"

[0161] Reply 2: "My dream when I was a child was to become a scientist and explore the mysteries of the universe. I wanted to know if there is other life in the universe, if there are other stars and planets, and what the most mysterious black hole in the universe looks like. I hope that one day I can take a spaceship and explore every corner of the universe to discover more unknown worlds."

[0162] (3) Generate rendering-related information according to the first input information and the second input information:

[0163] Figure 7 The flowchart of generating rendering-related information provided by an exemplary embodiment of the present application is shown.

[0164] In some embodiments, the rendering-related information includes digital human dynamic information, which is used to indicate the movement information of at least one body part of the digital human. The cloud device 200 includes an LLM, a TTS service, and a voice driving service. The input information includes first input information and second input information;

[0165] The cloud device 200 is used to generate dialogue information through the LLM according to the first input information and the dialogue style of the digital human. The dialogue information is used to respond to question content or request content;

[0166] Through the TTS service, according to the second voice color, the dialogue information is converted into third voice information corresponding to the second voice color, and the second voice color matches the virtual image type of the digital human;

[0167] Through the voice driving service, according to the second virtual model, the third voice information is converted into digital human dynamic information corresponding to the second virtual model, and the second virtual model matches the virtual image type of the digital human.

[0168] Suppose the virtual image type of the digital human is the virtual image of Little X. Then the second voice color is the simulated human voice color matching the virtual image of Little X, and the second virtual model is the simulated human model matching the virtual image of Little X. The TTS service converts the dialogue information into third voice information corresponding to the simulated human voice color according to the simulated human voice color; the voice driving service converts the third voice information into digital human dynamic information corresponding to the simulated human model according to the simulated human model.

[0169] Exemplarily, the first input information is "What's the weather like today?", the user selects the digital human "Little X", the dialogue style of Little X is to add "Little X personally thinks" before the dialogue, and the virtual model corresponding to Little X is a simulated human model. The LLM generates dialogue information "Little X personally thinks it will be a sunny day today" according to "What's the weather like today?" and the dialogue style of Little X. The voice color matching the virtual image of Little X is the simulated human voice color, and the TTS service converts the dialogue information into third voice information corresponding to the simulated human voice color according to the simulated human voice color. The voice driving service converts the third voice information into lip movement information corresponding to the simulated human model according to the simulated human model corresponding to Little X.

[0170] In some embodiments, the input information further includes dialogue context information, which is used to indicate the historical content of the digital human's dialogue;

[0171] The cloud device is further used to use the dialogue context information as reference information when generating dialogue information through the LLM. The dialogue information is used to respond to question content or request content.

[0172] Dialogue context information (context) can help the LLM maintain the coherence of the dialogue. Through the dialogue context information, the LLM can obtain the historical content of the digital human dialogue, which can be used as reference information to generate dialogue information with stronger relevance to the historical content.

[0173] For example, if it is known from the dialogue context information that sorting algorithms have been discussed in the historical dialogue, and the subsequent input information includes "quick sort", the LLM will generate dialogue information related to sorting algorithms.

[0174] Another example is that if programming-related issues have been discussed in the historical dialogue, then when the LLM receives input information subsequently, it will preferentially use information in the programming field as reference information to generate dialogue information, thereby improving the relevance of the dialogue information to the historical content.

[0175] In some embodiments, the digital human generation service includes at least one of the KWS (voice wake-up) service and the ASR (automatic speech recognition) service;

[0176] The cloud device 200 is used to wake up the digital human through the KWS service;

[0177] The cloud device 200 is used to receive input information through the ASR service and convert the voice information in the input information into text information.

[0178] For example, if the input information includes "Xiaox X, wake up", the KWS service will wake up the digital human corresponding to "Xiaox X"; if the input information provided by the user is not text information, but voice information input through voice instructions: "Provide a recipe for breakfast", after the ASR service receives this voice information, it will convert it into the corresponding text information.

[0179] Figure 8 Shows a schematic diagram of the digital human generation service provided by an exemplary embodiment of the present application.

[0180] The digital human generation service 220 includes: the KWS service 221, the ASR service 222, the LLM 223, the TTS service 224, and the voice drive service 225.

[0181] Each service in the digital human generation service 220 can directly transmit information, or can forward information through an access gateway to transmit information.

[0182] For example, in Figure 8 (a), the input information (voice information) includes "Xiaox X, wake up, what's the weather like today", indicating that the selected digital human is Xiaox X. The KWS service 221 will wake up the digital human corresponding to "Xiaox X" and send the input information to the ASR service 222, as Figure 8As shown in (b) of , a virtual model corresponding to the awakened "Xiaox" appears on the interface. Xiaox's conversation style is to add "Xiaox personally believes" before the conversation, and the virtual model corresponding to Xiaox is a human simulation model.

[0183] The ASR service 222 receives the input information, converts the voice information in the input information into text information, and then sends it to the LLM 223.

[0184] If the virtual image type of the digital human is not indicated, the LLM 223 will generate the conversation information "It will be a sunny day today" based on "What's the weather like today". When the selected digital human is Xiaox, the LLM 223 generates the conversation information "Xiaox personally believes that it will be a sunny day today" according to "What's the weather like today" and Xiaox's conversation style, and sends it to the TTS service 224. As Figure 8 As shown in (c) of , the dashed box indicates that the content that Xiaox is about to answer is "Xiaox personally believes that it will be a sunny day today".

[0185] The voice color matching Xiaox is a human simulation voice color. The TTS service 224 converts the conversation information into the third voice information corresponding to the human simulation voice color according to the human simulation voice color, and sends it to the voice driving service 225.

[0186] The voice driving service 225 converts the third voice information into lip movement information corresponding to the human simulation model according to the human simulation model corresponding to Xiaox. As Figure 8 As shown in (d) of , when the human simulation model corresponding to Xiaox plays the third voice information, the lip shape will change dynamically, and the change in the lip shape matches the played third voice information.

[0187] In some embodiments, the voice driving service 225 can also convert the third voice information into limb movement information, facial expression information, etc. corresponding to the human simulation model. For example, it can make the facial expression of the human simulation model corresponding to Xiaox become a "smiling face" when playing the third voice information. The embodiments of the present application do not limit this, and only the example of changing the lip shape is used for illustration.

[0188] In summary, the cloud device provided in this embodiment is used to generate rendering-related information according to at least one of the first input information and the second input information, so that the rendering-related information is associated with the input information, and finally the rendered digital human also has a stronger association with the input information, improving the interaction experience between the user and the digital human.

[0189] The cloud device provided in this embodiment enables the digital human to respond to the user's voice commands in real time through the text-to-speech service and speech recognition service; by using the LLM as the virtual brain of the digital human, the responses of the digital human are made more natural and effective; and through the voice-driven service, the facial expressions and body movements of the digital human are rendered in real time, enhancing the interaction experience between the user and the digital human.

[0190] Figure 9 FIG. 4 shows a flowchart of a computer system for digital human rendering provided by an exemplary embodiment of the present application. The computer system includes a local device 100 and a cloud device 200, and the computer system supports a local rendering mode and a cloud rendering mode; the local device 100 runs a client 110, the cloud device 200 also runs an access gateway 210, and at least one of the local device 100 and the cloud device 200 also runs a rendering engine 230;

[0191] In some embodiments, the client 110 is configured to send input information to the access gateway 210.

[0192] The access gateway 210 is used to perform information interaction with the client 110. After receiving the input information, the access gateway 210 may send the input information to the digital human generation service 220.

[0193] In some embodiments, the access gateway 210 is configured to call the digital human generation service 220 according to the input information to generate rendering-related information;

[0194] The digital human generation service 220 includes at least one service. The access gateway 210 calls the service corresponding to the input information according to the input information and finally generates rendering-related information.

[0195] In some embodiments, the access gateway 210 is configured to send the rendering-related information to the rendering engine 230; the rendering engine 230 is configured to render the rendering-related information to obtain the video stream data of the digital human and send the video stream data of the digital human to the client 110;

[0196] Among them, in the local rendering mode, the rendering engine 230 runs in the local device 100; in the cloud rendering mode, the rendering engine 230 runs in the cloud device 200.

[0197] Exemplarily, refer to Figure 8Embodiment: Taking the digital human as Little X, the virtual model corresponding to Little X is the simulation human model, the dialogue information is "Little X personally believes that today will be a sunny day", and taking the lip-sync information corresponding to the simulation human model as an example of the rendering-related information, the rendering engine 230 is used to render the lip-sync information corresponding to the simulation human model to obtain the video stream data of Little X. After the video stream data is sent to the client, the client can decode it to obtain a video. The content of the video includes a simulation human saying "Little X personally believes that today will be a sunny day", and the lip-sync matches the dialogue information.

[0198] Optionally, the rendering engine 230 sends the video stream data of the digital human and the voice information of the digital human to the client 110.

[0199] In the embodiments of the present application, it is usually described by taking the rendering engine 230 as the Unreal Engine as an example, and the specific implementation of the rendering engine 230 is not limited.

[0200] In summary, in the computer system provided in this embodiment, at least one of the local device and the cloud device also runs a rendering engine. In the local rendering mode, the rendering engine runs in the local device; in the cloud rendering mode, the rendering engine runs in the cloud device. In different modes, the rendering engine runs in different devices to adapt to the business requirements in different scenarios. For example, in the local rendering mode, the business requirement is mainly low latency, and running the rendering engine in the local device can render the digital human more quickly; in the cloud rendering mode, the business requirement is mainly high quality, and running the rendering engine in the cloud device can provide a digital human with higher precision.

[0201] In some embodiments, when the first switching condition is met, the local device and the cloud device are used to switch from the cloud rendering mode to the local rendering mode; when the second switching condition is met, the local device and the cloud device are used to switch from the local rendering mode to the cloud rendering mode.

[0202] The local device and the cloud device flexibly select to switch to the local rendering mode or the cloud rendering mode according to different scenarios, so as to meet the business requirements and improve the flexibility of the computer system to render digital humans.

[0203] Figure 10 Shows a schematic diagram of switching from the cloud rendering mode to the local rendering mode provided by an exemplary embodiment of the present application.

[0204] The scenarios where the local device 100 and the cloud device 200 switch from the cloud rendering mode to the local rendering mode include at least one of Scenario 1 and Scenario 2.

[0205] Scenario 1: When the first switching condition is met, the local device is used to send a first switching request to the cloud device, and the cloud device is used to send a first switching response to the local device. The first switching response is used to indicate that the cloud device agrees to the local device to switch from the cloud rendering mode to the local rendering mode.

[0206] In some embodiments, after receiving the first switching request, the cloud device switches from the cloud rendering mode to the local rendering mode; after receiving the first switching response, the local device switches from the cloud rendering mode to the local rendering mode.

[0207] Scenario 2: When the first switching condition is met, the cloud device is used to send a first switching instruction to the local device, and the first switching instruction is used to instruct the local device to switch from the cloud rendering mode to the local rendering mode.

[0208] In some embodiments, after the cloud device switches from the cloud rendering mode to the local rendering mode, it sends a first switching instruction to the local device; after receiving the first switching instruction, the local device switches from the cloud rendering mode to the local rendering mode.

[0209] In some embodiments, the first switching condition includes at least one of the following:

[0210] · The current idle computing resources of the local device are greater than (or greater than or equal to) the first idle threshold;

[0211] · The current idle bandwidth of the local device is greater than (or greater than or equal to) the second idle threshold;

[0212] · The data privacy level of the current service is higher than (or higher than or equal to) the first level, and the data privacy level is used to indicate the confidentiality requirements of the service;

[0213] · The latency requirement level of the current service is higher than (or higher than or equal to) the second level, and the latency requirement level is used to indicate the timeliness requirements of the service.

[0214] Wherein, the first idle threshold and the second idle threshold are preset values, and the first level and the second level are preset levels.

[0215] Optionally, the computer system in this application uses a full-function GPU, and the full-function GPU is used to implement artificial intelligence services, including large language models (LLMs), automatic speech recognition (ASR) services, text-to-speech (TTS) services, etc.

[0216] In some embodiments, the current idle computing resources are the current idle video memory size or the current idle video memory percentage of the full-function GPU.

[0217] In some embodiments, the encoding and decoding of the video stream data of the digital human are realized through an encoder and a decoder supported by the full-function GPU.

[0218] Rendering using a full - function GPU can improve the quality of the digital human appearance and the scene performance level.

[0219] Exemplarily, if the video memory size of the full - function GPU is 24 gigabytes (GB), the first idle threshold is 10 GB, and the currently used video memory size is 4 GB, then the currently idle video memory size is 24 - 4 = 20 GB, which is greater than the first idle threshold, meeting the first switching condition.

[0220] In some embodiments, the current idle bandwidth is the part of the total bandwidth allowed to be used by the local device that is not occupied by data transfer tasks.

[0221] In some embodiments, the higher the confidentiality requirement of the service, the higher the data privacy level. For example, the confidentiality requirement of Service A is extremely high, the data privacy level is 5, and the first level is 3. Therefore, when the current service is Service A, the first switching condition is met.

[0222] In some embodiments, the higher the timeliness requirement of the service, the higher the latency requirement level. For example, Service B has no requirement for timeliness, the latency requirement level is 1, and the second level is 3. Therefore, when the current service is Service B, the first switching condition is not met.

[0223] Figure 11 Shows a schematic diagram of switching from the local rendering mode to the cloud rendering mode provided by an exemplary embodiment of the present application.

[0224] The scenarios where the local device 100 and the cloud device 200 switch from the local rendering mode to the cloud rendering mode include at least one of Scenario 3 and Scenario 4.

[0225] Scenario 3: Under the condition of meeting the second switching condition, the local device is used to send a second switching request to the cloud device, and the cloud device is used to send a second switching response to the local device. The second switching response is used to indicate that the cloud device agrees that the local device switches from the local rendering mode to the cloud rendering mode.

[0226] In some embodiments, after receiving the second switching request, the cloud device switches from the local rendering mode to the cloud rendering mode; after receiving the second switching response, the local device switches from the local rendering mode to the cloud rendering mode.

[0227] Scenario 4: Under the condition of meeting the second switching condition, the cloud device is used to send a second switching instruction to the local device. The second switching instruction is used to indicate that the local device switches from the local rendering mode to the cloud rendering mode.

[0228] In some embodiments, after the cloud device switches from the local rendering mode to the cloud rendering mode, it sends a second switching instruction to the local device; after receiving the second switching instruction, the local device switches from the local rendering mode to the cloud rendering mode.

[0229] In some embodiments, the second switching condition includes at least one of the following:

[0230] · The current idle computing resources of the local device are less than the first idle threshold;

[0231] · The current idle bandwidth of the local device is less than the second idle threshold;

[0232] · The data privacy level of the current service is lower than the first level, and the data privacy level is used to indicate the confidentiality requirements of the service;

[0233] · The latency requirement level of the current service is lower than the second level, and the latency requirement level is used to indicate the timeliness requirements of the service.

[0234] For specific implementation details, refer to the first switching condition section and will not be elaborated here.

[0235] In summary, for the computer system provided in this embodiment, when the first switching condition is met, the local device and the cloud device are used to switch from the cloud rendering mode to the local rendering mode; when the second switching condition is met, the local device and the cloud device are used to switch from the local rendering mode to the cloud rendering mode, enabling the local device and the cloud device to switch the rendering mode according to different usage scenarios when the switching conditions are met, thereby meeting the service requirements.

[0236] Figure 12 Fig. shows a schematic diagram of a local device for digital human rendering provided by an exemplary embodiment of the present application. The local device supports the local rendering mode and the cloud rendering mode. The local device includes a client module 1210, or includes a client module 1210 and a rendering engine module 1220;

[0237] The client module 1210 is used to send input information to the cloud device, and the input information is used for the cloud device to generate rendering-related information associated with the digital human through the digital human generation service module;

[0238] In the local rendering mode, the client module 1210 is used to receive the rendering-related information sent by the cloud device, and the rendering-related information is used to provide the information required for rendering the digital human;

[0239] The rendering engine module 1220 is used to render the video stream data of the digital human according to the rendering-related information, and send the video stream data of the digital human to the client module 1210;

[0240] In the cloud rendering mode, the client module 1210 is used to receive the video stream data of the digital human, and the video stream data of the digital human is rendered by the cloud device according to the rendering-related information.

[0241] For specific implementation details, refer to Figure 3 the embodiments, which will not be elaborated here.

[0242] In some embodiments, the cloud device further includes an access gateway module, and at least one of the local device and the cloud device also runs a rendering engine module 1220;

[0243] The client module 1210 is used to send input information to the access gateway module;

[0244] The client module 1210 is used to receive the video stream data of the digital human sent by the rendering engine module 1220, and the video stream data of the digital human is rendered by the rendering engine module 1220 based on the rendering-related information;

[0245] Among them, in the local rendering mode, the rendering engine module 1220 runs in the local device; in the cloud rendering mode, the rendering engine module 1220 runs in the cloud device.

[0246] For specific implementation details, refer to Figure 9 the embodiments, which will not be elaborated here.

[0247] In some embodiments, the input information includes at least one of first input information, second input information, and dialogue context information;

[0248] Among them, the first input information is used to indicate the question content or request content that the digital human is to respond to, and the second input information is used to indicate at least one of the dialogue style and virtual image type of the digital human; the dialogue context information is used to indicate the historical content of the digital human's dialogue.

[0249] For specific implementation details, refer to Figure 4 the embodiments, which will not be elaborated here.

[0250] In some embodiments, when the first switching condition is met, the client module 1210 is used to send a first switching request to the cloud device and receive a first switching response sent by the cloud device, and the first switching response is used to indicate that the cloud device agrees to switch the local device from the cloud rendering mode to the local rendering mode.

[0251] In some embodiments, when the first switching condition is met, the client module 1210 is used to receive a first switching instruction sent by the cloud device, and the first switching instruction is used to indicate that the local device switches from the cloud rendering mode to the local rendering mode.

[0252] In some embodiments, the first switching condition includes at least one of the following:

[0253] · The current idle computing resources of the local device are greater than the first idle threshold;

[0254] · The current idle bandwidth of the local device is greater than the second idle threshold;

[0255] · The data privacy level of the current service is higher than the first level, and the data privacy level is used to indicate the confidentiality requirements of the service;

[0256] · The latency requirement level of the current service is higher than the second level, and the latency requirement level is used to indicate the timeliness requirements of the service.

[0257] In some embodiments, when the second switching condition is met, the client module 1210 is configured to send a second switching request to the cloud device and receive a second switching response sent by the cloud device, and the second switching response is used to indicate that the cloud device agrees that the local device switches from the local rendering mode to the cloud rendering mode.

[0258] In some embodiments, when the second switching condition is met, the client module 1210 is configured to receive a second switching indication sent by the cloud device, and the second switching indication is used to indicate that the local device switches from the local rendering mode to the cloud rendering mode.

[0259] In some embodiments, the second switching condition includes at least one of the following:

[0260] · The current idle computing resources of the local device are greater than the first idle threshold;

[0261] · The current idle bandwidth of the local device is less than the second idle threshold;

[0262] · The data privacy level of the current service is lower than the first level, and the data privacy level is used to indicate the confidentiality requirements of the service;

[0263] · The latency requirement level of the current service is lower than the second level, and the latency requirement level is used to indicate the timeliness requirements of the service.

[0264] For specific implementation details, refer to Figure 10 Embodiment and Figure 11 Embodiment, which will not be elaborated here.

[0265] In summary, the local device provided in this embodiment supports two rendering modes, namely the local rendering mode and the cloud rendering mode. In the local rendering mode, the rendering is performed by the rendering engine module on the local device, and in the cloud rendering mode, the rendering is performed by the cloud device, thereby meeting the service requirements in different scenarios and solving the problem that related technologies usually can only meet the requirements of a single service.

[0266] At least one of the local device and the cloud device provided in this embodiment further runs a rendering engine module. In the local rendering mode, the rendering engine module runs in the local device; in the cloud rendering mode, the rendering engine module runs in the cloud device. In different modes, the rendering engine module runs in different devices, so as to adapt to the business requirements in different scenarios. For example, in the local rendering mode, the business requirement is mainly low latency, and running the rendering engine module in the local device can render the digital human more quickly; in the cloud rendering mode, the business requirement is mainly high quality, and running the rendering engine module in the cloud device can provide a digital human with higher precision.

[0267] The local device provided in this embodiment is used to switch from the cloud rendering mode to the local rendering mode when the first switching condition is met; or, when the second switching condition is met, the local device is used to switch from the local rendering mode to the cloud rendering mode, so that the local device can switch the rendering mode according to different usage scenarios when the switching conditions are met, thereby meeting the business requirements.

[0268] Figure 13 The figure shows a schematic diagram of a cloud device for digital human rendering provided by an exemplary embodiment of the present application. The cloud device supports the local rendering mode and the cloud rendering mode. The cloud device includes an access gateway module 1310 and a digital human generation service module 1320, or includes an access gateway module 1310, a digital human generation service module 1320, and a rendering engine module 1301;

[0269] The digital human generation service module 1320 is configured to receive input information sent by the local device, and the input information is used for the cloud device to generate rendering-related information associated with the digital human through the digital human generation service module 1320;

[0270] The digital human generation service module 1320 is configured to generate rendering-related information according to the input information, and the rendering-related information is used to provide the information required for rendering the digital human;

[0271] In the local rendering mode, the access gateway module 1310 is configured to send the rendering-related information to the local device, and the rendering-related information is used for the local device to render the video stream data of the digital human;

[0272] In the cloud rendering mode, the rendering engine module 1301 is configured to render the video stream data of the digital human according to the rendering-related information, and send the video stream data of the digital human to the local device.

[0273] For specific implementation details, refer to Figure 3 the embodiment, which will not be elaborated here.

[0274] In some embodiments, the input information includes at least one of the first input information and the second input information. The first input information is used to indicate the problem content or request content that the digital human is to respond to, and the second input information is used to indicate at least one of the conversation style and virtual image type of the digital human;

[0275] The digital human generation service module 1320 is configured to generate rendering-related information based on at least one of the first input information and the second input information.

[0276] For specific implementation details, refer to Figure 4 the embodiments, which will not be elaborated here.

[0277] In some embodiments, the rendering-related information includes digital human dynamic information, which is used to indicate the motion information of at least one body part of the digital human. The cloud device includes LLM1323, TTS service module 1324, and voice driving service module 1325. The input information includes the first input information;

[0278] LLM1323 is configured to generate conversation information based on the first input information, and the conversation information is used to respond to the problem content or request content;

[0279] The TTS service module 1324 is configured to convert the conversation information into the first voice information;

[0280] The voice driving service module 1325 is configured to convert the first voice information into digital human dynamic information.

[0281] For specific implementation details, refer to Figure 5 the embodiments, which will not be elaborated here.

[0282] In some embodiments, the rendering-related information includes digital human dynamic information, which is used to indicate the motion information of at least one body part of the digital human. The cloud device includes LLM1323, TTS service module 1324, and voice driving service module 1325. The input information includes the second input information;

[0283] LLM1323 is configured to generate conversation information based on the preset text and the conversation style of the digital human, and the conversation information is used to respond to the problem content or request content;

[0284] The TTS service module 1324 is configured to convert the conversation information into the second voice information corresponding to the first voice color according to the first voice color, and the first voice color matches the virtual image type of the digital human;

[0285] The voice driving service module 1325 is configured to convert the second voice information into digital human dynamic information corresponding to the first virtual model according to the first virtual model, and the first virtual model matches the virtual image type of the digital human.

[0286] For specific implementation details, refer to Figure 6 the embodiments, which will not be elaborated here.

[0287] In some embodiments, the rendering-related information includes digital human dynamic information, which is used to indicate the movement information of at least one body part of the digital human. The cloud device includes LLM1323, TTS service module 1324, and voice driving service module 1325, and the input information includes first input information and second input information;

[0288] LLM1323 is used to generate dialogue information according to the first input information and the dialogue style of the digital human, and the dialogue information is used to respond to the question content or request content;

[0289] The TTS service module 1324 is used to convert the dialogue information into third voice information corresponding to the second voice according to the second voice, and the second voice matches the virtual image type of the digital human;

[0290] The voice driving service module 1325 is used to convert the third voice information into digital human dynamic information corresponding to the second virtual model according to the second virtual model, and the second virtual model matches the virtual image type of the digital human.

[0291] For specific implementation details, refer to Figure 7 the embodiments, which will not be elaborated here.

[0292] In some embodiments, the input information further includes dialogue context information, which is used to indicate the historical content of the digital human's dialogue;

[0293] LLM1323 is used to use the dialogue context information as reference information when generating dialogue information, and the dialogue information is used to respond to the question content or request content.

[0294] In some embodiments, the digital human generation service module 1320 includes at least one of the KWS service module 1321 and the ASR service module 1322;

[0295] The KWS service module 1321 is used to wake up the digital human;

[0296] The ASR service module 1322 is used to receive the input information and convert the voice information in the input information into text information.

[0297] For specific implementation details, refer to Figure 4 the embodiments, which will not be elaborated here.

[0298] In some embodiments, the local device includes a client module, and at least one of the local device and the cloud device further runs a rendering engine module 1301;

[0299] The access gateway module 1310 is used to receive the input information sent by the client module;

[0300] The access gateway module 1310 is used to call the digital human generation service module 1320 according to the input information to generate rendering-related information;

[0301] The access gateway module 1310 is used to send the rendering-related information to the rendering engine module 1301;

[0302] The rendering engine module 1301 is used to render the rendering-related information to obtain the video stream data of the digital human and send the video stream data of the digital human to the client module;

[0303] Among them, in the local rendering mode, the rendering engine module 1301 runs on the local device; in the cloud rendering mode, the rendering engine module 1301 runs on the cloud device.

[0304] For specific implementation details, refer to Figure 9 the embodiments, which will not be elaborated here.

[0305] In some embodiments, when the first switching condition is met, the cloud device is used to receive the first switching request sent by the local device; switch from the cloud rendering mode to the local rendering mode; and send a first switching response to the local device, where the first switching response is used to indicate that the cloud device agrees that the local device switches from the cloud rendering mode to the local rendering mode.

[0306] In some embodiments, when the first switching condition is met, the cloud device is used to send a first switching instruction to the local device and switch from the cloud rendering mode to the local rendering mode, where the first switching instruction is used to indicate that the local device switches from the cloud rendering mode to the local rendering mode.

[0307] In some embodiments, the first switching condition includes at least one of the following:

[0308] · The current idle computing resources of the local device are greater than the first idle threshold;

[0309] · The current idle bandwidth of the local device is greater than the second idle threshold;

[0310] · The data privacy level of the current service is higher than the first level, and the data privacy level is used to indicate the confidentiality requirements of the service;

[0311] · The latency requirement level of the current service is higher than the second level, and the latency requirement level is used to indicate the timeliness requirements of the service.

[0312] In some embodiments, when the second switching condition is met, the cloud device is configured to receive a second switching request sent by the local device; switch from the local rendering mode to the cloud rendering mode; and send a second switching response to the local device, where the second switching response is used to indicate that the cloud device agrees that the local device switches from the local rendering mode to the cloud rendering mode.

[0313] In some embodiments, when the second switching condition is met, the cloud device is configured to send a second switching indication to the local device and switch from the local rendering mode to the cloud rendering mode, where the second switching indication is used to indicate that the local device switches from the local rendering mode to the cloud rendering mode.

[0314] In some embodiments, the second switching condition includes at least one of the following:

[0315] · The current idle computing resources of the local device are less than the first idle threshold;

[0316] · The current idle bandwidth of the local device is less than the second idle threshold;

[0317] · The data privacy level of the current service is lower than the first level, and the data privacy level is used to indicate the confidentiality requirements of the service;

[0318] · The latency requirement level of the current service is lower than the second level, and the latency requirement level is used to indicate the timeliness requirements of the service.

[0319] For specific implementation details, refer to Figure 10 Embodiment and Figure 11 Embodiment, which will not be elaborated here.

[0320] In summary, the cloud device provided in this embodiment supports two rendering modes, namely the local rendering mode and the cloud rendering mode. In the local rendering mode, the local device performs rendering, and in the cloud rendering mode, the cloud device performs rendering, thereby meeting the service requirements in different scenarios and solving the problem that related technologies usually can only meet the requirements of a single service.

[0321] At least one of the local device and the cloud device provided in this embodiment also runs a rendering engine module. In the local rendering mode, the rendering engine module runs in the local device; in the cloud rendering mode, the rendering engine module runs in the cloud device. In different modes, the rendering engine module runs in different devices, so as to adapt to the service requirements in different scenarios. For example, in the local rendering mode, the service requirements are mainly for low latency, and running the rendering engine module in the local device can render the digital human more quickly; in the cloud rendering mode, the service requirements are mainly for high quality, and running the rendering engine module in the cloud device can provide a digital human with higher precision.

[0322] The cloud device provided in this embodiment is used to generate rendering-related information based on at least one of the first input information and the second input information, so that the rendering-related information is associated with the input information, and the finally rendered digital human also has a stronger association with the input information, improving the interaction experience between the user and the digital human.

[0323] The cloud device provided in this embodiment enables the local device to switch from the cloud rendering mode to the local rendering mode when the first switching condition is met; or, when the second switching condition is met, the local device is used to switch from the local rendering mode to the cloud rendering mode, so that the local device can switch the rendering mode according to different usage scenarios when the switching conditions are met, thereby meeting the business requirements.

[0324] Figure 14 The flowchart of the digital human rendering method provided by an exemplary embodiment of the present application is shown. This method is executed by a local device that supports the local rendering mode and the cloud rendering mode, and the method includes at least one of the following steps.

[0325] Step 1410: Send input information to the cloud device.

[0326] The input information is used for the cloud device to generate rendering-related information associated with the digital human through the digital human generation service.

[0327] A digital human refers to a virtual human image created by using digital technology that is close to the human image. The digital human appears as a two-dimensional image or a three-dimensional model and can have functions such as speech recognition and natural language processing. In this embodiment of the present application, the digital human is described by taking the three-dimensional model as an example.

[0328] The input information is a question or request for the digital human input by the user through the local device, including but not limited to text information and voice information. For example, the user makes a voice input "Little X, wake up" through the local device, or the user makes a text input "What's the weather like today?" through the local device 100.

[0329] Step 1420: Receive the rendering-related information sent by the cloud device in the local rendering mode.

[0330] The rendering-related information is used to provide the information required for rendering the digital human.

[0331] The rendering-related information includes digital human dynamic information, and the digital human dynamic information is used to indicate the motion information of at least one body part of the digital human. The digital human dynamic information includes at least one of the limb movement information, facial expression information, and mouth shape information of the digital human. This embodiment of the present application does not limit this.

[0332] The input information includes the user's question or request for the digital human. In order for the digital human to respond to the input information more flexibly, it is necessary to use the rendering-related information to render the digital human, improving the interactivity and flexibility of the digital human.

[0333] Step 1430: Render the video stream data of the digital human according to the rendering-related information.

[0334] As Figure 1 shown in (b) of, the local device 100 renders the video stream data of the digital human according to the rendering-related information. Since the rendering operation is executed in the local device 100, the rendering load of the cloud device 200 is reduced, and the response speed of the digital human is improved.

[0335] Step 1421: Receive the video stream data of the digital human in the cloud rendering mode.

[0336] Among them, the video stream data of the digital human is rendered by the cloud device according to the rendering-related information.

[0337] As Figure 1 shown in (a) of, the cloud device 200 renders the video stream data of the digital human according to the rendering-related information, and the local device 100 receives the video stream data of the digital human sent by the cloud device 200. Since the rendering operation is executed in the cloud device 200, the graphics card used by the cloud device 200 to render the digital human is better than the graphics card used by the local device 100, so the rendering quality of the digital human is higher.

[0338] In summary, the method provided in this embodiment is executed by the local device, which supports the local rendering mode and the cloud rendering mode. In the local rendering mode, the local device performs the rendering, and in the cloud rendering mode, the cloud device performs the rendering, so as to meet the business requirements in different scenarios and solve the problem that related technologies usually can only meet a single business requirement.

[0339] Figure 15 The flowchart shows the digital human rendering method provided by an exemplary embodiment of the present application. This method is executed by the local device, which supports the local rendering mode and the cloud rendering mode. This method includes at least one of the following steps.

[0340] Step 1410: Send the input information to the cloud device.

[0341] In some embodiments, the input information includes at least one of the first input information, the second input information, and the dialogue context information;

[0342] Among them, the first input information is used to indicate the question content or request content that the digital human is to respond to, and the second input information is used to indicate at least one of the conversation style and virtual image type of the digital human; the conversation context information is used to indicate the historical content of the digital human's conversation.

[0343] For specific implementation details, refer to Figure 4 the embodiments, which will not be elaborated here.

[0344] Step 1422: In the local rendering mode, receive the rendering-related information sent by the cloud device through the rendering engine.

[0345] In some embodiments, at least one of the local device and the cloud device also runs a rendering engine. In the local rendering mode, the rendering engine runs in the local device, and the local device receives the rendering-related information sent by the cloud device through the rendering engine.

[0346] For specific implementation details, refer to Figure 9 the embodiments, which will not be elaborated here.

[0347] Step 1432: Render the video stream data of the digital human according to the rendering-related information.

[0348] In some embodiments, in the local rendering mode, the rendering engine runs in the local device, and the local device renders the video stream data of the digital human according to the rendering-related information through the rendering engine.

[0349] For specific implementation details, refer to Figure 9 the embodiments, which will not be elaborated here.

[0350] Step 1421: In the cloud rendering mode, receive the video stream data of the digital human.

[0351] In some embodiments, at least one of the local device and the cloud device also runs a rendering engine. In the cloud rendering mode, the rendering engine runs in the cloud device. The rendering engine is used to render the rendering-related information to obtain the video stream data of the digital human and send the video stream data of the digital human to the local device.

[0352] For specific implementation details, refer to Figure 9 the embodiments, which will not be elaborated here.

[0353] Step 1440: Switch from the cloud rendering mode to the local rendering mode when the first switching condition is met; or switch from the local rendering mode to the cloud rendering mode when the second switching condition is met.

[0354] The local device flexibly selects to switch to the local rendering mode or the cloud rendering mode according to different scenarios, so as to meet the business requirements and improve the flexibility of rendering the digital human.

[0355] In some embodiments, when the first switching condition is satisfied, switching from the cloud rendering mode to the local rendering mode includes: when the first switching condition is satisfied, sending a first switching request to the cloud device and receiving a first switching response sent by the cloud device, where the first switching response is used to indicate that the cloud device agrees that the local device switches from the cloud rendering mode to the local rendering mode.

[0356] In some embodiments, when the first switching condition is satisfied, switching from the cloud rendering mode to the local rendering mode includes: when the first switching condition is satisfied, receiving a first switching instruction sent by the cloud device, where the first switching instruction is used to indicate that the local device switches from the cloud rendering mode to the local rendering mode.

[0357] By way of example and not limitation, the first switching condition includes at least one of the following:

[0358] The current idle computing resources of the local device are greater than (or greater than or equal to) a first idle threshold;

[0359] The current idle bandwidth of the local device is greater than (or greater than or equal to) a second idle threshold;

[0360] The data privacy level of the current service is higher than (or higher than or equal to) a first level, and the data privacy level is used to indicate the confidentiality requirements of the service;

[0361] The latency requirement level of the current service is higher than (or higher than or equal to) a second level, and the latency requirement level is used to indicate the timeliness requirements of the service.

[0362] For specific implementation details, refer to Figure 10 the embodiments, which will not be elaborated here.

[0363] In some embodiments, when the second switching condition is satisfied, switching from the local rendering mode to the cloud rendering mode includes: when the second switching condition is satisfied, sending a second switching request to the cloud device and receiving a second switching response sent by the cloud device, where the second switching response is used to indicate that the cloud device agrees that the local device switches from the local rendering mode to the cloud rendering mode.

[0364] In some embodiments, when the second switching condition is satisfied, switching from the local rendering mode to the cloud rendering mode includes: when the second switching condition is satisfied, receiving a second switching instruction sent by the cloud device, where the second switching instruction is used to indicate that the local device switches from the local rendering mode to the cloud rendering mode.

[0365] By way of example and not limitation, the second switching condition includes at least one of the following:

[0366] The current idle computing resources of the local device are greater than the first idle threshold;

[0367] The current available bandwidth of the local device is less than the second available threshold;

[0368] The data privacy level of the current service is lower than the first level, and the data privacy level is used to indicate the confidentiality requirements of the service;

[0369] The latency requirement level of the current service is lower than the second level, and the latency requirement level is used to indicate the timeliness requirements of the service.

[0370] For specific implementation details, refer to Figure 11 the embodiments, which will not be elaborated here.

[0371] In summary, in the method provided in this embodiment, at least one of the local device and the cloud device also runs a rendering engine. In the local rendering mode, the rendering engine runs in the local device; in the cloud rendering mode, the rendering engine runs in the cloud device. In different modes, the rendering engine runs in different devices, so as to adapt to the service requirements in different scenarios. For example, in the local rendering mode, the service requirements are mainly for low latency, and running the rendering engine in the local device can render the digital human more quickly; in the cloud rendering mode, the service requirements are mainly for high quality, and running the rendering engine in the cloud device can provide a higher-precision digital human.

[0372] The method provided in this embodiment also switches from the cloud rendering mode to the local rendering mode when the first switching condition is met; and switches from the local rendering mode to the cloud rendering mode when the second switching condition is met, so that the local device can switch the rendering mode according to different usage scenarios when the switching conditions are met, thereby meeting the service requirements.

[0373] Figure 16 The flowchart of the digital human rendering method provided by an exemplary embodiment of the present application is shown. This method is executed by a cloud device, and the cloud device supports the local rendering mode and the cloud rendering mode. This method includes at least one of the following steps.

[0374] Step 1610: Receive the input information sent by the local device.

[0375] Among them, the input information is used for the cloud device to generate rendering-related information associated with the digital human through the digital human generation service.

[0376] For specific implementation details, refer to Figure 14 step 1410 of the embodiment, which will not be elaborated here.

[0377] Step 1620: Generate rendering-related information according to the input information.

[0378] Among them, the rendering-related information is used to provide the information required for rendering the digital human.

[0379] The input information includes the user's questions or requests for the digital human, and the cloud device can generate corresponding rendering-related information according to different input information. In order for the digital human to respond to the input information more flexibly, it is necessary to use the rendering-related information to render the digital human, improving the interactivity and flexibility of the digital human.

[0380] Step 1630: In the local rendering mode, send the rendering-related information to the local device.

[0381] Among them, the rendering-related information is used for the local device to render the video stream data of the digital human.

[0382] For specific implementation details, refer to Figure 14 Steps 1420 and 1430 of the embodiment, which will not be elaborated here.

[0383] Step 1631: In the cloud rendering mode, render the video stream data of the digital human according to the rendering-related information; and send the video stream data of the digital human to the local device.

[0384] As Figure 1 shown in (a) of, the cloud device 200 renders the video stream data of the digital human according to the rendering-related information, and the cloud device 200 sends the video stream data of the digital human to the local device 100. Since the rendering operation is performed in the cloud device 200, the graphics card used by the cloud device 200 to render the digital human is better than the graphics card used by the local device 100, so the rendering quality of the digital human is higher.

[0385] In summary, the method provided in this embodiment is executed by the cloud device, which supports the local rendering mode and the cloud rendering mode. In the local rendering mode, the local device performs the rendering, and in the cloud rendering mode, the cloud device performs the rendering, so as to meet the business requirements in different scenarios and solve the problem that related technologies usually can only meet single business requirements.

[0386] Figure 17 The flowchart shows the digital human rendering method provided by an exemplary embodiment of the present application. This method is executed by the cloud device, which supports the local rendering mode and the cloud rendering mode. The method includes at least one of the following steps.

[0387] Step 1610: Receive the input information sent by the local device.

[0388] Step 1622: Generate the rendering-related information according to at least one of the first input information and the second input information.

[0389] In some embodiments, the input information includes at least one of first input information and second input information. The first input information is used to indicate the problem content or request content for which the digital human is to respond, and the second input information is used to indicate at least one of the conversation style and virtual image type of the digital human.

[0390] For specific implementation details, refer to Figure 4 the embodiments, which will not be elaborated here.

[0391] The cloud device can generate rendering-related information based on the first input information, or can generate rendering-related information based on the second input information, or can also generate rendering-related information based on the first input information and the second input information. The specific implementation details are as follows.

[0392] (1) Generating rendering-related information based on the first input information:

[0393] In some embodiments, the rendering-related information includes digital human dynamic information, which is used to indicate the movement information of at least one body part of the digital human. The cloud device includes an LLM, a TTS service, and a voice driving service, and the input information includes the first input information;

[0394] The cloud device generates conversation information based on the first input information through the LLM. The conversation information is used to respond to the problem content or request content; converts the conversation information into first voice information through the TTS service; and converts the first voice information into digital human dynamic information through the voice driving service.

[0395] The digital human dynamic information includes at least one of the limb movement information, facial expression information, and lip movement information of the digital human. The embodiments of the present application do not limit this.

[0396] For specific implementation details, refer to Figure 5 the embodiments, which will not be elaborated here.

[0397] (2) Generating rendering-related information based on the second input information:

[0398] In some embodiments, the rendering-related information includes digital human dynamic information, which is used to indicate the movement information of at least one body part of the digital human. The cloud device includes an LLM, a TTS service, and a voice driving service, and the input information includes the second input information;

[0399] The cloud device generates conversation information through the LLM according to the preset text and the conversation style of the digital human, and the conversation information is used to respond to the question content or request content; converts the conversation information into the second voice information corresponding to the first voice through the TTS service according to the first voice, and the first voice matches the virtual image type of the digital human; converts the second voice information into the digital human dynamic information corresponding to the first virtual model through the voice driving service according to the first virtual model, and the first virtual model matches the virtual image type of the digital human.

[0400] For specific implementation details, refer to Figure 6 the embodiments, which will not be elaborated here.

[0401] (3) Generate rendering-related information according to the first input information and the second input information:

[0402] In some embodiments, the rendering-related information includes digital human dynamic information, and the digital human dynamic information is used to indicate the motion information of at least one body part of the digital human. The cloud device includes an LLM, a TTS service, and a voice driving service, and the input information includes the first input information and the second input information;

[0403] The cloud device 2 generates conversation information through the LLM according to the first input information and the conversation style of the digital human, and the conversation information is used to respond to the question content or request content; converts the conversation information into the third voice information corresponding to the second voice through the TTS service according to the second voice, and the second voice matches the virtual image type of the digital human; converts the third voice information into the digital human dynamic information corresponding to the second virtual model through the voice driving service according to the second virtual model, and the second virtual model matches the virtual image type of the digital human.

[0404] For specific implementation details, refer to Figure 7 the embodiments, which will not be elaborated here.

[0405] In some embodiments, the input information further includes conversation context information, and the conversation context information is used to indicate the historical content of the digital human conversation. The method further includes:

[0406] The cloud device uses the conversation context information as reference information when generating conversation information through the LLM, and the conversation information is used to respond to the question content or request content.

[0407] In some embodiments, the digital human generation service includes at least one of a KWS (keyword spotting) service and an ASR (automatic speech recognition) service. The method further includes:

[0408] Wake up the digital human through the KWS service;

[0409] Receive the input information through the ASR service and convert the voice information in the input information into text information.

[0410] For specific implementation details, please refer to Figure 4 the embodiments, which will not be elaborated here.

[0411] Step 1630: In the local rendering mode, send rendering-related information to the local device.

[0412] Step 1631: In the cloud rendering mode, render the video stream data of the digital human according to the rendering-related information; and send the video stream data of the digital human to the local device.

[0413] In summary, the method provided in this embodiment generates rendering-related information based on at least one of the first input information and the second input information, making the rendering-related information associated with the input information. Finally, the rendered digital human also has a stronger association with the input information, improving the interaction experience between the user and the digital human.

[0414] The method provided in this embodiment also enables the digital human to respond to the user's voice commands in real time through the speech synthesis service and the speech recognition service; by using the LLM as the virtual brain of the digital human, the digital human's answers are more natural and effective; through the speech-driven service, the real-time rendering of the digital human's facial expressions and body movements is realized, improving the interaction experience between the user and the digital human.

[0415] Figure 18 The flowchart of the digital human rendering method provided by an exemplary embodiment of the present application is shown. This method is executed by a cloud device that supports the local rendering mode and the cloud rendering mode. The method includes at least one of the following steps.

[0416] Step 1610: Receive the input information sent by the local device.

[0417] Step 1620: Generate rendering-related information according to the input information.

[0418] Step 1632: In the local rendering mode, send the rendering-related information to the rendering engine of the local device.

[0419] In some embodiments, in the local rendering mode, the rendering engine runs in the local device. The rendering engine is used to render the rendering-related information to obtain the video stream data of the digital human.

[0420] For specific implementation details, please refer to Figure 9 the embodiments, which will not be elaborated here.

[0421] Step 1633: In the cloud rendering mode, use the rendering engine to render the video stream data of the digital human according to the rendering-related information; and use the rendering engine to send the video stream data of the digital human to the local device.

[0422] In some embodiments, in the cloud rendering mode, the rendering engine runs in a cloud device. The rendering engine is used to render rendering-related information to obtain the video stream data of the digital human.

[0423] For specific implementation details, refer to Figure 9 the embodiments, which will not be elaborated here.

[0424] Step 1640: Switch from the cloud rendering mode to the local rendering mode when the first switching condition is met; or switch from the local rendering mode to the cloud rendering mode when the second switching condition is met.

[0425] The cloud device can flexibly select whether to switch to the local rendering mode or the cloud rendering mode according to different scenarios, so as to meet the business requirements and improve the flexibility of rendering the digital human.

[0426] In some embodiments, switching from the cloud rendering mode to the local rendering mode when the first switching condition is met includes:

[0427] When the first switching condition is met, receiving a first switching request sent by the local device; switching from the cloud rendering mode to the local rendering mode; and sending a first switching response to the local device, where the first switching response is used to indicate that the cloud device agrees that the local device switches from the cloud rendering mode to the local rendering mode.

[0428] In some embodiments, switching from the cloud rendering mode to the local rendering mode when the first switching condition is met includes:

[0429] When the first switching condition is met, sending a first switching instruction to the local device and switching from the cloud rendering mode to the local rendering mode, where the first switching instruction is used to indicate that the local device switches from the cloud rendering mode to the local rendering mode.

[0430] By way of example and not limitation, the first switching condition includes at least one of the following:

[0431] The current idle computing resources of the local device are greater than (or greater than or equal to) the first idle threshold;

[0432] The current idle bandwidth of the local device is greater than (or greater than or equal to) the second idle threshold;

[0433] The data privacy level of the current service is higher than (or higher than or equal to) the first level, and the data privacy level is used to indicate the confidentiality requirements of the service;

[0434] The latency requirement level of the current service is higher than (or higher than or equal to) the second level, and the latency requirement level is used to indicate the timeliness requirements of the service.

[0435] For specific implementation details, refer toFigure 10 The embodiments are not described herein again.

[0436] In some embodiments, when the second switching condition is satisfied, switching from the local rendering mode to the cloud rendering mode includes:

[0437] When the second switching condition is satisfied, receiving a second switching request sent by the local device; switching from the local rendering mode to the cloud rendering mode; and sending a second switching response to the local device, where the second switching response is used to indicate that the cloud device agrees that the local device switches from the local rendering mode to the cloud rendering mode.

[0438] In some embodiments, when the second switching condition is satisfied, switching from the local rendering mode to the cloud rendering mode includes:

[0439] When the second switching condition is satisfied, sending a second switching instruction to the local device, and switching from the local rendering mode to the cloud rendering mode, where the second switching instruction is used to indicate that the local device switches from the local rendering mode to the cloud rendering mode.

[0440] By way of example and not limitation, the second switching condition includes at least one of the following:

[0441] The current idle computing resources of the local device are less than the first idle threshold;

[0442] The current idle bandwidth of the local device is less than the second idle threshold;

[0443] The data privacy level of the current service is lower than the first level, and the data privacy level is used to indicate the confidentiality requirements of the service;

[0444] The latency requirement level of the current service is lower than the second level, and the latency requirement level is used to indicate the timeliness requirements of the service.

[0445] For specific implementation details, refer to Figure 11 The embodiments are not described herein again.

[0446] In summary, in the method provided in this embodiment, at least one of the local device and the cloud device also runs a rendering engine. In the local rendering mode, the rendering engine runs in the local device; in the cloud rendering mode, the rendering engine runs in the cloud device. In different modes, the rendering engine runs in different devices, so as to adapt to the service requirements in different scenarios. For example, in the local rendering mode, the service requirements are mainly for low latency, and running the rendering engine in the local device can render the digital human more quickly; in the cloud rendering mode, the service requirements are mainly for high quality, and running the rendering engine in the cloud device can provide a digital human with higher precision.

[0447] The method provided in this embodiment also switches from the cloud rendering mode to the local rendering mode when the first switching condition is met; and switches from the local rendering mode to the cloud rendering mode when the second switching condition is met, enabling the cloud device to switch the rendering mode according to different usage scenarios when the switching conditions are met, so as to meet the business requirements.

[0448] Figure 19 FIG. 4 shows a block diagram of a computer device 1900 provided by an exemplary embodiment of the present application. This computer device can be used to implement the digital human rendering method provided in the above embodiments. The computer device 1900 includes a central processing unit (CPU) 1901, a system memory 1904 including a random access memory (RAM) 1902 and a read-only memory (ROM) 1903, and a system bus 1905 connecting the system memory 1904 and the central processing unit 1901. The computer device 1900 also includes an input / output (I / O) system 1906 for facilitating the transfer of information between various components within the computer device, and a mass storage device 1907 for storing an operating system 1913, application programs 1914, and other program modules 1915.

[0449] The input / output system 1906 includes a display 1908 for displaying information and input devices 1909 such as a mouse and a keyboard for user input. Both the display 1908 and the input devices 1909 are connected to the central processing unit 1901 through an input / output controller 1910 connected to the system bus 1905. The input / output system 1906 may also include an input / output controller 1910 for receiving and processing inputs from multiple other devices such as a keyboard, a mouse, or an electronic stylus. Similarly, the input / output controller 1910 also provides outputs to a display screen, a printer, or other types of output devices.

[0450] The mass storage device 1907 is connected to the central processing unit 1901 through a mass storage controller (not shown) connected to the system bus 1905. The mass storage device 1907 and its associated computer-readable storage medium provide non-volatile storage for the computer device 1900. That is to say, the mass storage device 1907 may include a computer-readable storage medium (not shown) such as a hard disk or a compact disc read-only memory (CD-ROM) drive.

[0451] Without loss of generality, the computer-readable storage medium may include a computer storage medium and a communication medium. The computer storage medium includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable storage instructions, data structures, program modules, or other data. The computer storage medium includes RAM, ROM, erasable programmable read-only memory (EPROM), electrically-erasable programmable read-only memory (EEPROM), flash memory or other solid-state storage technologies, CD-ROM, digital versatile disc (DVD) or other optical storage, magnetic tape cartridges, tapes, magnetic disk storage or other magnetic storage devices. Of course, those skilled in the art will appreciate that the computer storage medium is not limited to the above several types. The above-mentioned system memory 1904 and mass storage device 1907 may be collectively referred to as memory.

[0452] The memory stores one or more programs, and the one or more programs are configured to be executed by one or more central processing units 1901. The one or more programs include instructions for implementing the above method embodiments. The central processing unit 1901 executes the one or more programs to implement the digital human rendering method provided by each of the above method embodiments.

[0453] According to various embodiments of the present application, the computer device 1900 may also run by connecting to a remote terminal device on the network through a network such as the Internet. That is, the computer device 1900 may be connected to the network 1912 through the network interface unit 1911 connected to the system bus 1905, or in other words, the network interface unit 1911 may also be used to connect to other types of networks.

[0454] The memory further includes one or more programs. The one or more programs are stored in the memory and include steps for performing the method executed by the computer device provided in the embodiments of the present application.

[0455] The embodiments of the present application also provide a computer device, which includes: a processor and a memory, and at least one segment of program is stored in the memory; the processor is used to execute at least one segment of program in the memory to implement the digital human rendering method provided by each of the above method embodiments on the local device side.

[0456] An embodiment of the present application also provides a computer device, which includes: a processor and a memory, and at least one segment of program is stored in the memory; the processor is configured to execute at least one segment of program in the memory to implement the digital human rendering method provided by each of the above method embodiments on the cloud device side.

[0457] An embodiment of the present application also provides a computer-readable storage medium, in which at least one segment of program is stored, and the at least one segment of program is loaded and executed by a processor to implement the digital human rendering method provided by each of the above method embodiments.

[0458] An embodiment of the present application also provides a computer program product, which includes computer instructions, the computer instructions are stored in a computer-readable storage medium, the processor obtains the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to implement the digital human rendering method provided by each of the above method embodiments.

[0459] It should be understood that "a plurality of" mentioned herein refers to two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.

[0460] Those skilled in the art should be able to realize that in the above one or more examples, the functions described in the embodiments of the present application can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable storage medium or transmitted as one or more instructions or codes on a computer-readable storage medium. The computer-readable storage medium includes computer storage media and communication media, where the communication media includes any medium that facilitates the transmission of a computer program from one place to another. The storage medium can be any available medium accessible by a general-purpose or special-purpose computer.

[0461] The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A computer system for digital human rendering, characterized in that, The computer system includes a local device and a cloud device, and the computer system supports a local rendering mode and a cloud rendering mode; The local device is used to send input information to the cloud device, and the input information is used for the cloud device to generate rendering-related information associated with the digital human through a digital human generation service; The cloud device is used to generate the rendering-related information according to the input information, and the rendering-related information is used to provide the information required for rendering the digital human; In the local rendering mode, the cloud device is used to send the rendering-related information to the local device, and the local device is used to render the video stream data of the digital human according to the rendering-related information; In the cloud rendering mode, the cloud device is used to render the video stream data of the digital human according to the rendering-related information, and the local device is used to receive the video stream data of the digital human.

2. The computer system according to claim 1, wherein The input information includes at least one of first input information and second input information. The first input information is used to indicate the question content or request content that the digital human is to respond to, and the second input information is used to indicate at least one of the conversation style and virtual image type of the digital human; The cloud device is used to generate the rendering-related information according to at least one of the first input information and the second input information.

3. The computer system according to claim 2, wherein, The rendering-related information includes digital human dynamic information, and the digital human dynamic information is used to indicate the motion information of at least one body part of the digital human. The cloud device includes a large language model LLM, a text-to-speech TTS service, and a voice driving service, and the input information includes the first input information; The cloud device is used to generate conversation information according to the first input information through the LLM, and the conversation information is used to respond to the question content or the request content; Convert the conversation information into first voice information through the TTS service; Convert the first voice information into the digital human dynamic information through the voice driving service.

4. The computer system according to claim 2, wherein The rendering-related information includes digital human dynamic information, and the digital human dynamic information is used to indicate the motion information of at least one body part of the digital human. The cloud device includes an LLM, a TTS service, and a voice driving service, and the input information includes the second input information; The cloud device is used to generate conversation information according to a preset text and the conversation style of the digital human through the LLM, and the conversation information is used to respond to the question content or the request content; Convert the conversation information into second voice information corresponding to the first voice tone through the TTS service according to a first voice tone, and the first voice tone matches the virtual image type of the digital human; Convert the second voice information into digital human dynamic information corresponding to the first virtual model through the voice driving service according to a first virtual model, and the first virtual model matches the virtual image type of the digital human.

5. The computer system according to claim 2, wherein The rendering-related information includes digital human dynamic information, which is used to indicate the motion information of at least one body part of the digital human. The cloud device includes an LLM, a TTS service, and a voice driving service. The input information includes the first input information and the second input information; The cloud device is configured to generate conversation information through the LLM according to the first input information and the conversation style of the digital human, and the conversation information is used to respond to the question content or the request content; Convert the conversation information into third voice information corresponding to the second voice color through the TTS service according to the second voice color, and the second voice color matches the virtual image type of the digital human; Convert the third voice information into digital human dynamic information corresponding to the second virtual model through the voice driving service according to the second virtual model, and the second virtual model matches the virtual image type of the digital human.

6. The computer system according to any one of claims 2 to 5, characterized in that, The input information further includes conversation context information, which is used to indicate the historical content of the digital human's conversation; The cloud device is further configured to use the conversation context information as reference information when generating conversation information through the LLM, and the conversation information is used to respond to the question content or the request content.

7. The computer system according to claim 1, wherein The local device runs a client, the cloud device also runs an access gateway, and at least one of the local device and the cloud device also runs a rendering engine; The client is configured to send the input information to the access gateway; The access gateway is configured to call the digital human generation service according to the input information to generate the rendering-related information; The access gateway is configured to send the rendering-related information to the rendering engine; The rendering engine is configured to render the rendering-related information to obtain the video stream data of the digital human, and send the video stream data of the digital human to the client; Wherein, in the local rendering mode, the rendering engine runs in the local device; In the cloud rendering mode, the rendering engine runs in the cloud device.

8. The computer system according to any one of claims 1 to 7, characterized in that, The digital human generation service includes at least one of a voice wake-up KWS service and a speech recognition ASR service; The cloud device is configured to wake up the digital human through the KWS service; The cloud device is configured to receive the input information through the ASR service and convert the voice information in the input information into text information.

9. The computer system according to any one of claims 1 to 8, wherein In the case of satisfying the first switching condition, the local device is configured to send a first switching request to the cloud device, and the cloud device is configured to send a first switching response to the local device, and the first switching response is used to indicate that the cloud device agrees that the local device switches from the cloud rendering mode to the local rendering mode.

10. The computer system according to any one of claims 1 to 8, wherein When the first switching condition is met, the cloud device is configured to send a first switching instruction to the local device, and the first switching instruction is used to instruct the local device to switch from the cloud rendering mode to the local rendering mode.

11. The computer system according to claim 9 or 10, characterized in that The first switching condition includes at least one of the following: The current idle computing resources of the local device are greater than a first idle threshold; The current idle bandwidth of the local device is greater than a second idle threshold; The data privacy level of the current service is higher than a first level, and the data privacy level is used to indicate the confidentiality requirements of the service; The latency requirement level of the current service is higher than a second level, and the latency requirement level is used to indicate the timeliness requirements of the service.

12. The computer system according to any one of claims 1 to 8, wherein When the second switching condition is met, the local device is configured to send a second switching request to the cloud device, and the cloud device is configured to send a second switching response to the local device, and the second switching response is used to indicate that the cloud device agrees that the local device switches from the local rendering mode to the cloud rendering mode.

13. The computer system according to any one of claims 1 to 8, wherein When the second switching condition is met, the cloud device is configured to send a second switching instruction to the local device, and the second switching instruction is used to instruct the local device to switch from the local rendering mode to the cloud rendering mode.

14. The computer system according to claim 12 or 13, characterized in that, The second switching condition includes at least one of the following: The current idle computing resources of the local device are less than a first idle threshold; The current idle bandwidth of the local device is less than a second idle threshold; The data privacy level of the current service is lower than a first level, and the data privacy level is used to indicate the confidentiality requirements of the service; The latency requirement level of the current service is lower than a second level, and the latency requirement level is used to indicate the timeliness requirements of the service.

15. A local device for digital human rendering, characterized in that, The local device supports the local rendering mode and the cloud rendering mode. The local device includes a client module, or includes the client module and a rendering engine module; The client module is configured to send input information to the cloud device, and the input information is used for the cloud device to generate rendering-related information associated with the digital human through the digital human generation service module; In the local rendering mode, the client module is configured to receive the rendering-related information sent by the cloud device, and the rendering-related information is used to provide the information required for rendering the digital human; The rendering engine module is configured to render the video stream data of the digital human according to the rendering-related information, and send the video stream data of the digital human to the client module; In the cloud rendering mode, the client module is configured to receive the video stream data of the digital human, and the video stream data of the digital human is rendered by the cloud device according to the rendering-related information.

16. A cloud device for digital human rendering, characterized in that The cloud device supports the local rendering mode and the cloud rendering mode. The cloud device includes an access gateway module and a digital human generation service module, or includes the access gateway module, the digital human generation service module and a rendering engine module; The digital human generation service module is used to receive input information sent by a local device, where the input information is used for the cloud device to generate rendering-related information associated with the digital human through the digital human generation service module; The digital human generation service module is used to generate the rendering-related information according to the input information, where the rendering-related information is used to provide the information required for rendering the digital human; In the local rendering mode, the access gateway module is used to send the rendering-related information to the local device, where the rendering-related information is used for the local device to render the video stream data of the digital human; In the cloud rendering mode, the rendering engine module is used to render the video stream data of the digital human according to the rendering-related information, and send the video stream data of the digital human to the local device.

17. A digital human rendering method, characterized in that, The method is executed by a local device, where the local device supports the local rendering mode and the cloud rendering mode. The method includes: Sending input information to a cloud device, where the input information is used for the cloud device to generate rendering-related information associated with the digital human through the digital human generation service; In the local rendering mode, receiving the rendering-related information sent by the cloud device, where the rendering-related information is used to provide the information required for rendering the digital human; rendering the video stream data of the digital human according to the rendering-related information; In the cloud rendering mode, receiving the video stream data of the digital human, where the video stream data of the digital human is rendered by the cloud device according to the rendering-related information.

18. A digital human rendering method, characterized in that, The method is executed by a cloud device, where the cloud device supports the local rendering mode and the cloud rendering mode. The method includes: Receiving input information sent by a local device, where the input information is used for the cloud device to generate rendering-related information associated with the digital human through the digital human generation service; Generating the rendering-related information according to the input information, where the rendering-related information is used to provide the information required for rendering the digital human; In the local rendering mode, sending the rendering-related information to the local device, where the rendering-related information is used for the local device to render the video stream data of the digital human; In the cloud rendering mode, rendering the video stream data of the digital human according to the rendering-related information, and sending the video stream data of the digital human to the local device.

19. A computer-readable storage medium, characterized in that, At least one program is stored in the computer-readable storage medium, and the at least one program is loaded and executed by a processor to implement the digital human rendering method as claimed in claim 17; and / or, the digital human rendering method as claimed in claim 18.

20. A computer program product, characterized in that, The computer program product includes computer instructions, where the computer instructions are stored in a computer-readable storage medium, and the processor obtains the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to implement the digital human rendering method as claimed in claim 17; and / or, the digital human rendering method as claimed in claim 18.

Citation Information

Cited By

  • Task scheduling method and system based on hybrid cloud digital human processing architecture

    CN122179477A