Distributed digital human interaction method and distributed digital human interaction system
Through the distributed digital human interaction method, the multimodal data is streamed using the lip-synchronous animation generation model, which solves the problem of digital human interaction delay in the prior art, and realizes low-latency multimodal data processing and broadcasting.
Patent Information
- Application Number
- CN202510231551.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-20
AI Technical Summary
The existing digital human interaction system has obvious delays during interaction, which cannot meet the needs of low-latency interaction, especially in multi-user scenarios, the delay problem is more prominent.
The distributed digital human interaction method is adopted to obtain multimodal data and stream processing based on the lip-synchronous animation generation model. The processing results are generated for rendering digital human audio and video streaming data and displaying text content, realizing instant multimodal data processing and broadcasting.
Realize instant multimodal data processing and broadcasting, reduce the delay of digital human interaction, meet the needs of low-latency interaction, and can effectively solve the delay problem especially in multi-user scenarios.
Smart Images

Figure CN120179063A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of digital human interaction technology, and in particular to a distributed digital human interaction method, a distributed digital human interaction system, an electronic device, and a computer-readable storage medium. Background Art
[0002] With the continuous breakthroughs in science and technology and the accelerated evolution of artificial intelligence, the digital human industry is welcoming unprecedented development opportunities. From film and television entertainment to marketing, from e-commerce live broadcasting to financial services, digital humans have penetrated into various industries in multiple forms, reshaping the business ecology and user experience. At present, the digital human interaction system has obvious delays when interacting. Summary of the invention
[0003] The embodiments of the present application provide a distributed digital human interaction method, a distributed digital human interaction system, an electronic device and a computer-readable storage medium to solve at least one of the above-mentioned technical problems.
[0004] The distributed digital human interaction method of the present application embodiment includes:
[0005] Acquire multimodal data;
[0006] Performing streaming processing on the multimodal data based on a lip sync animation generation model to obtain processing results corresponding to the multimodal data;
[0007] Rendering is performed based on the processing result to obtain digital human audio and video stream data;
[0008] The text content is displayed based on the processing result, and the audio and video broadcast of the digital human is performed based on the digital human audio and video stream data.
[0009] In some embodiments, the multimodal data includes any one or more of audio stream data, video stream data, and click event data.
[0010] In some embodiments, the processing result includes recognized speech text, answer speech text, answer audio data and facial drive data, and the lip sync animation generation model is used to process the multimodal data to obtain the processing result corresponding to the multimodal data, including:
[0011] Calling a speech recognition model to obtain the recognized speech text according to the multimodal data;
[0012] Calling a natural language model to obtain the answer voice text according to the recognized voice text;
[0013] Calling a speech synthesis model to obtain the answer audio data according to the answer speech text;
[0014] Invoke the lip-sync animation generation model to stream generate the facial drive data according to the response audio data.
[0015] In some embodiments, the invoking the lip-sync animation generation model to stream generate the facial drive data according to the response audio data includes:
[0016] During the process of the lip-sync animation generation model obtaining the response audio data, perform non-global feature extraction on the response audio data through a transformer network to obtain Mel-frequency cepstral coefficients features;
[0017] Encode the Mel-frequency cepstral coefficients features through a temporal convolutional network to obtain temporal convolutional features;
[0018] Perform global calculation on the temporal convolutional features through the transformer network to obtain a first transformed feature;
[0019] Reconstruct the window and stride of the first transformed feature to obtain a second transformed feature;
[0020] Encode and decode the second transformed feature based on a multi-head attention mechanism to obtain the facial drive data.
[0021] In some embodiments, the performing non-global feature extraction on the response audio data through a transformer network to obtain Mel-frequency cepstral coefficients features includes:
[0022] Perform non-global feature extraction on the response audio data through the transformer network to obtain Mel-frequency cepstral coefficients;
[0023] Calculate corresponding first-order difference coefficients and second-order difference coefficients based on the Mel-frequency cepstral coefficients;
[0024] Determine the corresponding Mel-frequency cepstral coefficients features based on the Mel-frequency cepstral coefficients, the first-order difference coefficients, and the second-order difference coefficients.
[0025] In some embodiments, the rendering based on the processing result to obtain digital human audio-visual stream data includes:
[0026] Perform three-dimensional rendering on a three-dimensional model based on the response audio data and the facial drive data to obtain the digital human audio-visual stream data;
[0027] The displaying the text content based on the processing result and the audio-visual broadcast of the digital human based on the digital human audio-visual stream data includes:
[0028] Display the text content of the recognized speech text and the response speech text, and perform audio and video broadcast of the digital human based on the digital human audio and video stream data.
[0029] The distributed digital human interaction system according to the embodiment of the present application includes a front-end interaction module, a back-end server module, an algorithm server module, and a rendering and streaming module;
[0030] The front-end interaction module is used to obtain multimodal data;
[0031] The back-end server module is used to call the algorithm server module to process the multimodal data based on the lip-sync animation generation model to obtain the processing result corresponding to the multimodal data;
[0032] The rendering and streaming module is used to perform rendering based on the processing result to obtain digital human audio and video stream data;
[0033] The front-end interaction module is further used to display the text content based on the processing result and perform audio and video broadcast of the digital human based on the digital human audio and video stream data.
[0034] In some embodiments, the processing result includes a recognized speech text, a response speech text, response audio data, and facial drive data. The back-end server module includes a core server, and the algorithm server module includes a speech recognition model, a natural language model, a speech synthesis model, and a lip-sync animation generation model. The core server is used to call the speech recognition model, the natural language model, the speech synthesis model, and the lip-sync animation generation model;
[0035] The speech recognition model is used to obtain the recognized speech text according to the multimodal data;
[0036] The natural language model is used to obtain the response speech text according to the recognized speech text;
[0037] The speech synthesis model is used to obtain the response audio data according to the response speech text;
[0038] The lip-sync animation generation model is used to stream generate the facial drive data according to the response audio data.
[0039] In some embodiments, the front-end interaction module includes an interaction interface, and the rendering and streaming module includes an engine service unit;
[0040] The core server is used to output the recognized speech text and the response speech text to the front-end interaction module, and output the response audio data and the facial drive data to the engine service unit;
[0041] The engine service unit is used to perform 3D rendering on the 3D model based on the response audio data and the facial drive data to obtain the digital human audio-visual stream data;
[0042] The interaction interface is used to display the text content of the recognized speech text and the response speech text, and perform audio-visual broadcast of the digital human based on the digital human audio-visual stream data.
[0043] In some embodiments, the backend server module further includes a web server, and the rendering and streaming module further includes a signaling server;
[0044] The web server is used to send page data, the first address information of the signaling server, and the second address information of the core server to the front-end interaction module;
[0045] The front-end interaction module is used to obtain the interaction interface based on the page data, implement communication with the signaling server based on the first address data, and implement communication with the core server based on the second address data.
[0046] For the electronic device according to the embodiment of the present application, the electronic device includes one or more processors and a memory, and the memory stores a computer program. When the computer program is executed by the processor, the distributed digital human interaction method according to any one of the above embodiments is implemented.
[0047] For the computer-readable storage medium according to the embodiment of the present application, a computer program is stored thereon. When the program is executed by the processor, the distributed digital human interaction method according to any one of the above embodiments is implemented.
[0048] In the distributed digital human interaction method, distributed digital human interaction system, electronic device, and computer-readable storage medium according to the embodiments of the present application, multi-modal data is processed in a streaming manner based on a lip-sync animation generation model to obtain corresponding processing results; the digital human audio-visual stream data is rendered based on the processing results; the processing results are used for text content display, and the digital human audio-visual stream data is used for audio-visual broadcast of the digital human. In this way, the processing results of multi-modal data can be obtained immediately, and then text content display and digital human audio-visual broadcast can be performed faster, meeting the low-latency interaction requirements of digital human interaction.
[0049] The additional aspects and advantages of the embodiments of the present application will be partially given in the following description, partially become apparent from the following description, or be understood through the practice of the embodiments of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the description of the embodiments in conjunction with the following drawings, wherein:
[0051] Figure 1 is a schematic flowchart of a distributed digital human interaction method according to some embodiments of the present application;
[0052] Figure 2 is a schematic diagram of modules of a distributed digital human interaction system according to some embodiments of the present application;
[0053] Figure 3 is a schematic diagram of an interaction interface according to some embodiments of the present application;
[0054] Figure 4 is a schematic diagram of an interaction interface according to some embodiments of the present application;
[0055] Figure 5 is a schematic flowchart of a distributed digital human interaction method according to some embodiments of the present application;
[0056] Figure 6 is a schematic flowchart of a distributed digital human interaction method according to some embodiments of the present application;
[0057] Figure 7 is a schematic diagram of the working process of a lip synchronization animation generation model according to some embodiments of the present application;
[0058] Figure 8 is a schematic flowchart of a distributed digital human interaction method according to some embodiments of the present application;
[0059] Figure 9 is a schematic flowchart of a distributed digital human interaction method according to some embodiments of the present application;
[0060] Figure 10 is a schematic flowchart of a distributed digital human interaction method according to some embodiments of the present application;
[0061] Figure 11 is a schematic diagram of modules of an electronic device according to some embodiments of the present application;
[0062] Figure 12 is a schematic diagram of the connection state between a computer-readable storage medium and a processor according to some embodiments of the present application.
[0063] Reference numerals:
[0064] Distributed digital human interaction system 100, front-end interaction module 10, interaction interface 11, back-end server module 20, core server 21, web server 22, algorithm server module 30, speech recognition model 31, natural language model 32, speech synthesis model 33, lip-sync animation generation model 34, cache unit 35, rendering and streaming module 40, engine service unit 41, signaling server 42, electronic device 200, processor 210, memory 220, computer-readable storage medium 300, computer program 310, processor 320. Detailed implementation manners
[0065] The following further describes the implementation manners of the present application with reference to the accompanying drawings. The same or similar reference numerals in the drawings always denote the same or similar elements or elements with the same or similar functions. In addition, the implementation manners of the present application described below with reference to the accompanying drawings are exemplary and are only used to explain the implementation manners of the present application, and should not be construed as a limitation to the present application.
[0066] Please refer to Figure 1 and Figure 2 , the implementation manners of the present application provide a distributed digital human interaction method. The distributed digital human interaction method includes:
[0067] 010: Obtain multimodal data;
[0068] 020: Perform streaming processing on the multimodal data based on the lip-sync animation generation model 34 to obtain a processing result corresponding to the multimodal data;
[0069] 030: Perform rendering based on the processing result to obtain digital human audio-visual stream data;
[0070] 040: Display the text content based on the processing result, and perform audio-visual broadcast of the digital human based on the digital human audio-visual stream data.
[0071] In the distributed digital human interaction method of the implementation manners of the present application, streaming processing is performed on the multimodal data based on the lip-sync animation generation model 34 to obtain a corresponding processing result; digital human audio-visual stream data is obtained by rendering based on the processing result; the processing result is used for text content display, and the digital human audio-visual stream data is used for audio-visual broadcast of the digital human. In this way, the processing result of the multimodal data can be obtained immediately, and then text content display and digital human audio-visual broadcast can be performed faster, meeting the low-latency interaction requirements of digital human interaction.
[0072] Specifically, the distributed digital human interaction method can be applied to the distributed digital human interaction system 100 (such as Figure 2as shown). The distributed digital human interaction system 100 includes a front-end interaction module 10, a back-end server module 20, an algorithm server module 30, and a rendering and streaming module 40. The front-end interaction module 10 can be configured on electronic devices such as mobile phones, tablets, and laptops. The front-end interaction module 10 can display a digital human (such as Figure 3 as shown), and users can interact with the digital human through multimodal methods such as graphic buttons, microphones, and cameras on the front-end interaction module 10, so that the front-end interaction module 10 can obtain the multimodal data input by the user.
[0073] The front-end interaction module 10 is connected to the back-end server module 20 to send the multimodal data to the back-end server module 20. The back-end server module 20 is connected to the algorithm server module 30 to send the multimodal data to the algorithm server module 30 and call the algorithm server module 30 to process the multimodal data.
[0074] The algorithm server module 30 can include a variety of deep learning models, and the specific models can be configured according to the actual application situation. For example, the algorithm server module 30 can include a lip-sync animation generation model 34. Based on the lip-sync animation generation model 34, the multimodal data is processed in a streaming manner to obtain the corresponding processing result, and the processing result is sent to the back-end server module 20.
[0075] Streaming processing refers to the continuous and real-time processing of data streams, which has the characteristics of low latency and high throughput. Processing the multimodal data based on the lip-sync animation generation model 34 means that during the process of the algorithm server module 30 receiving the multimodal data, the lip-sync animation generation model 34 can process the multimodal data.
[0076] The back-end server module 20 can send the processing result to the front-end interaction module 10, and the front-end interaction module 10 can display the text content of the processing result. The back-end server module 20 is also connected to the rendering and streaming module 40 to send the processing result to the rendering and streaming module 40. The rendering and streaming module 40 can render the processing result to obtain the digital human audio-visual stream data.
[0077] The rendering and streaming module 40 is connected to the front-end interaction module 10 to send the digital human audio-visual stream data to the front-end interaction module 10. The front-end interaction module 10 can perform audio-visual broadcast of the digital human based on the digital human audio-visual stream data, thereby realizing the interaction between the user and the digital human. Audio-visual broadcast means that the front-end interaction module 10 can display the animation video of the digital human and broadcast the audio, and the animation video includes facial movements corresponding to the audio, such as lip movements.
[0078] In the embodiments of the present application, the multi-modal data is processed in a streaming manner based on the lip-sync animation generation model 34, and the corresponding processing results can be generated in a streaming manner, having an instant generation ability. In this way, the processing results of the multi-modal data can be obtained instantaneously, and then the text content can be displayed and the digital human audio and video can be broadcasted faster, meeting the low-latency interaction requirements of digital human interaction.
[0079] In some embodiments, the multi-modal data includes any one or more of audio stream data, video stream data, and click event data.
[0080] Specifically, the voice input by the user through the microphone is the audio stream data, and the video captured by the user through the camera is the video stream data. The front-end interaction module 10 may include an interaction interface 11, and the interaction interface 11 may be provided with interactive components such as graphic buttons and scroll bars for the user to operate. As Figure 4 shown, the interaction interface 11 is provided with an "intelligent Q&A" image button. When the user performs touch operations such as clicking on the graphic button, dragging the page, dragging the scroll bar, and long pressing, the front-end interaction module 10 can obtain click event data.
[0081] In practical applications, the user can click the "intelligent Q&A" image button on the interaction interface 11, and the front-end interaction module 10 sends the click event data to the back-end server module 20. The back-end server module 20 can send response data to the front-end interaction module 10 according to the click event data for corresponding feedback display on the interaction interface 11. For example, the lower-level interface of "intelligent Q&A" is displayed on the interaction interface 11. The specific feedback display can be set according to the actual application situation and is not limited here.
[0082] The back-end server module 20 can also send a driving instruction to the rendering and streaming module 40 according to the click event data. The rendering and streaming module 40 generates corresponding interactive feedback data according to the driving instruction and sends the interactive feedback data to the front-end interaction module 10 for corresponding visual feedback display on the interaction interface 11. The specific visual feedback display can be set according to the actual application situation and is not limited here.
[0083] The user can input voice through the microphone in the front-end interaction module 10, and the front-end interaction module 10 obtains the audio stream data. The audio stream data is processed by the back-end server module 20, the algorithm server module 30, and the rendering and streaming module 40, so that the front-end interaction module 10 can display the text content based on the processing results and perform audio and video broadcasts based on the digital human audio and video stream data to give a corresponding answer feedback to the voice input by the user, realizing the Q&A interaction between the user and the digital human.
[0084] In one example, while the user inputs voice through the microphone in the front-end interaction module 10, the user can also input video through the camera and / or operate on the interaction components. Thus, while the front-end interaction module 10 obtains audio stream data, it also obtains video stream data and / or click event data. The back-end server module 20, the algorithm server module 30, and the rendering and streaming module 40 process the audio stream data, the video stream data, and / or the click event data together to obtain a more accurate processing result.
[0085] Please refer to Figure 2 and Figure 5 , in some embodiments, the processing result includes recognized speech text, replied speech text, replied audio data, and facial drive data. Based on the lip-sync animation generation model 34, the multi-modal data is processed in a streaming manner to obtain the processing result corresponding to the multi-modal data (i.e., 020), including:
[0086] 021: Invoke the speech recognition model 31 to obtain the recognized speech text according to the multi-modal data;
[0087] 022: Invoke the natural language model 32 to obtain the replied speech text according to the recognized speech text;
[0088] 023: Invoke the speech synthesis model 33 to obtain the replied audio data according to the replied speech text;
[0089] 024: Invoke the lip-sync animation generation model 34 to stream-generate the facial drive data according to the replied audio data.
[0090] Specifically, the back-end server module 20 includes a core server 21, and the core server 21 can be a Server-Core server. The core server 21 is communicatively connected to the front-end interaction module 10, and the front-end interaction module 10 can send the multi-modal data to the core server 21.
[0091] The algorithm server module 30 includes a speech recognition model 31, a natural language model 32, a speech synthesis model 33, and a lip-sync animation generation model 34. The core server 21 is communicatively connected to the speech recognition model 31, the natural language model 32, the speech synthesis model 33, and the lip-sync animation generation model 34 respectively. The core server 21 can send the multi-modal data to the speech recognition model 31 and invoke the speech recognition model 31 to perform speech recognition on the multi-modal data to obtain the recognized speech text corresponding to the multi-modal data.
[0092] The speech recognition model 31 can send the recognized speech text to the core server 21, and the core server 21 then sends the recognized speech text to the natural language model 32, and calls the natural language model 32 to perform natural language processing (NLP) on the recognized speech text to obtain the response speech text corresponding to the speech recognition text.
[0093] The natural language model 32 can send the response speech text to the core server 21, and the core server 21 then sends the response speech text to the speech synthesis model 33, and calls the speech synthesis model 33 to perform speech synthesis on the response speech text to obtain the corresponding response audio data.
[0094] The speech synthesis model 33 can send the universally unique identifier (UUID) and address information of the response audio data to the core server 21. As Figure 2 shown, the algorithm server module 30 further includes a cache unit 35, and the speech synthesis model 33 can also cache the response audio data in the cache unit 35. The core server 21 can send the UUID and address information of the response audio data to the lip-sync animation generation model 34 and call the lip-sync animation generation model 34. The lip-sync animation generation model 34 reads the response audio data from the cache unit 35 based on the UUID and address information, and generates facial drive data in a streaming manner based on the response audio data. After generating the facial drive data, the lip-sync animation generation model 34 can send the facial drive data to the core server 21.
[0095] The specific process of generating the facial drive data is described in detail below.
[0096] Please refer to Figure 6 and Figure 7 , in some embodiments, calling the lip-sync animation generation model 34 to generate facial drive data (i.e., 024) in a streaming manner according to the response audio data includes:
[0097] 0241: In the process of the lip-sync animation generation model 34 obtaining the response audio data, non-global feature extraction is performed on the response audio data through a converter network to obtain Mel-frequency cepstral coefficient features;
[0098] 0242: Encoding the Mel-frequency cepstral coefficient features through a temporal convolutional network to obtain temporal convolutional features;
[0099] 0243: Performing global calculation on the temporal convolutional features through a converter network to obtain a first conversion feature;
[0100] 0244: Reconstruct the windowing and step size of the first conversion feature to obtain the second conversion feature;
[0101] 0245: Encode and decode the second conversion feature based on the multi-head attention mechanism to obtain the facial driving data.
[0102] Please refer to Figure 8 , in some embodiments, non-global feature extraction is performed on the response audio data through a transformer network to obtain Mel-Frequency Cepstral Coefficients features (i.e., 0241), including:
[0103] 02411: Perform non-global feature extraction on the response audio data through a transformer network to obtain Mel-Frequency Cepstral Coefficients;
[0104] 02412: Calculate the corresponding first-order difference coefficients and second-order difference coefficients based on the Mel-Frequency Cepstral Coefficients;
[0105] 02413: Determine the corresponding Mel-Frequency Cepstral Coefficients features based on the Mel-Frequency Cepstral Coefficients, first-order difference coefficients, and second-order difference coefficients.
[0106] Specifically, during the process of the lip-sync animation generation model 34 obtaining the response audio data from the cache unit 35, the lip-sync animation generation model 34 can process the obtained response audio data, thereby enabling the streaming generation of facial driving data. The lip-sync animation generation model 34 includes a transformer network, which can perform non-global feature extraction on the response audio data to obtain Mel-Frequency Cepstral Coefficients (MFCC). The calculation formula for Mel-Frequency Cepstral Coefficients is as follows:
[0107]
[0108] Among them, C(n) is the Mel-Frequency Cepstral Coefficient. n = 0, 1, 2,..., L; L is a coefficient of the Mel-Frequency Cepstral Coefficient feature dimension. In one example, L = 15. m is the Mel frequency. M is the number of filters. Based on the Mel-Frequency Cepstral Coefficients, the first-order difference coefficients and second-order difference coefficients of the Mel-Frequency Cepstral Coefficients can be obtained as shown in the following formula:
[0109] ΔC(n) = C(n) - C(n - 1)
[0110] Δ 2 C(n) = ΔC(n) - ΔC(n - 1)
[0111] Among them, ΔC(n) is the first-order difference coefficient of the Mel-Frequency Cepstral Coefficient. Δ 2C(n) is the second-order difference coefficient of the Mel-frequency cepstral coefficients. Based on the Mel-frequency cepstral coefficients, the first-order difference coefficients, and the second-order difference coefficients, a 45-dimensional Mel-frequency cepstral coefficient feature can be constructed.
[0112] After that, the Mel-frequency cepstral coefficient feature can be encoded by a Temporal Convolutional Network (TCN) to obtain a temporal convolutional feature. In the Temporal Convolutional Network, the feature encoder consists of seven blocks. Each block's temporal convolutional layer has 512 channels, with strides of (5, 2, 2, 2, 2, 2, 2) respectively, and kernel widths of (10, 3, 3, 3, 3, 2, 2) respectively. The output frequency of the feature encoder is 49 Hz, the stride between each sample is approximately 20 milliseconds, and the receptive field is 400 input samples, that is, 25 milliseconds of audio.
[0113] The temporal convolutional feature is input to the Transformer network. Through the Transformer network, global calculation can be performed on the temporal convolutional feature to obtain a first transformed feature. In the Transformer network, the kernel size of the convolutional layer used for modeling relative position embedding is 128 and is divided into 16 groups.
[0114] In the first transformed feature, the feature vector is as shown in the following formula:
[0115] F tcn ={N T ,H,K}
[0116] Among them, N T , H, and K all represent dimensions. N T is the dimension of the time length T, H is 16, and K is 512. To achieve the streaming generation of face-driven data, it is necessary to reconstruct the windowing and stride of F tcn to obtain a second transformed feature, as shown in the following formula:
[0117]
[0118] Among them, the windowing queue length is 3000 milliseconds, M is the number of frames corresponding to 3000 milliseconds (that is, 3 seconds). For example, when the frame rate is 25, M is 75.
[0119] After reconstruction, encoding and decoding the second transformed feature based on the multi-head attention mechanism can obtain the face-driven data. The expression of the multi-head attention mechanism is as follows:
[0120]
[0121] Among them, Q represents the Query matrix, K represents the Key matrix, and V represents the Value matrix. B is the bias term. d k is the scaling factor. CombineFigure 7 , the second transformation feature passes through three different linear transformation layers to obtain a query matrix, a key matrix, and a value matrix respectively. The query matrix, the key matrix, and the value matrix are divided into multiple heads (i.e., multiple subspaces), and each head has different linear transformation parameters. For each head, a dot product operation is performed once, and then an activation function (softmax) is used to obtain the attention weights. The weights are used to weight the value matrix, and the weighted sum is generated as the output of each head. The outputs of multiple heads are concatenated together to fuse the information of multiple heads, and the final multi-head attention output is obtained, which is the facial driving data. The facial driving data includes controlsRig and BlendShape based on the Metahuman plugin and the ARKit (Augmented Reality Kit).
[0122] In related technologies, the lip-sync algorithm needs to perform global feature extraction based on the complete audio, which requires a huge amount of computing power. In addition, when applied to digital human interaction, there is an obvious delay, which cannot meet the low-latency interaction requirements of digital humans, and in a multi-user scenario, the delay will be more obvious.
[0123] In the embodiments of the present application, during the process that the lip-sync animation generation model 34 obtains the response audio data from the cache unit 35, the lip-sync animation generation model 34 can perform non-global feature extraction on the obtained response audio data and generate the facial driving data in a streaming manner. In this way, the instant generation of the facial driving data can be realized, which can meet the low-latency interaction requirements of digital human interaction, and even can meet the digital human interaction requirements in a multi-user scenario.
[0124] Please refer to Figure 2 and Figure 9 , in some embodiments, based on the processing result, rendering is performed to obtain the digital human audio-visual stream data (i.e., 030), including:
[0125] 031: Perform three-dimensional rendering on the three-dimensional model based on the response audio data and the facial driving data to obtain the digital human audio-visual stream data;
[0126] At this time, text content display is performed based on the processing result, and audio-visual broadcast of the digital human is performed based on the digital human audio-visual stream data (i.e., 040), including:
[0127] 041: Perform text content display on the recognized speech text and the response speech text, and perform audio-visual broadcast of the digital human based on the digital human audio-visual stream data.
[0128] Specifically, after the core server 21 processes to obtain the recognized speech text and the response speech text by invoking the algorithm server module 30, it can output the recognized speech text and the response speech text to the front-end interaction module 10. The rendering and streaming module 40 includes an engine service unit 41, and the core server 21 is communicatively connected to the engine service unit 41. After the core server 21 processes to obtain the response audio data and the facial drive data by invoking the algorithm server module 30, it can output the response audio data and the facial drive data to the engine service unit 41.
[0129] The engine service unit 41 may be the Unreal Engine (UE). The engine service unit 41 can perform 3D rendering on the 3D digital human based on the response audio data and the facial drive data to obtain the digital human audio-visual stream.
[0130] The rendering and streaming module 40 is communicatively connected to the front-end interaction module 10 and can output the digital human audio-visual stream to the front-end interaction module 10. After receiving the recognized speech text, the front-end interaction module 10 can display the text content of the recognized speech text on the interaction interface 11 for the user to view.
[0131] After receiving the response speech text, the front-end interaction module 10 can display the text content of the response speech text on the interaction interface 11. When displaying the response speech text, the interaction interface 11 can also perform audio-visual broadcast on the received digital human audio-visual stream. In this way, the multi-modal output of the digital human is realized.
[0132] Please refer to Figure 2 and Figure 10 , in some embodiments, the back-end server module 20 further includes a web server 22. The rendering and streaming module 40 further includes a signaling server 42. Before obtaining the multi-modal data (i.e., 010), the distributed digital human interaction method further includes:
[0133] 050: Sending page data, the first address data of the signaling server 42, and the second address data of the core server 21 to the front-end interaction module 10 through the web server 22;
[0134] 060: Obtaining the interaction interface 11 based on the page data by the front-end interaction module 10, and implementing communication with the signaling server 42 based on the first address data, and implementing communication with the core server 21 based on the second address data.
[0135] Specifically, the user only needs a browser in the front-end interaction module 10 to use the digital human and related services. The back-end server module 20 further includes a web server 22. By entering an Internet Protocol (IP) address or a domain name address in the browser, the front-end interaction module 10 can send a Hypertext Transfer Protocol (HTTP) request to the web server 22. The web server 22 can be an HTTP World Wide Web (Web) server.
[0136] The web server 22 can send HyperText Markup Language (HTML) including page data to the front-end interaction module 10 according to the HTTP request. The browser decodes the page data to display the interaction interface 11. The interaction interface 11 is implemented by the HyperText Markup Language 5 (HTML5, H5) technology and is built based on front-end languages such as Node.js. In this way, the user access method is simplified, and the convenience and flexibility of digital human interaction are improved.
[0137] In addition, the rendering and streaming module 40 further includes a signaling server 42, and the signaling server 42 can be a Web Real-Time (WeRTC) signaling server. The HTML sent by the web server 22 also includes the first address data of the signaling server 42 and the second address data of the core server 21. The front-end interaction module 10 can communicate with the signaling server 42 through the first address data and communicate with the core server 21 through the second address data.
[0138] After establishing communication with the front-end interaction module 10, the signaling server 42 can transmit the digital human audio-video stream to the front-end interaction module 10 after the digital human audio-video stream is generated in the engine service unit 41. The signaling server 42 can also send the interaction feedback data corresponding to the driving instruction to the front-end interaction module 10 when the engine service unit 41 receives the driving instruction. In this way, the digital human audio-video stream is transmitted to the front-end interaction module 10 through web instant messaging, and the function separation from the front-end interaction module 10 can support cluster deployment.
[0139] After establishing communication with the front-end interaction module 10, the core server 21 can output the recognized speech text and the answering speech text to the front-end interaction module 10 after obtaining the recognized speech text and the answering speech text. The core server 21 can also send the corresponding response data to the front-end interaction module 10 after receiving the click event data sent by the front-end interaction module 10 for feedback display on the interaction interface 11.
[0140] Please refer to Figure 2 , an embodiment of the present application also provides a distributed digital human interaction system 100. The distributed digital human interaction system 100 includes a front-end interaction module 10, a back-end server module 20, an algorithm server module 30, and a rendering and streaming module 40. The front-end interaction module 10 is used to obtain multimodal data. The back-end server module 20 is used to call the algorithm server module 30 to perform streaming processing on the multimodal data based on the lip-sync animation generation model 34 to obtain a processing result corresponding to the multimodal data. The rendering and streaming module 40 is used to perform rendering based on the processing result to obtain digital human audio-visual stream data. The front-end interaction module 10 is further used to display text content based on the processing result and perform audio-visual broadcast of the digital human based on the digital human audio-visual stream data.
[0141] In the distributed digital human interaction system 100 of the embodiment of the present application, the multimodal data is subjected to streaming processing based on the lip-sync animation generation model 34 to obtain a corresponding processing result; the digital human audio-visual stream data is obtained by rendering based on the processing result; the processing result is used for text content display, and the digital human audio-visual stream data is used for audio-visual broadcast of the digital human. In this way, the processing result of the multimodal data can be obtained immediately, and then the text content display and the audio-visual broadcast of the digital human can be performed faster, meeting the low-latency interaction requirements of digital human interaction.
[0142] In the related art, the digital human interaction system relies on the overall architecture and needs to be developed as a whole. When deploying algorithms and systems, they must also be deployed as a whole, unable to be compatible with cross-terminal and cross-platform issues, and the computing power resource allocation is unbalanced.
[0143] In the embodiment of the present application, the distributed digital human interaction system 100 adopts a distributed architecture with functional modularization. Multiple modules can be developed independently in parallel, and the algorithms and system deployment can also be carried out separately. Multiple modules can be decoupled in development, deployment, and application, thereby realizing the reasonable allocation and use of hardware computing power resources and meeting the digital human interaction requirements in multi-user scenarios.
[0144] For example, front-end interaction requires the computing power of a Central Processing Unit (CPU); deep learning models require the computing power of the Compute Unified Device Architecture (CUDA) cores of a Graphics Processing Unit (GPU); digital human rendering requires the 3D graphics rendering ability of a Graphic Process Unit (GPU).
[0145] The above three capabilities are respectively deployed in different modules, that is, in the front-end interaction module 10, the algorithm server module 30, and the rendering and streaming module 40. Compared with being deployed in the same device, the computing power resources are effectively balanced. In addition, the CUDA resources required by the algorithm server module 30 and the 3D graphics rendering capabilities required by the rendering and streaming module 40 can be independently and elastically expanded according to requirements, realizing high-concurrency bearing capacity.
[0146] The distributed digital human interaction system 100 can also be compatible with cloud service deployment solutions, edge-side deployment solutions, and terminal integrated deployment solutions, so as to meet the deployment requirements of various digital human applications and provide support for the diversification of business forms. In this way, it has good compatibility with problems such as cross-terminal, cross-platform, and unbalanced computing power resources in actual scenarios.
[0147] In some embodiments, the processing results include recognized speech text, response speech text, response audio data, and facial drive data. The back-end server module 20 includes a core server 21. The algorithm server module 30 includes a speech recognition model 31, a natural language model 32, a speech synthesis model 33, and a lip-sync animation generation model 34. The core server 21 is used to call the speech recognition model 31, the natural language model 32, the speech synthesis model 33, and the lip-sync animation generation model 34. The speech recognition model 31 is used to obtain recognized speech text based on multi-modal data. The natural language model 32 is used to obtain response speech text based on the recognized speech text. The speech synthesis model 33 is used to obtain response audio data based on the response speech text. The lip-sync animation generation model 34 is used to stream generate facial drive data based on the response audio data.
[0148] In some embodiments, the front-end interaction module 10 includes an interaction interface 11. The rendering and streaming module 40 includes an engine service unit 41. The core server 21 is used to output the recognized speech text and the response speech text to the front-end interaction module 10, and output the response audio data and the facial drive data to the engine service unit 41. The engine service unit 41 is used to perform 3D rendering on the 3D model based on the response audio data and the facial drive data to obtain digital human audio-visual stream data. The interaction interface 11 is used to display the text content of the recognized speech text and the response speech text, and perform audio-visual broadcast of the digital human based on the digital human audio-visual stream data.
[0149] In some embodiments, the backend server module 20 further includes a web server 22. The rendering and streaming module 40 further includes a signaling server 42. The web server 22 is configured to send page data, the first address information of the signaling server 42, and the second address information of the core server 21 to the front-end interaction module 10. The front-end interaction module 10 is configured to obtain an interaction interface 11 based on the page data, implement communication with the signaling server 42 based on the first address data, and implement communication with the core server 21 based on the second address data.
[0150] It should be noted that the explanations of the distributed digital human interaction method in the foregoing embodiments are equally applicable to the distributed digital human interaction system 100 in the embodiments of the present application, and will not be elaborated herein.
[0151] In some embodiments, the core server 21 is further configured to perform access management and resource management for multi-user access.
[0152] Among them, access management means that after the user obtains the interaction interface 11 through the web server 22, the core server 21 can manage user login and information verification. Resource management means managing the concurrent resources of the algorithm server module 30 and the rendering and streaming module 40. The management includes: hardware resource management, memory management, and sub-service process management. The scenarios of resource management include server clusters, edge computing devices, multi-node computing resources, single offline host devices, etc.
[0153] The following will Figure 2 describe in detail the complete communication process inside the distributed digital human interaction system 100 according to the embodiments of the present application.
[0154] a represents the process of the front-end interaction module 10 establishing a connection and communicating with other modules, where:
[0155] a1 represents that the user inputs an IP address or a domain name address to access the web server 22;
[0156] a2 represents that the web server 22 sends HTML to the front-end interaction module 10;
[0157] a3 represents that the user establishes a communication connection with the signaling server 42 according to the first address data;
[0158] a4 represents that the front-end interaction module 10 establishes a normal data connection and an audio / video stream connection with the signaling server 42.
[0159] b represents the voice interaction process of the digital human, where:
[0160] b1-1 represents that the front-end interaction module 10 sends audio stream data to the core server 21;
[0161] b1-2 means that the core server 21 sends the audio stream data to the speech recognition model 31;
[0162] b1-3 means that the speech recognition model 31 sends the recognized speech text to the core server 21;
[0163] b1-4 means that the core server 21 sends the recognized speech text to the front-end interaction module 10;
[0164] b2-1 means that the core server 21 sends the recognized speech text to the natural language model 32;
[0165] b2-2 means that the natural language model 32 sends the answering speech text to the core server 21;
[0166] b2-3 means that the core server 21 sends the answering speech text to the front-end interaction module 10;
[0167] b3-1 means that the core server 21 sends the answering speech text to the speech synthesis model 33;
[0168] b3-2 means that the speech synthesis model 33 sends the UUID and address information of the answering audio data to the core server 21;
[0169] b3-3 means that the speech synthesis model 33 stores the answering audio data in the cache unit 35;
[0170] b4-1 means that the cache unit 35 sends the answering audio data to the core server 21;
[0171] b4-2 means that the core server 21 sends the answering audio data to the rendering and streaming module 40;
[0172] b5-1 means that the core server 21 sends the UUID and address information of the answering audio data to the lip-sync animation generation model 34;
[0173] b5-2 means that the lip-sync animation generation model 34 reads the answering audio data from the cache unit 35;
[0174] b5-3 means that the lip-sync animation generation model 34 sends the facial drive data to the core server 21;
[0175] b5-4 means that the core server 21 sends the facial drive data to the rendering and streaming module 40;
[0176] b6 means that the rendering and streaming module 40 sends the digital human audio-video stream to the front-end interaction module 10 through a4.
[0177] c represents the event interaction process of the digital human, where:
[0178] c-1 means that the front-end interaction module 10 sends click event data to the core server 21;
[0179] c-2 means that the core server 21 sends digital human driving instructions to the rendering and streaming module 40;
[0180] c-3 means that the core server 21 sends response data to the front-end interaction module 10;
[0181] c-4 means that the rendering and streaming module 40 sends interaction feedback data to the front-end interaction module 10 through a4.
[0182] As can be seen from Figure 2 it, the communication inside the distributed digital human interaction system 100 includes web instant messaging, ordinary data communication, and long-connection data transmission. In this way, an efficient data communication and transmission method is constructed in the communication process, and the overall interaction delay is small. For example, the overall interaction delay in the edge-side deployment scheme and the terminal integrated deployment scheme is less than 1.2 seconds.
[0183] Please refer to Figure 11 , this embodiment of the present application also provides an electronic device 200. The electronic device 200 includes one or more processors 210 and a memory 220. When the computer program stored in the memory 220 is executed by the processor 210, the distributed digital human interaction method of any of the above embodiments is implemented.
[0184] For example, when the program is executed by the processor 210, the following distributed digital human interaction method is implemented:
[0185] 010: Obtain multimodal data;
[0186] 020: Perform streaming processing on the multimodal data based on the lip-sync animation generation model 34 to obtain a processing result corresponding to the multimodal data;
[0187] 030: Perform rendering based on the processing result to obtain digital human audio-visual stream data;
[0188] 040: Display text content based on the processing result, and perform audio-visual broadcast of the digital human based on the digital human audio-visual stream data.
[0189] For another example, when the program is executed by the processor 210, the following distributed digital human interaction method is implemented:
[0190] 021: Call the speech recognition model 31 to obtain the recognized speech text according to the multimodal data;
[0191] 022: Call the natural language model 32 to obtain the answering speech text according to the recognized speech text;
[0192] 023: Invoke the speech synthesis model 33 to obtain response audio data based on the response speech text;
[0193] 024: Invoke the lip-sync animation generation model 34 to stream generate facial drive data based on the response audio data.
[0194] It should be noted that the explanations of the distributed digital human interaction method and the distributed digital human interaction system 100 in the foregoing embodiments are equally applicable to the electronic device 200 in the embodiments of the present application, and will not be elaborated herein.
[0195] Please refer to Figure 12 , the embodiments of the present application further provide a computer-readable storage medium 300, on which a computer program 310 is stored. When the program is executed by the processor 320, the distributed digital human interaction method of any of the foregoing embodiments is implemented.
[0196] For example, when the program is executed by the processor 320, the following distributed digital human interaction method is implemented:
[0197] 010: Obtain multimodal data;
[0198] 020: Perform streaming processing on the multimodal data based on the lip-sync animation generation model 34 to obtain a processing result corresponding to the multimodal data;
[0199] 030: Perform rendering based on the processing result to obtain digital human audio-visual stream data;
[0200] 040: Display the text content based on the processing result, and perform audio-visual broadcast of the digital human based on the digital human audio-visual stream data.
[0201] For another example, when the program is executed by the processor 320, the following distributed digital human interaction method is implemented:
[0202] 021: Invoke the speech recognition model 31 to obtain the recognized speech text based on the multimodal data;
[0203] 022: Invoke the natural language model 32 to obtain the response speech text based on the recognized speech text;
[0204] 023: Invoke the speech synthesis model 33 to obtain response audio data based on the response speech text;
[0205] 024: Invoke the lip-sync animation generation model 34 to stream generate facial drive data based on the response audio data.
[0206] It should be noted that the explanations of the distributed digital human interaction method and the distributed digital human interaction system 100 in the foregoing embodiments are equally applicable to the computer-readable storage medium 300 of the embodiments of the present application, and will not be elaborated herein.
[0207] In summary, in the distributed digital human interaction method, the distributed digital human interaction system 100, the electronic device 200, and the computer-readable storage medium 300 of the embodiments of the present application, the multi-modal data is stream-processed based on the lip-sync animation generation model 34 to obtain corresponding processing results; the digital human audio-visual stream data is rendered based on the processing results; the processing results are used for text content display, and the digital human audio-visual stream data is used for the audio-visual broadcast of the digital human. In this way, the processing results of the multi-modal data can be obtained immediately, and then the text content display and the digital human audio-visual broadcast can be performed faster, meeting the low-latency interaction requirements of digital human interaction.
[0208] In the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples.
[0209] Any process or method description in the flowchart or described in other ways herein can be understood as representing a module, segment, or part of the code including one or more executable instructions for implementing a specific logical function or process, and the scope of the preferred embodiments of the present application includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in the reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the art of the embodiments of the present application.
[0210] The logic and / or steps represented in the flowchart or otherwise described herein can, for example, be considered as a definable sequence of executable instructions for implementing logical functions, and can be embodied specifically in any computer-readable storage medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device. For the purposes of this specification, a computer-readable storage medium can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with the instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of the computer-readable storage medium include the following: an electrical connection portion having one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable storage medium can even be paper or other suitable medium on which the program can be printed, as the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpretation, or otherwise processing as appropriate, and then storing it in a computer memory.
[0211] It should be understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGA), field programmable gate arrays (FPGA), etc.
[0212] Those of ordinary skill in the art can understand that all or part of the steps carried out in implementing the above-described embodiment methods can be completed by instructing relevant hardware through a program. The said program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment. In addition, in each of the various embodiments of the present application, each functional unit can be integrated in a processing module, or each unit can exist physically alone, or two or more units can be integrated in one module. The above-mentioned integrated module can be implemented in the form of hardware or in the form of a software functional module. If the above-mentioned integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. The storage medium mentioned above can be a read-only memory, a magnetic disk, an optical disc, or the like.
[0213] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application. The scope of the present application is defined by the claims and their equivalents.
Claims
1. A distributed digital human interaction method, characterized in that: include: Acquire multimodal data; Performing streaming processing on the multimodal data based on a lip sync animation generation model to obtain processing results corresponding to the multimodal data; Rendering is performed based on the processing result to obtain digital human audio and video stream data; The text content is displayed based on the processing result, and the audio and video broadcast of the digital human is performed based on the digital human audio and video stream data.
2. The distributed digital human interaction method according to claim 1, characterized in that: The multimodal data includes any one or more of audio stream data, video stream data, and click event data.
3. The distributed digital human interaction method according to claim 1 or 2, characterized in that: The processing result includes recognized speech text, answer speech text, answer audio data and facial drive data, and the multimodal data is stream-processed based on the lip sync animation generation model to obtain the processing result corresponding to the multimodal data, including: Calling a speech recognition model to obtain the recognized speech text according to the multimodal data; Calling a natural language model to obtain the answer voice text according to the recognized voice text; Calling a speech synthesis model to obtain the answer audio data according to the answer speech text; The lip sync animation generation model is called to generate the facial driving data according to the answer audio data stream.
4. The distributed digital human interaction method according to claim 3, characterized in that: The calling of the lip sync animation generation model to stream-generate the facial drive data according to the answer audio data comprises: In the process of the lip sync animation generation model acquiring the answer audio data, performing non-global feature extraction on the answer audio data through a converter network to obtain Mel-frequency cepstral coefficient features; Encoding the Mel-frequency cepstral coefficient feature through a temporal convolutional network to obtain a temporal convolution feature; Performing global calculation on the temporal convolution feature through the converter network to obtain a first conversion feature; Performing window division and step-size reconstruction on the first conversion feature to obtain a second conversion feature; The second conversion feature is encoded and decoded based on a multi-head attention mechanism to obtain the facial driving data.
5. The distributed digital human interaction method according to claim 4, characterized in that: The non-global feature extraction of the answer audio data by the converter network to obtain Mel frequency cepstral coefficient features includes: Performing non-global feature extraction on the answer audio data through the converter network to obtain Mel-frequency cepstrum coefficients; Calculating corresponding first-order difference coefficients and second-order difference coefficients based on the Mel-frequency cepstrum coefficients; The corresponding Mel-frequency cepstral coefficient features are determined based on the Mel-frequency cepstral coefficients, the first-order difference coefficients and the second-order difference coefficients.
6. The distributed digital human interaction method according to claim 3, characterized in that: The rendering based on the processing result to obtain the digital human audio and video stream data includes: Performing three-dimensional rendering on the three-dimensional model based on the answer audio data and the facial drive data to obtain the digital human audio and video stream data; The displaying of text content based on the processing result and the broadcasting of the audio and video of the digital human based on the audio and video stream data of the digital human include: The text content of the recognized voice book and the answer voice text is displayed, and the audio and video of the digital human is broadcast based on the audio and video stream data of the digital human.
7. A distributed digital human interaction system, characterized in that: It includes front-end interaction module, back-end server module, algorithm server module and rendering and streaming module; The front-end interaction module is used to obtain multimodal data; The backend server module is used to call the algorithm server module to perform streaming processing on the multimodal data based on the lip sync animation generation model to obtain processing results corresponding to the multimodal data; The rendering and streaming module is used to render based on the processing result to obtain the digital human audio and video stream data; The front-end interaction module is also used to display text content based on the processing result, and to broadcast the audio and video of the digital human based on the audio and video stream data of the digital human.
8. The distributed digital human interaction system according to claim 7, characterized in that: The processing result includes recognition speech text, answer speech text, answer audio data and facial drive data, the back-end server module includes a core server, the algorithm server module includes a speech recognition model, a natural language model, a speech synthesis model and a lip sync animation generation model, and the core server is used to call the speech recognition model, the natural language model, the speech synthesis model and the lip sync animation generation model; The speech recognition model is used to obtain the recognized speech text according to the multimodal data; The natural language model is used to obtain the answer voice text according to the recognized voice text; The speech synthesis model is used to obtain the answer audio data according to the answer speech text; The lip sync animation generation model is used to stream-generate the facial drive data according to the answer audio data.
9. The distributed digital human interaction system according to claim 8, characterized in that: The front-end interaction module includes an interaction interface, and the rendering and streaming module includes an engine service unit; The core server is used to output the recognized voice text and the answer voice text to the front-end interaction module, and output the answer audio data and the facial drive data to the engine service unit; The engine service unit is used to perform three-dimensional rendering on the three-dimensional model based on the answer audio data and the facial drive data to obtain the digital human audio and video stream data; The interactive interface is used to display the text content of the recognized voice book and the answer voice text, and to broadcast the audio and video of the digital human based on the audio and video stream data of the digital human.
10. The distributed digital human interaction system according to claim 9, characterized in that: The backend server module also includes a web server, and the rendering and streaming module also includes a signaling server; The web page server is used to send page data, the first address information of the signaling server, and the second address information of the core server to the front-end interaction module; The front-end interaction module is used to obtain the interaction interface based on the page data, and to achieve communication with the signaling server based on the first address data, and to achieve communication with the core server based on the second address data.