Digital human audio and video stream data acquisition method and device

By employing pre-rendering technology and streaming communication in the digital human system, the high latency problem of the digital human system was solved, resulting in faster user response and a better user experience.

CN121924327APending Publication Date: 2026-04-24IND BANK CO
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
IND BANK CO
Filing Date
2025-12-29
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing digital human systems have high end-to-end latency, which leads to a decline in user experience, especially when the user information response latency exceeds 2 seconds, resulting in serious user churn.

Method used

Initialization operations are achieved through interaction between the digital human front-end and back-end. Pre-rendering technology is used to cache offline broadcast audio and video data. Combined with bidirectional streaming communication and real-time streaming data transmission, the latency of speech recognition and audio and video streaming data is reduced.

Benefits of technology

It effectively reduces the latency of digital human systems during initialization, speech recognition, speech synthesis, and audio/video streaming, improving user response speed and experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121924327A_ABST
    Figure CN121924327A_ABST
Patent Text Reader

Abstract

The invention provides a digital human audio and video stream data acquisition method and device, and relates to the technical field of data processing. The method comprises the following steps: performing initialization operation on the digital human system based on interaction between the digital human front end and the digital human rear end; based on interaction among the digital human front end, the digital human central control and the digital human rear end, obtaining voice recognition data and replying synthetic voice data of voice input of a user; and obtaining digital human audio and video stream data based on interaction between the digital human central controller and the digital human rear end, and sending the digital human audio and video stream data to an audio and video platform. According to the digital human audio and video stream data acquisition method and device provided by the embodiment of the invention, the time delay in the process of acquiring the digital human audio and video stream data can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, specifically to a method and apparatus for acquiring digital human audio and video stream data. Background Technology

[0002] A complete digital human system comprises multiple stages, including automatic speech recognition, image recognition, natural language processing, semantic understanding, speech synthesis, digital human driving, and image rendering. Because each stage involves AI algorithms that consume significant computing power, and these stages have a clear sequential relationship, data processing can only be performed serially, resulting in high end-to-end latency. According to surveys, when the latency from sending a message to receiving a response exceeds 2 seconds, the user experience deteriorates, leading to significant user churn. Summary of the Invention

[0003] To address the problems in the prior art, embodiments of the present invention provide a method and apparatus for acquiring digital human audio and video stream data, which can at least partially solve the problems existing in the prior art.

[0004] On one hand, this invention proposes a method for acquiring audio and video stream data of a digital human. This method is applied to a digital human system, which includes a digital human front-end, a digital human central control unit, and a digital human back-end. The method for acquiring audio and video stream data of a digital human includes:

[0005] The initialization operation of the digital human system is achieved based on the interaction between the digital human front end and the digital human back end.

[0006] Based on the interaction between the digital human front end, the digital human central control unit, and the digital human back end, synthesized voice data is used to acquire speech recognition data and respond to user voice input.

[0007] Based on the interaction between the digital human central control unit and the digital human backend, the digital human audio and video stream data is acquired and sent to the audio and video platform.

[0008] The initialization operation of the digital human system based on the interaction between the digital human front end and the digital human back end includes:

[0009] In response to a user accessing the digital human system, the digital human front end plays pre-stored offline audio and video data of the digital human to the user, so that the digital human front end can obtain user feedback information generated by the user in response to the offline audio and video data of the digital human.

[0010] The digital human front end sends the user feedback information to the digital human back end;

[0011] The digital human backend performs initialization operations on the digital human system based on the user feedback information.

[0012] The user feedback information includes user audiovisual sensory feedback information; correspondingly, the digital human backend performs initialization operations on the digital human system based on the user feedback information, including:

[0013] The digital human backend identifies the user's audiovisual sensory feedback information to obtain user identification information and user intent information, and performs initialization operations on the digital human system based on the user identification information and the user intent information.

[0014] The method of acquiring speech recognition data and responding to user voice input based on the interaction between the digital human front-end, the digital human central control unit, and the digital human back-end includes:

[0015] The digital human front end sends the real-time collected user audio data to the digital human central control unit;

[0016] The digital human central control unit sends the user's audio data to the digital human backend in real time;

[0017] The digital human backend recognizes the user's audio data to obtain speech recognition data, and then sends the speech recognition data to the digital human central control unit.

[0018] The method of acquiring speech recognition data and responding to user voice input based on the interaction between the digital human front-end, the digital human central control unit, and the digital human back-end includes:

[0019] The digital human central control unit sends the voice recognition data to the digital human front end;

[0020] After the user inputs voice data, the digital human central control unit sends the text data corresponding to all user audio data to the digital human backend.

[0021] The digital human backend generates synthesized speech data to respond to the user's voice input based on the text data corresponding to all user audio data, and sends the synthesized speech data to the digital human central control in the form of data fragments.

[0022] The step of acquiring digital human audio and video stream data based on the interaction between the digital human central control unit and the digital human backend includes:

[0023] The digital human central control unit sends the synthesized voice data to the digital human backend;

[0024] The digital human backend drives and renders the digital human based on the synthesized speech data to obtain digital human audio and video stream data.

[0025] On one hand, this invention proposes a digital human audio and video stream data acquisition device, which is applied to a digital human system, the digital human system including a digital human front end, a digital human central control unit, and a digital human back end; the digital human audio and video stream data acquisition device includes:

[0026] An initialization unit is used to perform initialization operations on the digital human system based on the interaction between the digital human front end and the digital human back end.

[0027] The first acquisition unit is used to acquire speech recognition data and synthesized speech data in response to user voice input based on the interaction between the digital human front end, the digital human central control unit and the digital human back end;

[0028] The second acquisition unit is used to acquire digital human audio and video stream data based on the interaction between the digital human central control unit and the digital human backend, and send the digital human audio and video stream data to the audio and video platform.

[0029] In another aspect, embodiments of the present invention provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the following method:

[0030] The initialization operation of the digital human system is achieved based on the interaction between the digital human front end and the digital human back end.

[0031] Based on the interaction between the digital human front end, the digital human central control unit, and the digital human back end, synthesized voice data is used to acquire speech recognition data and respond to user voice input.

[0032] Based on the interaction between the digital human central control unit and the digital human backend, the digital human audio and video stream data is acquired and sent to the audio and video platform.

[0033] This invention provides a computer-readable storage medium, comprising:

[0034] The computer-readable storage medium stores a computer program that, when executed by a processor, implements the following method:

[0035] The initialization operation of the digital human system is achieved based on the interaction between the digital human front end and the digital human back end.

[0036] Based on the interaction between the digital human front end, the digital human central control unit, and the digital human back end, synthesized voice data is used to acquire speech recognition data and respond to user voice input.

[0037] Based on the interaction between the digital human central control unit and the digital human backend, the digital human audio and video stream data is acquired and sent to the audio and video platform.

[0038] This invention also provides a computer program product, which includes a computer program that, when executed by a processor, implements the following method:

[0039] The initialization operation of the digital human system is achieved based on the interaction between the digital human front end and the digital human back end.

[0040] Based on the interaction between the digital human front end, the digital human central control unit, and the digital human back end, synthesized voice data is used to acquire speech recognition data and respond to user voice input.

[0041] Based on the interaction between the digital human central control unit and the digital human backend, the digital human audio and video stream data is acquired and sent to the audio and video platform.

[0042] The digital human audio and video stream data acquisition method and apparatus provided in this invention initializes the digital human system based on the interaction between the digital human front end and the digital human back end; acquires speech recognition data and synthesized speech data in response to user voice input based on the interaction between the digital human front end, the digital human central control unit, and the digital human back end; and acquires digital human audio and video stream data and sends the digital human audio and video stream data to the audio and video platform based on the interaction between the digital human central control unit and the digital human back end, thereby reducing latency in the process of acquiring digital human audio and video stream data. Attached Figure Description

[0043] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings:

[0044] Figure 1 This is a flowchart illustrating a method for acquiring digital human audio and video stream data according to an embodiment of the present invention.

[0045] Figure 2 This is a schematic diagram of the structure of the digital human system provided in an embodiment of the present invention.

[0046] Figure 3 This is a schematic diagram of the modular structure of the digital human backend function provided in an embodiment of the present invention.

[0047] Figure 4 This is a schematic diagram of the structure of a digital human audio and video stream data acquisition device provided in an embodiment of the present invention.

[0048] Figure 5 This is a schematic diagram of the physical structure of a computer device provided in an embodiment of the present invention. Detailed Implementation

[0049] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings. Here, the illustrative embodiments and descriptions of the present invention are used to explain the present invention, but are not intended to limit the present invention. It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of this application can be arbitrarily combined with each other.

[0050] Figure 1 This is a flowchart illustrating a digital human audio / video stream data acquisition method according to an embodiment of the present invention, as shown below. Figure 1 As shown, the digital human audio and video stream data acquisition method provided in this embodiment of the invention is applied to a digital human system, which includes a digital human front-end, a digital human central control unit, and a digital human back-end; the digital human audio and video stream data acquisition method includes:

[0051] Step S1: Initialize the digital human system based on the interaction between the digital human front end and the digital human back end.

[0052] Step S2: Based on the interaction between the digital human front end, the digital human central control unit, and the digital human back end, acquire speech recognition data and synthesized speech data to respond to user voice input.

[0053] Step S3: Based on the interaction between the digital human central control unit and the digital human backend, acquire digital human audio and video stream data, and send the digital human audio and video stream data to the audio and video platform.

[0054] In step S1 above, the device initializes the digital human system based on the interaction between the digital human front end and the digital human back end. The device can be a computer device executing this method. The acquisition, storage, use, and processing of data in this application's technical solution all comply with relevant regulations. The digital human system and audio / video platform to which the method is applied are, for example... Figure 2 As shown.

[0055] Before each broadcast, the interactive digital human pre-caches general offline broadcast audio and video data in its backend. When the user enters the application, this offline audio and video data is played, enabling the digital human to proactively initiate communication. Through this proactive communication, the system gathers the specific services the user requires. At this point, the digital human's backend initializes accordingly, reducing the latency of the system's response to the user. Details are as follows:

[0056] For specific scenarios, the pre-rendered digital human backend is used to create offline audio and video data for digital human broadcasting.

[0057] Once a user enters the digital human system, the system proactively acquires and broadcasts offline audio and video data to attract the user's attention and collect their intent. Specifically, it can identify the user's visual and auditory sensory feedback to obtain information such as facial expressions and emotions while watching and listening to the offline audio and video data. This information can be used to determine the user's intent. For example, if the user's facial expressions and emotions change when the offline audio and video data plays information about high-yield financial products, this change can be identified to determine if the user is interested in such products.

[0058] In addition, user facial information can be obtained through facial recognition and compared with pre-stored user identity information to determine user identification information.

[0059] By using user intent information, the digital person serving the user can be identified, and initialization operations such as generating and rendering the digital person model can be performed. The generated digital person is then associated with the user's identification information. This association establishes a binding between the user and the digital person reflecting the user's preferences, facilitating subsequent business operations.

[0060] It should be noted that the initialization operation can be performed when the user is watching and listening to the digital human's offline audio and video broadcast data. This ensures that the digital human system initialization operation is completed when the offline audio and video broadcast data is finished, saving the user the time of waiting for the digital human system to initialize.

[0061] In step S2 above, the device acquires speech recognition data and synthesized speech data in response to user voice input based on the interaction between the digital human front end, the digital human central control unit and the digital human back end.

[0062] It employs bidirectional streaming communication to receive and process user-generated audio data, sending the real-time collected audio data to an automatic speech recognition service and displaying and correcting the recognition results to the user in real time. Specific details are as follows:

[0063] The digital human front end can send binary data of the collected user audio data to the digital human central control every 20 milliseconds.

[0064] The digital human central control unit sends the received user audio data to the voice recognition service on the back end of the digital human in real time.

[0065] The digital human central control unit activates a listening service to receive real-time voice recognition data from the voice recognition service on the digital human's backend.

[0066] The digital human backend sends the voice recognition data back to the digital human frontend via the digital human central control unit, allowing users to see the text data corresponding to their voice data.

[0067] After the user finishes inputting their voice, the digital human central control unit can immediately send the text data corresponding to all the user's audio data to the natural language understanding module, thereby reducing the latency in the speech recognition stage.

[0068] After obtaining the text to be replied to the user, the digital human backend generates synthesized speech data in real time and returns the synthesized speech data to the user in real time using streaming communication. Details are as follows:

[0069] The digital human central control unit sends the full text data corresponding to the speech data to be synthesized to the speech synthesis service at once.

[0070] The digital human central control unit activates a monitoring service to receive audio data from the digital human's backend speech synthesis service in real time.

[0071] The digital human backend speech synthesis service generates synthesized speech data to respond to the user's voice input, and uses data slicing to return the synthesized speech data to the digital human central control unit in real time.

[0072] In step S3 above, the device acquires digital human audio and video stream data based on the interaction between the digital human central control unit and the digital human backend, and sends the digital human audio and video stream data to the audio and video platform.

[0073] The digital human central control unit sends synthesized speech data to the digital human backend. The digital human central control unit then sends the synthesized speech data to the digital human driving module and the digital human rendering module, thereby realizing the driving and rendering of the digital human. Real-time streaming data transmission reduces the latency of the digital human system in the speech synthesis, image driving, and image rendering stages.

[0074] After rendering the digital human's audio and video stream data to be displayed to the user in real time, the digital human uses an RTC-based streaming media service to transmit the audio and video stream data, thereby reducing the transmission latency of the media stream data. Details are as follows:

[0075] The digital human backend pushes the digital human's audio and video stream data to the audio and video platform in real time.

[0076] The audio and video platform will push the acquired digital human audio and video stream data to the digital human front end in real time.

[0077] After acquiring the audio and video stream data of the digital human, the front end decodes it in real time and plays the audio and video information. By transmitting the audio and video data to the front end in real time, the latency of the audio and video data transmission stage is reduced.

[0078] The modular implementation of the digital human backend functions in this invention is as follows: Figure 3 As shown.

[0079] The technical problem solved by the digital human audio and video stream data acquisition method provided in this embodiment of the invention is as follows:

[0080] 1) By using pre-rendering technology to generate audio and video data that can be cached offline, the problem of the inability to respond to user interactions in a timely manner during the initialization phase of digital humans is solved, and the system's response latency to users during the initialization phase is reduced.

[0081] 2) By adopting a two-way streaming communication method to receive and process user-generated audio and video data, the latency of the digital human system in the automatic speech recognition stage is reduced.

[0082] 3) By using streaming communication to return audio data to the digital human front end in real time, the latency of the digital human system in the speech synthesis stage is reduced.

[0083] 4) By using real-time communication technology to push the digital human audio and video stream data generated by the system, the latency of audio and video data transmission is reduced.

[0084] The digital human audio and video stream data acquisition method provided in this embodiment of the invention initializes the digital human system based on the interaction between the digital human front end and the digital human back end; acquires speech recognition data and synthesized speech data in response to user voice input based on the interaction between the digital human front end, the digital human central control unit, and the digital human back end; and acquires digital human audio and video stream data and sends the digital human audio and video stream data to the audio and video platform based on the interaction between the digital human central control unit and the digital human back end, thereby reducing the latency in the process of acquiring digital human audio and video stream data.

[0085] In the above optional embodiments, the initialization operation of the digital human system based on the interaction between the digital human front end and the digital human back end includes:

[0086] In response to a user accessing the digital human system, the digital human front end plays pre-stored offline audio and video data of the digital human to the user, so that the digital human front end can obtain user feedback information generated by the user in response to the offline audio and video data of the digital human; the above embodiments can be referred to for description, and will not be repeated here.

[0087] The digital human front end sends the user feedback information to the digital human back end; this can be described with reference to the above embodiments, and will not be repeated here.

[0088] The digital human backend performs initialization operations on the digital human system based on the user feedback information. This can be referred to the above embodiments for further explanation and will not be repeated here.

[0089] In the above optional embodiments, the user feedback information includes user audiovisual sensory feedback information; correspondingly, the digital human backend performs initialization operations on the digital human system based on the user feedback information, including:

[0090] The digital human backend identifies the user's audiovisual sensory feedback information to obtain user identification information and user intent information, and performs initialization operations on the digital human system based on the user identification information and user intent information. This can be referred to the above embodiments for further explanation and will not be repeated here.

[0091] In the above optional embodiments, the step of acquiring speech recognition data and responding to user voice input based on the interaction between the digital human front end, the digital human central control unit, and the digital human back end includes:

[0092] The digital human front end sends the real-time collected user audio data to the digital human central control unit; this can be referred to the above embodiment for explanation, and will not be repeated here.

[0093] The digital human central control unit sends the user's audio data to the digital human backend in real time; this can be referred to the above embodiment for explanation, and will not be repeated here.

[0094] The digital human backend recognizes the user's audio data to obtain speech recognition data, and then sends the speech recognition data to the digital human central control unit. This can be referred to the above embodiment for further explanation, and will not be repeated here.

[0095] In the above optional embodiments, the step of acquiring speech recognition data and responding to user voice input based on the interaction between the digital human front end, the digital human central control unit, and the digital human back end includes:

[0096] The central control unit of the digital human sends the voice recognition data to the front end of the digital human; this can be referred to the above embodiment for explanation, and will not be repeated here.

[0097] After the user inputs voice data, the digital human central control sends the text data corresponding to all user audio data to the digital human backend; this can be referred to the above embodiment for explanation, and will not be repeated here.

[0098] The digital human backend generates synthesized speech data to respond to user voice input based on the text data corresponding to all user audio data, and sends the synthesized speech data to the digital human central control unit in a data fragmentation manner. This can be referred to the above embodiment for further explanation and will not be repeated here.

[0099] In the above optional embodiments, the step of acquiring digital human audio and video stream data based on the interaction between the digital human central control unit and the digital human backend includes:

[0100] The digital human central control unit sends the synthesized voice data to the digital human backend; this can be referred to the above embodiment for explanation, and will not be repeated here.

[0101] The digital human backend drives and renders the digital human based on the synthesized speech data to obtain digital human audio and video stream data. This can be referred to the above embodiments for further explanation and will not be repeated here.

[0102] Figure 4 This is a schematic diagram of the structure of a digital human audio and video stream data acquisition device according to an embodiment of the present invention, as shown below. Figure 4 As shown, the digital human audio and video stream data acquisition device provided in this embodiment of the invention is applied to a digital human system, which includes a digital human front-end, a digital human central control unit, and a digital human back-end; the digital human audio and video stream data acquisition device includes an initialization unit 401, a first acquisition unit 402, and a second acquisition unit 403, wherein:

[0103] The initialization unit 401 is used to perform initialization operations on the digital human system based on the interaction between the digital human front end and the digital human back end; the first acquisition unit 402 is used to acquire speech recognition data and synthesized speech data in response to user voice input based on the interaction between the digital human front end, the digital human central control unit and the digital human back end; the second acquisition unit 403 is used to acquire digital human audio and video stream data based on the interaction between the digital human central control unit and the digital human back end, and send the digital human audio and video stream data to the audio and video platform.

[0104] Specifically, the initialization unit 401 in the device is used to perform initialization operations on the digital human system based on the interaction between the digital human front end and the digital human back end; the first acquisition unit 402 is used to acquire speech recognition data and synthesized speech data in response to user voice input based on the interaction between the digital human front end, the digital human central control unit, and the digital human back end; the second acquisition unit 403 is used to acquire digital human audio and video stream data based on the interaction between the digital human central control unit and the digital human back end, and send the digital human audio and video stream data to the audio and video platform.

[0105] The digital human audio and video stream data acquisition device provided in this embodiment of the invention initializes the digital human system based on the interaction between the digital human front end and the digital human back end; acquires speech recognition data and synthesized speech data in response to user voice input based on the interaction between the digital human front end, the digital human central control unit, and the digital human back end; and acquires digital human audio and video stream data and sends the digital human audio and video stream data to the audio and video platform based on the interaction between the digital human central control unit and the digital human back end, thereby reducing the latency in the process of acquiring digital human audio and video stream data.

[0106] The embodiments of the present invention provide a digital human audio and video stream data acquisition device that can be used to execute the processing flow of the above method embodiments. Its functions will not be repeated here, but can be referred to the detailed description of the above method embodiments.

[0107] Figure 5 This is a schematic diagram of the physical structure of a computer device provided in an embodiment of the present invention, such as... Figure 5 As shown, the computer device includes: a memory 501, a processor 502, and a computer program stored in the memory 501 and executable on the processor 502. When the processor 502 executes the computer program, it implements the following method:

[0108] The initialization operation of the digital human system is achieved based on the interaction between the digital human front end and the digital human back end.

[0109] Based on the interaction between the digital human front end, the digital human central control unit, and the digital human back end, synthesized voice data is used to acquire speech recognition data and respond to user voice input.

[0110] Based on the interaction between the digital human central control unit and the digital human backend, the digital human audio and video stream data is acquired and sent to the audio and video platform.

[0111] This embodiment discloses a computer program product, which includes a computer program that, when executed by a processor, implements the following method:

[0112] The initialization operation of the digital human system is achieved based on the interaction between the digital human front end and the digital human back end.

[0113] Based on the interaction between the digital human front end, the digital human central control unit, and the digital human back end, synthesized voice data is used to acquire speech recognition data and respond to user voice input.

[0114] Based on the interaction between the digital human central control unit and the digital human backend, the digital human audio and video stream data is acquired and sent to the audio and video platform.

[0115] This embodiment provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the following method:

[0116] The initialization operation of the digital human system is achieved based on the interaction between the digital human front end and the digital human back end.

[0117] Based on the interaction between the digital human front end, the digital human central control unit, and the digital human back end, synthesized voice data is used to acquire speech recognition data and respond to user voice input.

[0118] Based on the interaction between the digital human central control unit and the digital human backend, the digital human audio and video stream data is acquired and sent to the audio and video platform.

[0119] Compared with existing technical solutions, the digital human audio and video stream data acquisition method provided by this invention initializes the digital human system based on the interaction between the digital human front-end and the digital human back-end; acquires speech recognition data and synthesized speech data in response to user voice input based on the interaction between the digital human front-end, the digital human central control unit, and the digital human back-end; and acquires digital human audio and video stream data and sends the digital human audio and video stream data to the audio and video platform based on the interaction between the digital human central control unit and the digital human back-end, thereby reducing latency in the process of acquiring digital human audio and video stream data.

[0120] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0121] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0122] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0123] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0124] In the description of this specification, the references to terms such as "an embodiment," "a specific embodiment," "some embodiments," "for example," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0125] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for acquiring digital human audio and video stream data, characterized in that, The digital human audio and video stream data acquisition method is applied to a digital human system, which includes a digital human front-end, a digital human central control unit, and a digital human back-end; the digital human audio and video stream data acquisition method includes: The initialization operation of the digital human system is achieved based on the interaction between the digital human front end and the digital human back end. Based on the interaction between the digital human front end, the digital human central control unit, and the digital human back end, synthesized voice data is used to acquire speech recognition data and respond to user voice input. Based on the interaction between the digital human central control unit and the digital human backend, the digital human audio and video stream data is acquired and sent to the audio and video platform.

2. The method for acquiring digital human audio and video stream data according to claim 1, characterized in that, The initialization operation of the digital human system based on the interaction between the digital human front end and the digital human back end includes: In response to a user accessing the digital human system, the digital human front end plays pre-stored offline audio and video data of the digital human to the user, so that the digital human front end can obtain user feedback information generated by the user in response to the offline audio and video data of the digital human. The digital human front end sends the user feedback information to the digital human back end; The digital human backend performs initialization operations on the digital human system based on the user feedback information.

3. The method for acquiring digital human audio and video stream data according to claim 2, characterized in that, The user feedback information includes user audiovisual sensory feedback information; correspondingly, the digital human backend performs initialization operations on the digital human system based on the user feedback information, including: The digital human backend identifies the user's audiovisual sensory feedback information to obtain user identification information and user intent information, and performs initialization operations on the digital human system based on the user identification information and the user intent information.

4. The method for acquiring digital human audio and video stream data according to claim 3, characterized in that, The method of acquiring speech recognition data and responding to user voice input based on the interaction between the digital human front-end, the digital human central control unit, and the digital human back-end includes: The digital human front end sends the real-time collected user audio data to the digital human central control unit; The digital human central control unit sends the user's audio data to the digital human backend in real time; The digital human backend recognizes the user's audio data to obtain speech recognition data, and then sends the speech recognition data to the digital human central control unit.

5. The method for acquiring digital human audio and video stream data according to claim 4, characterized in that, The method of acquiring speech recognition data and responding to user voice input based on the interaction between the digital human front-end, the digital human central control unit, and the digital human back-end includes: The digital human central control unit sends the voice recognition data to the digital human front end; After the user inputs voice data, the digital human central control unit sends the text data corresponding to all user audio data to the digital human backend. The digital human backend generates synthesized speech data to respond to the user's voice input based on the text data corresponding to all user audio data, and sends the synthesized speech data to the digital human central control in the form of data fragments.

6. The method for acquiring digital human audio and video stream data according to claim 5, characterized in that, The method of acquiring digital human audio and video stream data based on the interaction between the digital human central control unit and the digital human backend includes: The digital human central control unit sends the synthesized voice data to the digital human backend; The digital human backend drives and renders the digital human based on the synthesized speech data to obtain digital human audio and video stream data.

7. A digital human audio and video stream data acquisition device, characterized in that, The digital human audio and video stream data acquisition device is applied to a digital human system, which includes a digital human front-end, a digital human central control unit, and a digital human back-end; the digital human audio and video stream data acquisition device includes: An initialization unit is used to perform initialization operations on the digital human system based on the interaction between the digital human front end and the digital human back end. The first acquisition unit is used to acquire speech recognition data and synthesized speech data in response to user voice input based on the interaction between the digital human front end, the digital human central control unit and the digital human back end; The second acquisition unit is used to acquire digital human audio and video stream data based on the interaction between the digital human central control unit and the digital human backend, and send the digital human audio and video stream data to the audio and video platform.

8. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method of any one of claims 1 to 6.

10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the method of any one of claims 1 to 6.