Live broadcast stream pushing method based on intelligent digital human model
By combining voice interaction, digital character generation and live streaming technology, a live streaming method based on intelligent digital human model is developed, which solves the problems of insufficient semantic understanding and generation capabilities of voice interaction systems, lack of interactivity in digital character models, and lack of intelligent interaction in the live streaming mode in the existing technology, real-time interactive live broadcast driven by voice is realized, improving user experience and expanding the application field.
Patent Information
- Application Number
- CN202510055273.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-06
AI Technical Summary
In the prior art, voice interaction systems lack extensive semantic understanding and generation capabilities, digital character models lack interactivity, and existing live broadcast modes also lack intelligent interaction, making it difficult to realize voice-driven real-time interactive live broadcast.
By combining voice interaction, digital character generation and live streaming technology, a live streaming method based on intelligent digital human model is developed to realize that users can obtain real-time image responses through natural voice interaction and enhance user experience. Specific steps include preloading the lip model, recording user voice in real time, voice recognition, generating reply text and video, and pushing it to the front end in real time through the RTMP server.
The combination of live streaming technology and intelligent human-computer interaction system has been realized. Users can obtain real-time image responses through natural voice interaction, which improves user experience and expands the application fields of voice interaction and digital character models.
Smart Images

Figure CN119946313A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of digital human model technology and live streaming technology, and in particular to a live streaming method based on an intelligent digital human model. Background Art
[0002] The origins of voice interaction technology can be traced back to the 1950s, when the technology was limited to recognizing an extremely limited vocabulary. Early speech recognition systems relied on template matching technology, where users had to speak in a specific way and rhythm so that the system could compare with pre-recorded word templates. The limitations of this approach are very obvious. It cannot adapt to the pronunciation differences between different users, nor can it process continuous natural language.
[0003] With the advancement of computing technology, hidden Markov models (HMMs) began to be used in speech recognition in the late 1990s, marking an important shift in speech recognition technology. HMMs provide an effective way to process time series data, allowing the system to better handle continuous speech. However, these systems still require a lot of manual tuning and expertise to optimize performance.
[0004] Entering the 21st century, with the rise of deep learning, speech recognition technology has undergone revolutionary changes. Neural networks, especially convolutional neural networks (CNN) and recurrent neural networks (RNN), have shown great potential in the field of speech recognition. These technologies can automatically learn features from large amounts of data, significantly improving the accuracy of recognition.
[0005] Nowadays, speech recognition technology has become an integral part of our daily lives. Smartphones, smart speakers, car systems and other devices have integrated speech recognition capabilities, allowing users to interact with devices through voice. The most advanced speech recognition systems can recognize natural language with an accuracy rate of over 95%, and can maintain high accuracy even in noisy environments.
[0006] This progress is attributed to several key factors. The first is the availability of big data. Massive amounts of speech data allow machine learning models to better understand and predict human language. The second is the increase in computing power, especially the emergence of specialized hardware such as GPUs and TPUs, which greatly accelerates the training process of deep learning models. In addition, the continued development of new algorithms and architectures is also driving progress in this field.
[0007] The development of digital human model technology is closely linked to the history of computer graphics. As early as the 1970s, computer graphics had begun to attempt to simulate human appearance and movement. Initially, these models were very simple, just some basic geometric shapes and lines. As technology developed, these models became more and more complex and realistic. Early breakthroughs included 3D modeling of human faces and bodies, and simulation of body movements.
[0008] In the 1990s, as personal computers became more powerful and graphics processing technology advanced, digital characters began to appear in video games and movies. During this period, digital character models gradually changed from a stiff, rigid appearance to a more natural and realistic image. The introduction of motion capture technology was an important milestone, allowing digital characters to more naturally imitate the movements and expressions of real people.
[0009] Modern digital character modeling technology has achieved amazing realism. Advanced facial modeling technology can accurately capture and reproduce human expressions, even subtle eyes and smiles. The application of these technologies is not limited to entertainment fields, such as movies and games, but has also begun to expand to other fields, such as virtual reality, distance education, and even medical simulation.
[0010] The progress of technology is mainly reflected in two aspects. First, the improvement of modeling technology can create more and more detailed and realistic digital character models. Second, the progress of motion capture technology can capture and reproduce human movements more accurately and naturally. In addition, the integration of artificial intelligence technology enables digital characters to perform more complex interactions and responses, providing users with a richer experience.
[0011] Despite the great progress, digital character model technology still faces many challenges. The first is the issue of realism. Although modern technology can create very realistic models, it is still a challenge to completely reach the point where it is impossible to distinguish between the real and the fake. In addition, real-time rendering of highly complex digital character models requires huge computing resources, which places higher demands on the hardware. There is also the issue of interactivity. How to enable digital characters to interact with the real world and users more naturally is still an issue to be solved.
[0012] The history of live streaming can be traced back to the mid-1990s, when video quality was often poor and latency was very high due to bandwidth limitations. However, with the development of Internet technology, especially the popularity of broadband Internet, live streaming technology began to develop rapidly. In the early 2000s, the first batch of dedicated live streaming platforms emerged, providing better video quality and lower latency.
[0013] With the popularization of smartphones and mobile Internet, the live broadcast industry has experienced explosive growth in the 2010s. Users can not only watch live broadcasts, but also easily broadcast them through their own devices. During this period, the number of live broadcast platforms increased dramatically, and the live broadcast content became increasingly diversified, ranging from game live broadcasts to education, e-commerce and other fields.
[0014] Current live broadcast technology is already very mature, providing high-definition video, low-latency transmission, and a stable streaming experience. A variety of live broadcast software and platforms make live broadcasting extremely easy, and anyone with a smartphone can become a live broadcaster. In addition to individual users, many companies and organizations have also begun to use live broadcast technology for product demonstrations, education and training, remote meetings, etc.
[0015] The diversification of live broadcast content is also a significant trend. From the initial game and entertainment content, it has developed to education, fitness, cooking, travel and almost every aspect of life. In addition, e-commerce live broadcast has also become an important trend, especially in the Asian market. Through live broadcast, merchants can intuitively display products, interact with viewers in real time, and promote sales.
[0016] Although technology has made significant progress, the live broadcasting field still faces some challenges. The first is the bandwidth and traffic problem. High-quality live broadcasting requires a high upload speed, which is still a limiting factor in some regions. The second is the content supervision issue. With the increase in live broadcast content, how to effectively manage and filter inappropriate content has become a challenge.
[0017] In summary, the current voice interaction, digital character modeling and live broadcast technology have made great progress, but the cases applied to the live broadcast field are still relatively rare. The voice interaction system lacks the ability to understand and generate semantics with wide coverage, the digital character model lacks interactivity, and the existing live broadcast mode lacks intelligent interaction. Therefore, it is very necessary to develop an intelligent system that can combine voice interaction, digital character generation and live broadcast streaming. In order to realize voice-driven real-time interactive live broadcast, an intelligent live broadcast interaction system based on voice input and digital characters is proposed. Summary of the invention
[0018] In view of the above technical problems, the present invention provides a live streaming method based on an intelligent digital human model, which combines the live streaming technology with an intelligent human-computer interaction system, enabling users to obtain real-time image responses through natural voice interaction, thereby greatly enhancing the user experience.
[0019] The technical solution of the present invention is as follows:
[0020] The live streaming method based on the intelligent digital human model has the following specific steps:
[0021] S100: preloads the lip shape model when the system is deployed and preprocesses the character video;
[0022] S200: The front-end part records user voice in real time and calls iFlytek ASR through websocket for real-time voice recognition;
[0023] S300: After the speech recognition is successful, the recognized text is sent to the streaming interface via http, and the conversation ID, response text and http-flv link are obtained;
[0024] S400: After receiving the http-flv link, wait, obtain and process the video stream data, and display it synchronously with the text;
[0025] S500: The user chooses to stop answering midway, and the current conversation ID is sent to the stop streaming interface via http;
[0026] S600: the streaming interface creates a conversation ID after receiving the text sent by the front-end part, and generates a reply text through the language model;
[0027] S700: The streaming interface creates a sub-thread based on the conversation ID, and processes the video stream through the sub-thread;
[0028] S800: calling the speech synthesis interface in the child thread to convert the reply text into speech;
[0029] S900: The subthread of the streaming interface combines the voice and the pre-processed video data to generate a lip-aligned video;
[0030] S1000: In the subthread of the streaming interface, FFmpeg is used to push the processed video to the RTMP server in real time. The RTMP server converts the received RTMP stream into an http-flv stream in real time and pushes it to the front-end.
[0031] S1100: After receiving the conversation ID requested by the front-end part, a termination signal is sent to the thread corresponding to the ID to stop pushing the video stream corresponding to the ID.
[0032] Specifically, the lip shape model in step S100 is a neural network model based on deep learning, which can generate lip animation matching the input voice data. Preloading the model can improve the response speed and efficiency of the system.
[0033] Specifically, the specific content of the character video preprocessing in step S100 is: preprocessing the video to be put into the digital character model, such as extracting action information, background denoising, etc., to improve the quality of the generated video.
[0034] Specifically, the specific content of the streaming interface creating a sub-thread based on the conversation ID in step S700 is: when the streaming interface receives the text information sent by the front-end part, it will create a new sub-thread. By creating the sub-thread, the system's concurrent processing capability is ensured, so that the system can process requests from multiple users at the same time, thereby improving the system's processing efficiency.
[0035] Specifically, the specific content of calling the speech synthesis interface in the sub-thread in step S800 to convert the reply text into speech is: using a large language model to process the received user conversation text and generate a corresponding response, and the processed response is converted into speech data through the speech synthesis interface, wherein the large language model is a neural network model based on deep learning, which can generate a matching answer based on the input conversation text.
[0036] Beneficial effects of the present invention:
[0037] The real-time streaming system based on voice input of the present invention effectively realizes the combination of live streaming technology and intelligent human-computer interaction system. The system uses the front-end part to collect the user's voice in real time, convert the language into text, and then generate voice and video images matching the voice in real time through the streaming interface. The video stream is then pushed to the front-end part in real time through the RTMP server, which improves the user's interactive experience. When the user wants to stop answering, he can stop the video push in time by terminating the streaming interface operation, avoiding unnecessary waste of resources.
[0038] The present invention can understand user voice input in real time, use digital characters to intelligently generate corresponding voice and video, and push them to users in real time, thus achieving a live broadcast effect driven by voice interaction. This can not only improve the interactivity and fun of live broadcasts, but also expand the application fields of voice interaction and digital character models.
[0039] This system has broad application prospects and can be applied to multiple fields such as online education, virtual anchors, e-commerce live broadcasts, etc., providing users with a richer and more interactive live broadcast experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 It is a schematic diagram of the overall process of the present invention;
[0041] Figure 2 It is a schematic diagram of the flow chart of the front end part of the present invention;
[0042] Figure 3 It is a schematic diagram of the flow of the streaming interface of the present invention;
[0043] Figure 4 The figure is a flow chart of stopping the streaming interface of the present invention. DETAILED DESCRIPTION
[0044] The technical solution of the present invention will be described clearly and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0045] Example 1
[0046] In the embodiment of the live streaming method based on the digital human model, the entire process architecture can be described as follows: the front-end part records the user's voice input, processes it in real time through the voice recognition service, and displays the results. When the voice recognition is successful, the text is sent to the streaming interface, and a conversation with text and streaming link responses is started. During the live broadcast, the video is displayed synchronously with the text. It also involves pre-loading the lip synchronization model and pre-processing the digital human video. The streaming interface generates a response using a large language model by receiving the text request from the front-end part, and creates a video synchronized with the lip shape. The video is then broadcast live to the server and converted into a streaming format displayed on the front end. At the same time, an interface is provided to stop the live broadcast according to the user's request. This system example demonstrates the integration of voice recognition, language processing and video live broadcast to create an interactive digital human live broadcast experience.
[0047] The specific implementation of the present invention is further described in detail below in conjunction with the accompanying drawings. Figure 1 As shown, the specific steps of the method are as follows:
[0048] S100: preloads the lip shape model when the system is deployed and preprocesses the character video;
[0049] S200: The front-end part records user voice in real time and calls iFlytek ASR through websocket for real-time voice recognition;
[0050] S300: After the speech recognition is successful, the recognized text is sent to the streaming interface via http, and the conversation ID, response text and http-flv link are obtained;
[0051] S400: After receiving the http-flv link, wait, obtain and process the video stream data, and display it synchronously with the text;
[0052] S500: The user chooses to stop answering midway, and the current conversation ID is sent to the stop streaming interface via http;
[0053] S600: the streaming interface creates a conversation ID after receiving the text sent by the front-end part, and generates a reply text through the language model;
[0054] S700: The streaming interface creates a sub-thread based on the conversation ID, and processes the video stream through the sub-thread;
[0055] S800: calling the speech synthesis interface in the child thread to convert the reply text into speech;
[0056] S900: The subthread of the streaming interface combines the voice and the pre-processed video data to generate a lip-aligned video;
[0057] S1000: In the subthread of the streaming interface, FFmpeg is used to push the processed video to the RTMP server in real time. The RTMP server converts the received RTMP stream into an http-flv stream in real time and pushes it to the front-end.
[0058] S1100: After receiving the conversation ID requested by the front-end part, a termination signal is sent to the thread corresponding to the ID to stop pushing the video stream corresponding to the ID.
[0059] like Figure 2 ,The front-end part plays a key role in the live streaming method of the intelligent digital human model. The functions of the front-end part include receiving user input, calling the intelligent digital human model in real time, displaying the model's results, and sending the results to the streaming interface, obtaining the conversation id, response text and http-flv link. The user inputs voice in the front-end interface. When the user inputs, the front-end plays a listening animation to enhance the user experience and let the user know that the system is processing their input.
[0060] The front-end part uses websocket to call iFlytek ASR in real time for speech recognition. websocket is a protocol for full-duplex communication on a single TCP connection, which allows data to be quickly transmitted between the client and the server. The intelligent digital human model generates the corresponding video based on the user's input. After the text recognition is successful, the front-end part sends the text to the push stream interface via http. http is an application layer protocol that enables communication between the client and the server through the exchange of requests and responses. The front-end part obtains the conversation id, response text and http-flv link. The conversation id is used to identify each conversation. The response text is the text content of the video generated by the intelligent digital human model based on the user's input, and the http-flv link is the link of the video stream. The front-end part waits for the video stream data of the http-flv link. During the waiting process, the front-end part plays a waiting animation to enhance the user experience and let the user know that the system is processing their request. After the data is obtained, the front-end part will display the video stream synchronously. The user clicks to stop answering, and the front-end part sends the current conversation id to the stop push stream interface via http. The front-end part changes to play the listening animation and waits for the user's next input. The user can stop the current video playback at any time and start a new input. After the video stream is played, the front-end plays the listening animation again and waits for the user's next input. The system can continuously provide services to users and meet their needs. The workflow of the front-end is a cyclic process, always waiting for user input, processing user input, displaying processing results, and waiting for the user's next input. This process is both efficient and flexible and can meet a variety of user needs.
[0061] like Figure 3, the main responsibility of the push streaming interface is to receive the text requested by the front-end part, create the current conversation ID, answer the text through the large language model, get the reply text, and create a sub-thread based on the conversation ID for video processing. After receiving the text requested by the front-end part, the push streaming interface immediately creates the current conversation ID. The conversation ID is a unique identifier used to identify each conversation and ensure the independence and coherence of each conversation. After creating the conversation ID, the push streaming interface answers the text through the large language model to get the reply text. Among them, the large language model is an artificial intelligence model that can understand and generate human language to provide users with accurate and natural replies. After getting the reply text, the push streaming interface creates a sub-thread based on the conversation ID for video processing. The sub-thread is an independent execution process that can process tasks in parallel to improve the processing efficiency of the system. In the sub-thread, the push streaming interface uses an intelligent digital human model to synthesize the video corresponding to the reply text. The intelligent digital human model is an artificial intelligence model that can generate corresponding videos based on text to provide users with a rich and vivid visual experience. After generating the video, the push streaming interface uses FFmpeg to push the video to the RTMP server in real time. FFmpeg is a multimedia processing tool that can process audio and video in various formats. RTMP server is a streaming media server that can transmit audio and video in real time. The RTMP server converts the RTMP stream into an http-flv stream in real time and pushes it to the front-end. Among them, the http-flv stream is a streaming media format that can transmit audio and video in real time, providing users with a smooth and high-definition viewing experience. After receiving the conversation ID requested by the front-end, the push stream interface sends a termination signal to the thread corresponding to the ID and responds to the front-end part with the interrupt result. In this way, the user can stop the current video playback at any time and start a new input. In general, the workflow of the push stream interface is an efficient and flexible process that can meet a variety of user needs.
[0062] like Figure 4The main responsibility of the stop push interface is to receive the conversation ID requested by the front-end, send a termination signal to the thread corresponding to the ID, and respond to the front-end with the interrupt result. The workflow of the stop push interface starts with receiving the conversation ID requested by the front-end. The conversation ID is a unique identifier used to identify each conversation to ensure the independence and continuity of each conversation. When the user wants to stop the current video playback, the front-end will send the current conversation ID to the stop push interface through http. After receiving the conversation ID, the stop push interface sends a termination signal to the thread corresponding to the ID. This termination signal is a special signal used to notify the child thread to stop the current task. After receiving the termination signal, the child thread will immediately stop the current task and release related resources. After sending the termination signal, the stop push interface responds to the front-end with the interrupt result. The interrupt result is a special response used to notify the front-end that the current video playback has been successfully stopped. After receiving the interrupt result, the front-end will immediately stop text output, play the listening animation instead, and wait for the user's next input. In general, the workflow of the stop push interface is a concise and efficient process. It can quickly respond to user requests and immediately stop the current video playback, providing users with more flexible and personalized services.
[0063] The above shows and describes the basic principles and main features of the present invention and the advantages of the present invention. It should be understood by those skilled in the art that the present invention is not limited to the above embodiments. The above embodiments and descriptions are only for explaining the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, which fall within the scope of the present invention to be protected. The scope of protection of the present invention is defined by the attached claims and their equivalents.
Claims
1. The live streaming method based on the intelligent digital human model has the following specific steps: S100: preloads the lip shape model when the system is deployed and preprocesses the character video; S200: The front-end part records user voice in real time and calls iFlytek ASR through websocket for real-time voice recognition; S300: After the speech recognition is successful, the recognized text is sent to the streaming interface via http, and the conversation ID, response text and http-flv link are obtained; S400: After receiving the http-flv link, wait, obtain and process the video stream data, and display it synchronously with the text; S5 00: The user chooses to stop answering, and the current conversation ID is sent to the stop streaming interface via http; S600: the streaming interface creates a conversation ID after receiving the text sent by the front-end part, and generates a reply text through the language model; S700: The streaming interface creates a sub-thread based on the conversation ID, and processes the video stream through the sub-thread; S800: calling the speech synthesis interface in the child thread to convert the reply text into speech; S900: The subthread of the streaming interface combines the voice and the pre-processed video data to generate a lip-aligned video; S1000: In the subthread of the streaming interface, FFmpeg is used to push the processed video to the RTMP server in real time. The RTMP server converts the received RTMP stream into an http-flv stream in real time and pushes it to the front-end. S1100: After receiving the conversation ID requested by the front-end part, a termination signal is sent to the thread corresponding to the ID to stop pushing the video stream corresponding to the ID.
2. The live streaming method based on the intelligent digital human model according to claim 1 is characterized in that The lip shape model in step S100 is a neural network model based on deep learning.
3. The live streaming method based on the intelligent digital human model according to claim 1 is characterized in that The specific content of the character video preprocessing in step S100 is: preprocessing the video to be put into the digital character model.
4. The live streaming method based on the intelligent digital human model according to claim 1 is characterized in that The specific content of the streaming interface creating a sub-thread based on the conversation ID in step S700 is: when the streaming interface receives the text information sent by the front-end part, a new sub-thread will be created.
5. The live streaming method based on the intelligent digital human model according to claim 1 is characterized in that The specific content of calling the speech synthesis interface in the sub-thread in step S800 to convert the reply text into speech is: using a large language model to process the received user conversation text and generate a corresponding response, and the processed response is converted into speech data through the speech synthesis interface, wherein the large language model is a neural network model based on deep learning, which can generate a matching answer based on the input conversation text.