Video-based real-time multi-mode digital human AIGC interaction technology
By training a general big model and external camera to obtain image frames, the problem that digital human videos cannot interact in real time and multimodal interaction is solved, real-time application and multimodal capabilities are realized in multiple scenarios, and the interaction efficiency and information input capabilities of digital human technology are improved.
Patent Information
- Application Number
- CN202510873060.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-08-22
AI Technical Summary
Existing digital human videos cannot interact in real time and lack multimodal interaction capabilities, which limits their application in multiple scenarios.
By recording and processing target character videos, train a general-purpose big model with a resolution of 256x256, combine external cameras to acquire image frames, realize real-time interaction and multimodal interaction, use streaming to process speech synthesis, generate real-time interactive videos and perform multimodal tasks.
It realizes the application of digital people in real-time interactive scenarios, such as customer service and display exhibitions, enhances information input capabilities, improves rendering speed and clarity, and supports multimodal tasks such as image recognition.
Smart Images

Figure CN120526795A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of AIGC interaction technology, and in particular to a video-based real-time multimodal digital human AIGC interaction technology. Background Art
[0002] In recent years, with the development of deep learning technology, the technology for generating speaking videos by driving two-dimensional digital human faces with voice has rapidly developed and has been applied in various fields. These speaking videos are AIGC (Artificial Intelligence Generated Content). Because they are generated by replacing mouth shapes with real-life videos, they appear realistic to ordinary users and are therefore widely used in short videos and other fields.
[0003] However, existing digital human videos have the following two problems: first, they are non-interactive, take a long time to generate, and cannot respond to real-time user needs; second, they only have voice or text as input, and lack input methods with richer information such as cameras, which limits the applicability of digital humans. Summary of the Invention
[0004] In response to the shortcomings of the existing technology, the purpose of the present invention is to provide a CAT water meter data collection method and device to solve the problems raised in the above-mentioned background technology. The present invention has a novel structure. After possessing real-time interaction capabilities, the digital human technology can be applied in more scenarios where real-time interaction is required, such as customer service, exhibitions, tour guides, etc.; after possessing multimodal interaction capabilities, the present invention enables the digital human to have more interaction methods and receive more information input, greatly enhancing the application scope of the digital human.
[0005] To achieve the above objectives, the present invention is implemented through the following technical solutions: a video-based real-time multimodal digital human AIGC interaction technology, the real-time multimodal digital human AIGC interaction technology includes digital human real-time interaction and multimodal interaction, and the digital human real-time interaction includes the following steps:
[0006] (1) Record a video of the target person for at least 2 minutes;
[0007] (2) The characters in the recorded video do not show their teeth or speak, and the frame rate is 25 frames per second;
[0008] (3) Cut the video into pictures;
[0009] (4) Perform face recognition on each frame and crop the face image to save it separately;
[0010] (5) Extract the coordinates x1, y1, x2, y2 of the face image relative to the original image and save it as a pkl file;
[0011] (6) looping and reversing the original image to form a complete closed-loop video;
[0012] (7) Text input, calling the speech synthesis interface to synthesize audio data with a sampling rate of 16000Hz;
[0013] (8) Split the received audio data into 12.5ms audio blocks, which contain 16000×12.5÷1000=200 audio data;
[0014] (9) Input the above audio data block and the previous ten frames of images into the 256×256 model, and the model generates a new image;
[0015] (10) The generated image is used to replace the original image, and we obtain a video of the target person speaking, completing real-time interaction;
[0016] Multimodal interaction includes the following steps:
[0017] (1) First, the user inputs text or voice;
[0018] (2) After receiving user input, the program calls the camera to acquire images;
[0019] (3) The program then feeds the captured images, as well as multimodal data such as text and speech, into the large model;
[0020] (4) The large model provides text feedback and answers user questions;
[0021] (5) After the answer is completed, the program resumes text and voice monitoring and waits for the next input to complete the closed loop.
[0022] Furthermore, in the digital human real-time interaction step 5, the video segmentation results are subjected to face recognition, and face pictures and coordinate information are obtained and stored locally.
[0023] Furthermore, in the multimodal interaction step 9, the audio data block, face image and face coordinate information will be uniformly input into the digital human model, and the digital human model will generate a mouth shape replacement image corresponding to the current audio block based on this information.
[0024] Furthermore, in the multimodal interaction step 1, the device used by the user for text or voice input includes an input device such as a microphone, a keyboard, and a camera.
[0025] Furthermore, in the multimodal interaction step 2, the image acquisition is performed by calling an external camera to acquire a series of image frames, and then these image frames are input into the visual macro model.
[0026] Furthermore, in the multimodal interaction step 3, the output result of the large model will enter the streaming speech synthesis module for streaming speech output, and the output result will enter the digital human rendering module for real-time digital human rendering output, and finally the rendering result will be output to the output device in real time.
[0027] Furthermore, the output device includes a display, AR glasses, etc.
[0028] Beneficial effects of the present invention:
[0029] 1. This invention trains a universal large model with a resolution of 256x256. This eliminates the need for separate training for each digital human, improving the clarity and rendering speed of the digital human. Furthermore, for speech synthesis, this solution uses streaming processing, enabling the synthesis results of the digital human to be output at a speed of 25 frames per second, ultimately achieving real-time effects.
[0030] 2. This invention acquires a series of image frames by calling an external camera, and then inputs these image frames into a large visual model to achieve multimodal interaction capabilities, allowing the digital human to perform image recognition, thereby completing multimodal tasks such as document recognition, face recognition, and image classification. This solves the problems of existing digital human solutions such as the inability to interact in real time and the lack of multimodal capabilities.
[0031] 3. Compared with existing technologies, the present invention, with its real-time interaction capabilities, enables digital human technology to be applied in more scenarios requiring real-time interaction, such as customer service, exhibitions, and tour guides. With its multimodal interaction capabilities, the present invention enables digital humans to have more interaction methods and receive more information input, greatly enhancing the application scope of digital humans. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 This is a principle flow chart of a real-time digital human based on a video-based real-time multimodal digital human AIGC interaction technology of the present invention;
[0033] Figure 2 This is a principle flow chart of multimodal interaction of the video-based real-time multimodal digital human AIGC interaction technology of the present invention. DETAILED DESCRIPTION
[0034] In order to make the technical means, creative features, objectives and effects achieved by the present invention easier to understand, the present invention is further described below in conjunction with specific implementation methods.
[0035] See also Figures 1 to 2 , the present invention provides a technical solution:
[0036] A video-based real-time multimodal digital human AIGC interaction technology, the real-time multimodal digital human AIGC interaction technology includes real-time digital human interaction and multimodal interaction, and the real-time digital human interaction includes the following steps:
[0037] (1) Record a video of the target person for at least 2 minutes;
[0038] (2) The characters in the recorded video do not show their teeth or speak, and the frame rate is 25 frames per second;
[0039] (3) Cut the video into pictures;
[0040] (4) Perform face recognition on each frame and crop the face image to save it separately;
[0041] (5) Extract the coordinates x1, y1, x2, y2 of the face image relative to the original image and save it as a pkl file;
[0042] (6) looping and reversing the original image to form a complete closed-loop video;
[0043] (7) Text input, calling the speech synthesis interface to synthesize audio data with a sampling rate of 16000Hz;
[0044] (8) Split the received audio data into 12.5ms audio blocks, which contain 16000×12.5÷1000=200 audio data;
[0045] (9) Input the above audio data block and the previous ten frames of images into the 256×256 model, and the model generates a new image;
[0046] (10) The generated image is used to replace the original image, and we obtain a video of the target person speaking, completing real-time interaction;
[0047] In this embodiment, first, the user records a real-life video and segments the video into pictures. The segmentation results are then used for face recognition to obtain face pictures and coordinate information, which are then saved locally. Subsequently, when a speech text request is received, the text is first converted into audio data and then segmented into audio data blocks according to the length of 12.5ms. Finally, the audio data blocks, face pictures, and face coordinate information are uniformly input into the digital human model. Based on this information, the digital human model generates a mouth shape replacement picture corresponding to the current audio block. This picture replaces the original picture to obtain a speaking digital human video stream, realizing real-time digital human interaction.
[0048] Multimodal interaction includes the following steps:
[0049] (1) First, the user inputs text or voice;
[0050] (2) After receiving user input, the program calls the camera to acquire images;
[0051] (3) The program then feeds the captured images, as well as multimodal data such as text and speech, into the large model;
[0052] (4) The large model provides text feedback and answers user questions;
[0053] (5) After the answer is completed, the program resumes text and voice monitoring and waits for the next input to complete the closed loop.
[0054] In this embodiment, first, the user inputs multimodal information through input devices such as microphones, keyboards, and cameras. Then, this information will be input into the multimodal large model. The output result of the large model will then enter the streaming speech synthesis module for streaming speech output. The output result will then enter the digital human rendering module for real-time digital human rendering output. Finally, the rendering result will be output in real time to output devices such as displays, AR glasses, etc., and wait for the user's subsequent output to achieve an interactive closed loop.
[0055] For the real-time interaction of the digital human, this solution trains a universal large model with a resolution of 256x256. This eliminates the need for separate training for each digital human, improving their clarity and rendering speed. Furthermore, for speech synthesis, this solution uses streaming processing, enabling the output of the digital human's synthesis results at a rate of 25 frames per second, ultimately achieving real-time performance. For multimodal interaction, this solution first acquires a series of image frames by invoking an external camera. These frames are then input into the large visual model to achieve multimodal interaction capabilities, enabling the digital human to perform image recognition, thereby completing multimodal tasks such as document recognition, face recognition, and image classification. This solution addresses the problems of existing digital human solutions, such as the inability to interact in real time and the lack of multimodal capabilities. With real-time interaction capabilities, this solution enables digital human technology to be applied in more scenarios requiring real-time interaction, such as customer service, exhibitions, and tour guides. With multimodal interaction capabilities, this solution enables digital humans to have more interactive modes and receive more information input, greatly expanding their application areas.
[0056] The basic principles, main features and advantages of the present invention are shown and described above. It is obvious to those skilled in the art that the present invention is not limited to the details of the above exemplary embodiments, and that the present invention can be implemented in other specific forms without departing from the spirit or basic features of the present invention.
[0057] In addition, it should be understood that although this specification is described in terms of implementation methods, not every implementation method contains only one independent technical solution. This narrative method of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment can also be appropriately combined to form other implementation methods that can be understood by those skilled in the art.
Claims
1. A video-based real-time multimodal digital human AIGC interaction technology, characterized by: The real-time multimodal digital human AIGC interaction technology includes real-time digital human interaction and multimodal interaction. The real-time digital human interaction includes the following steps: (1) Record a video of the target person for at least 2 minutes; (2) The characters in the recorded video do not show their teeth or speak, and the frame rate is 25 frames per second; (3) Cut the video into pictures; (4) Perform face recognition on each frame and crop the face image to save it separately; (5) Extract the coordinates x1, y1, x2, y2 of the face image relative to the original image and save it as a pkl file; (6) looping and reversing the original image to form a complete closed-loop video; (7) Text input, calling the speech synthesis interface to synthesize audio data with a sampling rate of 16000Hz; (8) Split the received audio data into 12.5ms audio blocks, which contain 16000×12.5÷1000=200 audio data; (9) Input the above audio data block and the previous ten frames of images into the 256×256 model, and the model generates a new image; (10) The generated image is used to replace the original image, and we obtain a video of the target person speaking, completing real-time interaction; Multimodal interaction includes the following steps: (1) First, the user inputs text or voice; (2) After receiving user input, the program calls the camera to acquire images; (3) The program then feeds the captured images, as well as multimodal data such as text and speech, into the large model; (4) The large model provides text feedback and answers user questions; (5) After the answer is completed, the program resumes text and voice monitoring and waits for the next input to complete the closed loop.
2. The video-based real-time multimodal digital human AIGC interaction technology according to claim 1, characterized in that: In the digital human real-time interaction step 5, the video segmentation results are respectively subjected to face recognition, and face pictures and coordinate information are obtained and stored locally.
3. The video-based real-time multimodal digital human AIGC interaction technology according to claim 2, characterized in that: In the multimodal interaction step 9, the audio data block, face image and face coordinate information will be uniformly input into the digital human model, and the digital human model will generate a mouth shape replacement image corresponding to the current audio block based on this information.
4. The video-based real-time multimodal digital human AIGC interaction technology according to claim 1, characterized in that: In the multimodal interaction step 1, the device used by the user to input text or voice includes an input device such as a microphone, a keyboard, and a camera.
5. The video-based real-time multimodal digital human AIGC interaction technology according to claim 4, characterized in that: In the multimodal interaction step 2, the image acquisition is performed by calling an external camera to acquire a series of image frames, and then these image frames are input into the visual macro model.
6. The video-based real-time multimodal digital human AIGC interaction technology according to claim 5, characterized in that: In the multimodal interaction step 3, the output result of the large model will enter the streaming speech synthesis module for streaming speech output, and the output result will enter the digital human rendering module for real-time digital human rendering output, and finally the rendering result will be output to the output device in real time.
7. The video-based real-time multimodal digital human AIGC interaction technology according to claim 6, characterized in that: The output devices include displays, AR glasses, etc.