Real-time video communication method based on AIGC

By extracting audio and keyframes on the sending end and synchronous videos on the receiving end, the problem of high bandwidth and privacy leakage in AIGC video transmission is solved, high-quality video transmission is achieved, adapting to multi-network environments, and improving user experience.

CN120358377APending Publication Date: 2025-07-22RENMIN UNIVERSITY OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510537931.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing AIGC video transmission methods have high bandwidth consumption under harsh network conditions, and there are problems such as leakage of tone privacy and poor video quality.

Method used

By extracting audio data and keyframes on the sending end, synchronous videos are synthesized on the receiving end using the video generation model to avoid tone reconstruction, and combining dynamic keyframe extraction and lip motion synthesis to ensure video quality and privacy protection.

Benefits of technology

Significantly reduce bandwidth consumption, improve video quality, improve user experience, adapt to various network environments, and protect user privacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120358377A_ABST
    Figure CN120358377A_ABST
Patent Text Reader

Abstract

The invention relates to a real-time video communication method based on AIGC, and the method comprises the steps: S1, a transmitting end obtains complete video content to be subjected to video communication, and extracts all audio data and key frames in image frames from the complete video content; s2, the sending end transmits the extracted audio data and the key frame to a receiving end; s3, inputting a preset video generation model by the receiving end based on the received audio data and the key frame, and generating video content of which the voice and the picture are synchronous and the picture content and the key frame reach a set similarity; and S4, the receiving end caches the generated video content, and automatically plays the video content or plays the video content according to an instruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video transmission, and particularly to a real-time video communication method based on AIGC. Background Art

[0002] With the rapid development of Internet technology, video streaming has become an important way for people to obtain information, entertainment, and education.

[0003] Traditional video streaming mainly focuses on transmitting complete raw audio and video data, which usually results in high bandwidth consumption. For example, transmitting a 720p video requires approximately 8 Mbps of bandwidth. Therefore, in harsh network conditions such as high-speed trains, elevators, and underground garages, this traditional mode will cause problems such as latency, stuttering, and quality degradation. These problems seriously affect the user experience and limit the wide application of video streaming.

[0004] Although some video transmission methods based on Artificial Intelligence Generated Content (AIGC) have emerged in the current market, these methods attempt to reduce bandwidth requirements by generating synthetic videos, but there are still several technical bottlenecks.

[0005] The defects of existing AIGC video transmission include: (1) Many existing solutions use audio-to-text methods to reduce the amount of transmitted data and resynthesize audio at the data receiving end. This process involves voice color simulation and audio synthesis, which not only increases the computational cost but also may pose a risk of privacy leakage.

[0006] (2) Some AIGC methods only rely on single-frame images to generate videos, which results in a large difference between the generated videos and the original video content, and the video quality often fails to meet user expectations. Summary of the Invention

[0007] The present invention provides a real-time video communication method based on AIGC, which solves the problems of voice color privacy and video quality while reducing bandwidth consumption. The present invention aims to improve the adaptability and user experience of video streaming in various network environments through innovative design and implementation, and provide a new solution for future video transmission technology.

[0008] To achieve the above object, the present invention adopts the following technical solutions: In a first aspect, the present application provides a real-time video communication method based on AIGC, including: S1. The sending end obtains the complete video content to be video communicated, and extracts all the audio data and the key frames in the image frames from the complete video content; S2. The sender transmits the extracted audio data and key frames to the receiver; S3. The receiver inputs the received audio data and key frames into a preset video generation model to generate video content with synchronized voice and picture, and the picture content reaches a set similarity with the key frames; S4. The receiver caches the generated video content and plays the video content automatically or according to an instruction.

[0009] In one implementation, in S1, key frames in the image frames are extracted from the complete video content based on video quality constraints and bandwidth constraints.

[0010] In one implementation, in the video quality constraint, the structural similarity index SSIM is used to evaluate a frame of image and the video frame sent last time The similarity between them, when and the picture sent in the previous frame The SSIM value between them is lower than the set threshold Then, this video frame Meets the video quality constraint.

[0011] In one implementation, in the setting of bandwidth constraint, the current network bandwidth and the highest bandwidth limit preset by the user are monitored to obtain the minimum average time interval for key frame extraction , To ensure that the transmission bandwidth of video transmission does not exceed the bandwidth limit preset by the user, it is necessary to ensure that the average time interval for the key frame extraction module to finally extract and transmit all key frames is greater than or equal to this minimum average time interval .

[0012] In one implementation, in S3, the receiver uses a model that synthesizes lip movements with audio, inputs the processed image and the original audio into the model, and generates the final video output through the model, and ensures that the lip movements are synchronized with the sound.

[0013] In one implementation, in S4, the generated video is stored in a cache queue, and when the number of generated videos reaches a certain number, the videos in the queue start to be played at a constant speed.

[0014] Due to the above technical solutions, the present invention has the following advantages: The technical solution of the present invention proposes a method for compressing and extracting original video frames, selectively extracting and compressing key frames according to the importance of video content to reduce the amount of transmitted data, adopting dynamic video frame extraction technology to automatically select key frames according to important information in the original video, providing high-quality input for subsequent processing; and designing an audio synthesis process that does not rely on timbre reconstruction, directly combining the original audio with the extracted video frames, thus avoiding involving user biometric information; an efficient video frame synthesis method for synthesizing lip movements based on audio, realizing fast and accurate synthesis through synchronous processing of the extracted key frames and corresponding audio. Description of the Drawings

[0015] Figure 1 It is a schematic diagram of the application scenario of the real-time video communication method based on AIGC in the embodiments of the present invention; Figure 2 It is a schematic diagram of the method flow in the detailed embodiments of the present invention; Figure 3 It is a flowchart of key frame sending in the detailed embodiments of the present invention; Figure 4a and Figure 4b It is an effect diagram in an example. Detailed Embodiments

[0016] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present invention fall within the protection scope of the present invention.

[0017] Aiming at the defects and problems of the prior art, the present application provides a real-time video communication method based on AIGC, based on Figure 1 the schematic diagram, the method includes: S1. The sending end obtains the complete video content to be video communicated, and extracts all the audio data and key frames in the image frames from the complete video content; S2. The sending end transmits the extracted audio data and key frames to the receiving end; S3. The receiving end inputs the received audio data and key frames into a preset video generation model to generate video content with synchronized voice and picture, and the picture content reaches a set similarity with the key frames; S4. The receiving end caches the generated video content and plays the video content automatically or according to an instruction.

[0018] The above method will be described in a more detailed embodiment with reference to more attached drawings. Subsequently, its effects will be illustrated in a numerical example.

[0019] Detailed Embodiment The present invention provides a real-time video communication method based on AIGC, and its method flow can be referred to Figure 2 .

[0020] Specifically, it includes the following steps: Step 1. Bandwidth Budget and Quality Tolerance Collection When the system starts, the user needs to define the bandwidth budget (BW) and the video quality tolerance (α). The bandwidth budget refers to the maximum data transmission rate acceptable in a specific network environment, usually in Kbps. The quality tolerance represents the user's acceptance of video quality and is quantified using the Structural Similarity Index (SSIM), which ranges from 0 to 1, where 1 means exactly the same. The system sets the benchmarks for subsequent operations based on these parameters input by the user and makes dynamic adjustments according to these benchmarks to minimize bandwidth consumption while ensuring video quality.

[0021] Step 2. Dynamically Extract Video Key Frames The key frame extraction module is responsible for dynamically extracting key frames in the video according to real-time network conditions and parameters set by the user. The specific process is as follows: Initial calculation of the picture time interval: According to the bandwidth budget provided by the user, the appropriate picture interval ( ) is calculated using a formula. For each video frame input to the key frame extraction module, the picture interval should satisfy the bandwidth constraint, that is, ensure that the amount of data sent each time does not exceed the set bandwidth limit.

[0022] Judge whether the quality constraint is satisfied: The SSIM index is used to evaluate the similarity between consecutive frames. When the SSIM value between the newly input video frame and the previous key frame sent is lower than the preset threshold, it indicates that and have a large content gap, then the system considers this frame to be a key frame.

[0023] Judge whether the bandwidth constraint is satisfied: Use as an indicator to measure whether the time interval between pictures meets the bandwidth constraint requirements. For the video frame for which it has been judged that the quality constraint is satisfied, calculate the time interval between the input time of the video frame and the transmission time of the previous key frame whether it is greater than or equal to . When and only when When it is considered a key frame It meets both the quality constraint and the bandwidth constraint and can be transmitted.

[0024]

[0025]

[0026] Specifically, in the key frame extraction module, whenever a frame of the source video is received , the time interval between the current time and the image of the previous frame transmitted will be obtained. Only when and , the key frame extraction module determines that the current video frame meets the bandwidth constraint (Bandwidth Constraint). When and only when a frame of the source video meets both the bandwidth constraint and the video quality constraint, the key frame extraction module transmits the key image of this frame. The schematic diagram of this step can be referred to Figure 3 .

[0027] Step 3. Transmission of original audio data Audio data transmission: When the key frame extraction module transmits a key frame , the audio transmission module will directly forward a segment of the original audio corresponding to this frame from the sender to the receiver without any processing, in order to reduce latency and protect user privacy.

[0028] Step 4. Video synthesis and playback At the receiver, the video generation module is responsible for combining the received key video frames with the original audio to achieve the final video synthesis. The specific steps are as follows: Input processing: The receiver first receives the important images and original audio data from the sender for subsequent processing.

[0029] Synchronization processing: Ensure synchronization between the received video frames and audio signals, and align them through timestamps or other methods.

[0030] Generate the final video: Use the method model of audio synthesizing lip movement to combine the synchronized images and audio to generate the final video.

[0031] Cache mechanism playback: The video synthesized by the method model of audio synthesizing lip movement is first temporarily stored in the cache queue. When the number of synthesized videos in the cache queue reaches a certain amount, the videos in the cache queue will be played smoothly to ensure the coherence of the video.

[0032] Step 5. Continuous monitoring and optimization Continuous performance monitoring: The system continuously monitors network dynamics during each round of transmission to ensure the system operates efficiently in a new environment.

[0033] Iterative optimization adjustment: Continuously adjust parameters according to the real-time monitored data to adapt to the changing network conditions.

[0034] The quantization effect of the present invention is given in a numerical example of an application below.

[0035] Numerical example: The performance comparison between the present invention and several baseline methods is evaluated in terms of both bandwidth consumption and video quality (measured by average SSIM). As Figure 4a shown in a and b, compared with the traditional RTC system (2.613 Mbps), the present invention only requires 566 Kbps, reducing the bandwidth requirement by more than 78%. At the same time, the average SSIM value of the present invention is 0.661, which is an improvement compared with the existing model Txt2Vid (0.539). In summary, the present invention demonstrates good performance in maintaining video quality and transmission efficiency, providing a solid foundation for future applications.

[0036] In several embodiments provided by the present invention, it should be understood that the disclosed method can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above division of units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.

[0037] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

Claims

1. A real-time video communication method based on AIGC, characterized in that, Including: S1. The sending end obtains the complete video content to be video - communicated, extracts all the audio data from the complete video content, and the key frames in the image frames; S2. The sending end transmits the extracted audio data and key frames to the receiving end; S3. The receiving end inputs the received audio data and key frames into a preset video generation model to generate a video content with synchronized voice and picture, and the picture content reaches a set similarity with the key frames; S4. The receiving end caches the generated video content and plays the video content automatically or according to an instruction.

2. The real-time video communication method based on AIGC according to claim 1, wherein In S1, based on video quality constraints and bandwidth constraints, key frames in the image frames are extracted from the complete video content.

3. The real-time video communication method based on AIGC according to claim 2, characterized in that, In video quality constraints, the Structural Similarity Index (SSIM) is used to evaluate a frame of an image with the previously transmitted video frame for similarity. When the SSIM value between the current frame and the previously transmitted picture is lower than the set threshold then this video frame meets the video quality constraint.

4. The real-time video communication method based on AIGC according to claim 3, wherein In the bandwidth-constrained setting, by monitoring the current network bandwidth and the highest bandwidth limit preset by the user, the minimum average time interval for key frame extraction can be obtained. To ensure that the transmission bandwidth of video transmission does not exceed the bandwidth limit preset by the user, it is necessary to ensure that the average time interval for the key frame extraction module to finally extract and transmit all key frames is greater than or equal to this minimum average time interval. .

5. The real-time video communication method based on AIGC according to claim 1, characterized in that, In S3, the receiving end uses a model that synthesizes lip movements with audio, inputs the processed image and the original audio into the model, and generates the final video output through the model, ensuring that the lip movements are synchronized with the sound.

6. The real-time video communication method based on AIGC according to claim 1, characterized in that In S4, the generated video is stored in a cache queue. After the number of generated videos reaches a certain amount, the videos in the queue start to be played at a uniform speed.