Audio and video joint encoding and decoding method and system based on generative artificial intelligence
By employing a generative artificial intelligence-based audio-video joint coding method, and utilizing cross-modal attention and dynamic adaptive weight allocation techniques, the challenges of traditional audio-video encoding and decoding technologies in terms of high-quality and personalized generation are addressed, enabling efficient and personalized audio-video content generation.
Patent Information
- Application Number
- CN202411787828.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-06
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2044-12-06
AI Technical Summary
Existing audio and video encoding and decoding technologies cannot meet diverse needs, especially in the challenge of generating high-quality and personalized audio and video content.
We employ a generative artificial intelligence-based audio-video joint coding method. By combining cross-modal attention mechanisms and dynamic adaptive weight allocation with techniques such as sentiment analysis, sound event detection, timbre recognition, and scene transition detection, we extract multimodal features from audio and video signals, and then fuse and decode them to ensure the consistency and high quality of the generated audio and video content.
It improves the efficiency and quality of audio and video encoding and decoding, can adjust the style and quality according to user needs, adapt to different task scenarios, enhances the robustness and generalization ability of the model, and realizes high-quality personalized audio and video generation.
Smart Images

Figure CN119583873B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of audio and video generation, and particularly to audio and video joint coding and decoding technology based on generative artificial intelligence. BACKGROUND
[0002] With the rapid development of multimedia technology, the demand for audio and video content generation and processing is increasing. The development of new audio and video generation technology provides a new solution for audio and video coding. Especially in the application scenarios that require high-quality audio and video content generation and real-time processing, these new technologies have shown great potential. At present, with the rapid development of artificial intelligence technology, speech and video recognition and generation technology based on artificial intelligence has been very mature, which lays a solid technical foundation for generative artificial intelligence audio and video joint coding. Through the use of generative artificial intelligence technology, efficient and accurate audio and video coding and decoding can be realized to meet higher demands.
[0003] In the process of audio and video content recognition and generation, AI generation technology can maintain high accuracy in multiple languages, multiple emotions and complex environments. Through multi-modal data fusion technology, high-level features of audio and video are extracted and fused to generate a comprehensive feature vector, ensuring the consistency and coherence of the generated audio and video content in vision and hearing. The progress of these technologies not only improves the efficiency and quality of audio and video coding, but also provides new technical means for multimedia content generation and editing, with wide application prospects.
[0004] Generative artificial intelligence technology can generate high-quality audio and video content by learning a large amount of audio and video data. Compared with traditional coding technology, generative artificial intelligence technology has higher efficiency and accuracy in processing complex scenes and multi-modal data. Through multi-modal data fusion technology, generative artificial intelligence can extract and fuse high-level features of audio and video to generate a comprehensive feature vector, thereby realizing efficient audio and video coding and decoding.
[0005] With the further development of generation technology, the demand for customized audio and video generation systems is also increasing. Modern users want to generate high-quality, personalized audio and video content according to their personal preferences and needs. This demand is not limited to entertainment and social media, but also extends to education, medicine, advertising and corporate training. In the entertainment industry, users want to generate customized music videos, movie clips and game scenes to meet personalized viewing and interactive experiences. In the education field, teachers and students want to generate customized teaching videos and interactive courseware to improve learning effectiveness and participation. In the medical field, doctors and patients want to generate customized medical images and surgical simulations to improve diagnostic and treatment accuracy. In the advertising and corporate training field, businesses want to generate customized promotional videos and training materials to improve brand influence and employee skills. The development of generation technology makes these customized needs possible. By using generative artificial intelligence technology, audio and video content that meets user needs can be efficiently and accurately generated. This not only improves the efficiency of audio and video content generation and processing, but also provides new technical means for the fusion and analysis of multi-modal data, ensuring the consistency and coherence of the generated audio and video content in terms of vision and hearing. In summary, with the continuous progress of generation technology, the demand for customized audio and video generation systems will become increasingly strong, which will drive the further development and application of related technologies, bringing more innovation and possibilities.
[0006] However, current audio and video encoding and decoding technologies face serious challenges in the face of today's digital media explosion. Traditional audio and video encoding methods usually process audio and video separately, and this independent encoding method has been unable to meet the growing demand for diversity. SUMMARY
[0007] The technical problem to be solved by the present application is how to improve the efficiency and quality of audio and video transmission.
[0008] The present application solves the above technical problems by the following technical means: an audio and video joint encoding method based on generative artificial intelligence, comprising the following steps:
[0009] S9, extracting various modal features from the audio signal and the video;
[0010] S10, performing fusion in the cross-modal attention, comprising:
[0011] S101, first, input the extracted video features into the cross-modal attention;
[0012] S102, input the audio features into the cross-modal attention through the auxiliary residual auxiliary network, and analyze the features of different modalities through the cross-modal attention mechanism;
[0013] S103, the cross-modal attention mechanism of the audio feature is the same as S101 and S102; the features of different modalities are connected with the output results through the cross-modal attention;
[0014] S11, task recognition, first identify the current task type, determine the task type, the system will analyze the specific requirements of the task;
[0015] S12, dynamic adaptive weight distribution;
[0016] S13, fusion features, the multi-modal features with different weights are fused.
[0017] As a further optimized technical solution, the step S12, dynamic adaptive weight distribution, specifically includes:
[0018] S121, feature importance evaluation: according to the task requirements, the system evaluates the importance of audio and video features, determines the importance weight of each feature by calculating the contribution of each feature in the current task;
[0019] S122, weight adjustment strategy: the system dynamically adjusts the weights of audio and video features according to the feature importance evaluation results;
[0020] S123, adaptive weight distribution: the system monitors the progress of the task and the performance of the features in real time during the execution of each task, and dynamically adjusts the weight distribution;
[0021] S124, modal data loss processing: in the case of data loss or absence of a certain modality, the system generates data for the missing modality using existing modal data.
[0022] As a further optimized technical solution, various modal features are extracted from the audio signal, including:
[0023] S1, emotion analysis, including:
[0024] Audio data preprocessing: including noise reduction, normalization; emotion state recognition: using emotion analysis algorithm, emotion state recognition is performed on the preprocessed audio data, and the emotion state in the audio is classified;
[0025] S2, sound event detection, including:
[0026] Sound event recognition: sound event detection is performed on the audio data, and various sound events in the audio are recognized;
[0027] Event labeling: detailed labeling is performed on the detected sound events, including the starting time, duration and type of the event;
[0028] S3, timbre recognition, including:
[0029] Timbre judgment: through timbre recognition technology, judge whether the audio segment is of the same timbre, if different timbre is detected, send the original voice segment for a predetermined time;
[0030] Timbre consistent processing: if the timbre is consistent, perform voice recognition and extract voice content;
[0031] Time labeling, time labeling of voice;
[0032] S4, audio feature extraction, including feature parameter extraction and feature integration, integrating time labeling, sound event and emotion analysis into voice features to form a comprehensive audio feature dataset;
[0033] S5, data sending, sending the original voice segment and the audio feature dataset after feature fusion.
[0034] As a further optimized technical solution, various modal features are extracted from the video signal, including:
[0035] S6, scene switching detection
[0036] S61, for the Nth frame extracted in the video, the scene discriminator compares it with the previous frame, analyzes the content of the Nth frame and the N-1th frame, and judges whether there is a scene switching;
[0037] S62, if the scene switching is detected, the frame is set as a key frame, the Nth frame is marked as a key frame, and the frame number is initialized to 1;
[0038] S63, feature extraction and feature fusion are performed on the key frame to obtain the features of the frame, various video features are extracted from the key frame; the extracted features are preprocessed to generate processed features;
[0039] S7, set key frame regularly
[0040] S71, if no scene switching is detected, the video encoder increments the frame number by 1, and the current frame number N is incremented by 1;
[0041] S72, multiple check: judge whether the current frame number is an integer multiple of GOP; if the frame number is an integer multiple of GOP, set the frame as a key frame, and reset the frame number to 1;
[0042] S73, feature extraction and feature preprocessing are performed on the key frame, various video features are extracted from the key frame, and the extracted features are preprocessed to generate processed features;
[0043] S8, inter-frame motion vector calculation
[0044] S81, for the case that the frame number is not an integer multiple of the GOP, the video feature extractor calculates the inter-frame motion vector between the frame and the previous frame using the optical flow method;
[0045] S82, in combination with the features of the previous frame, the features of the Nth frame are generated, and the features of the Nth frame are generated using the features of the previous frame and the calculated motion vector.
[0046] The application also provides an audio-video joint decoding method based on generative artificial intelligence, comprising the following steps:
[0047] S15, feature decoding and preprocessing
[0048] The receiving end first needs to preprocess the multi-modal features transmitted from the sending end after being processed by the audio-video joint encoding method of any one of claims 1-4, the preprocessing steps including normalization, denoising, interpolation, and feature alignment and synchronization, the feature alignment including time alignment and space alignment, and then decoding the preprocessed multi-modal features;
[0049] S16, joint decoding
[0050] The multi-modal features of audio and video decoded in step S15 are input into the joint decoding module, which processes the decoded audio and video features through an AI-driven multi-modal deep learning model. This module combines the audio and video generation module to generate the final multi-modal content through feature fusion and deep network: Video generation module: based on the encoder-decoder architecture of deep learning, used to generate video frames from video features. This module extracts key features of the video from the input features through AI video generation technology, especially the Transformer encoder, to help generate more natural and smooth visual content. Audio generation module: AI timbre encoder is used to extract and fuse the weighted multi-modal features and the timbre features of the original audio. The fused features are converted into semantic token sequences through autoregressive Transformer, and the audio content with high consistency and clarity is gradually generated. Finally, the AI synthesis algorithm HiFNet is used to ensure that the generated audio meets the expected sound quality and semantics;
[0051] S17, merge audio and video, and use AI-driven synchronization technology to ensure that audio and video are aligned on the time axis.
[0052] As a further optimized technical solution, the step S16, joint decoding specifically comprises:
[0053] S161, the video decoding module first performs word embedding to convert the input video features into vector representation, and then performs position encoding on the vector;
[0054] S162, the audio decoding module converts the fused features into semantic tokens and generates the final audio content based on AI-driven audio generation technology. It includes: generating timbre features, fusing multi-modal features with the timbre features of the original audio through an AI timbre encoder to ensure the consistency and richness of audio generation; autoregressive Transformer is used to convert input features into semantic token sequences, and the AI model ensures the coherence and high quality of the audio content through the recursive generation process in this stage; the semantic token sequence generated by the autoregressive Transformer is input into the stream matching model, and the stream matching model is used to convert the generated semantic token sequence into a mel spectrum graph; HiFNet is based on AI spectrum conversion technology, which converts the mel spectrum graph into a high-fidelity audio waveform. The model uses the optimization algorithm of the spectrum feature and waveform feature generation process to ensure that the output audio conforms to the effect of natural sound production.
[0055] As a further optimized technical solution, the step S17, after merging the audio and video, further includes:
[0056] S18, objective evaluation: an AI-generated content evaluation model is used to evaluate the quality of the audio and video, and the evaluation indicators include: peak signal-to-noise ratio, the higher the peak signal-to-noise ratio value, the better the quality of the generated video; structural similarity index, the closer the structural similarity index value is to 1, the more similar the structure of the generated video is to the reference video; audio quality evaluation: using perceptual speech quality evaluation or speech transmission index indicators to evaluate the quality of the generated audio, if the objective evaluation result is qualified, the video is directly output, if not, the weight adjustment module is entered, and the weight distribution of the multi-modal features is dynamically adjusted until a qualified video is generated;
[0057] S19, dynamic adjustment.
[0058] As a further optimized technical solution, the step S19, dynamic adjustment includes:
[0059] S191, weight adjustment module: the AI model dynamically adjusts the weight distribution of the audio and video features using an optimization algorithm (such as gradient descent or genetic algorithm) to improve the quality of the content;
[0060] S192, regenerate content: after each weight adjustment, the new multi-modal features regenerate the audio and video content through the joint decoding module, and the AI model evaluates the output quality again until it meets the expected standard.
[0061] The application also provides an audio and video coding and decoding method based on generative artificial intelligence, including the following steps:
[0062] Audio feature extraction;
[0063] Video feature extraction;
[0064] Audio-video joint encoding, using any of the above audio-video joint encoding methods to jointly encode the extracted audio features and video features;
[0065] Audio-video joint decoding, using any of the above audio-video joint decoding methods to jointly decode the received audio features and video features.
[0066] The application also provides an audio-video encoding system based on generative artificial intelligence, comprising an audio feature extractor, a video feature extractor, an audio-video joint encoder, and an audio-video joint decoder, wherein the audio-video joint encoder uses any of the above audio-video joint encoding methods to jointly encode the extracted audio features and video features; the audio-video joint decoder uses any of the above audio-video joint decoding methods to jointly decode the received audio features and video features.
[0067] The advantages of the application are as follows: the audio-video feature extraction, video feature extraction and AI-based generation technology are first applied to the joint encoding of audio and video, which is a new innovation in the audio coding architecture. It mainly solves the problems of low compression efficiency and inability to pursue higher quality in the traditional method of video coding. At the same time, the generative coding method can flexibly adjust the style and quality according to the personal needs of users. Moreover, the generative audio-video joint coding features can fully utilize the features of audio and video to complement different modal features, improve compression efficiency, and help generate more realistic original videos (decoding) in the process. It opens up new possibilities for future multimedia applications, especially in meeting personalized needs, which shows great potential. Specifically, it includes:
[0068] Firstly, the audio feature extractor process optimization proposes a high-efficiency audio feature extraction method, which can comprehensively analyze and process audio data. Through steps such as sentiment analysis, sound event detection, timbre recognition and speech recognition, the accuracy and efficiency of audio feature extraction are improved, and strong technical support is provided for audio data analysis in multiple fields.
[0069] Secondly, the video feature extractor process optimization improves the efficiency and quality of video coding through scene change detection, key frame setting and optical flow calculation technology. This method sets key frames regularly to ensure the reasonable distribution of key frames in the video coding process, improving the compression efficiency and decoding quality of the video. This method not only applies to regular video coding, but also can be applied to high frame rate and high resolution video coding scenarios, with wide application prospect and market value. Through optimization, the pressure on joint coding and the generation of generative artificial intelligence is reduced, thereby improving the overall coding efficiency.
[0070] Furthermore, a novel multimodal fusion framework is proposed for multi-task multimodal joint encoding, aiming to dynamically balance and optimize the contribution of each modality. Through a cross-modal attention mechanism and a residual auxiliary network, the model's prediction accuracy and robustness are improved. This framework can generate other modalities from the existing modality when other modalities are missing or absent, ensuring data integrity and consistency. The dynamic weighting strategy adapts to different task scenarios, enhancing the model's generalization ability and resulting in better performance across various tasks. Specifically, when data for a certain modality is missing, the system can generate data for the missing modality using existing modality data, ensuring overall data integrity. In addition, the dynamically adaptive weight adjustment mechanism automatically adjusts the weights of each modality according to the needs of different tasks, enabling the model to perform well across various tasks.
[0071] Finally, the generative AI-based joint audio-video decoder is an advanced technology that achieves high-quality audio and video generation through the decoding and fusion of multimodal features. First, this technology decodes the multimodal features transmitted from the sender and performs normalization, denoising, and interpolation to improve data quality. Next, it ensures temporal and spatial alignment of audio and video features, avoiding desynchronization issues. Then, the decoded audio and video multimodal features are input into the joint decoding module. The video generation module generates video through word embedding, positional encoding, and multi-head self-attention mechanisms, while the audio generation module fuses the weighted multimodal features with the timbre features extracted from the original audio using a timbre encoder, and then uses an autoregressive Transformer for stream matching and HiFNet to generate high-quality audio. Finally, this technology ensures that audio and video are aligned on the timeline using timestamps or other synchronization techniques. Through multiple iterations and optimizations, the quality of the generated content is continuously improved, ultimately achieving high-quality audio and video generation. This technology not only improves the quality and consistency of audio and video generation but also provides an effective solution for the fusion of multimodal features. Attached Figure Description
[0072] Figure 1 This is a general framework diagram of the audio-visual fusion method based on generative artificial intelligence according to an embodiment of the present invention;
[0073] Figure 2 This is a flowchart of the audio encoding process in an embodiment of the present invention;
[0074] Figure 3 This is a flowchart of the video encoding process in an embodiment of the present invention;
[0075] Figure 4 This is a diagram of the audio and video joint coding module in an embodiment of the present invention;
[0076] Figure 5This is a diagram of the audio and video joint decoding module in an embodiment of the present invention. Detailed Implementation
[0077] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0078] Please see Figure 1 This invention relates to audio and video coding technologies, aiming to significantly improve the efficiency and quality of data processing. By optimizing the audio feature extractor process, video feature extractor process, multi-task multimodal joint coding, and the design of the audio-video joint decoder, it provides comprehensive technical support for the analysis and processing of multimedia data. Specifically, the optimized audio feature extractor process proposes an efficient audio feature extraction method capable of comprehensively analyzing and processing audio data; the optimized video feature extractor process improves the efficiency and quality of video coding through techniques such as scene transition detection, keyframe setting, and optical flow calculation; the audio-video joint coding proposes an innovative multimodal fusion framework aimed at dynamically balancing and optimizing the contribution of each modality; and the audio-video joint decoder ensures the synchronization and consistency of audio and video signals through feature fusion and time alignment techniques.
[0079] The audio feature extractor process includes the following steps: First, sentiment analysis is performed on the audio to identify and classify emotional states. Next, sound event detection is performed, and the detected events are annotated in detail for subsequent processing. Then, timbre recognition technology is used to determine if audio segments share the same timbre. If different timbres are detected, a 3-second original audio segment is sent for later timbre reconstruction. If the timbres are consistent, speech recognition is performed to extract the speech content. Based on this, time annotation is applied to the speech to ensure the accuracy of the time information for each segment. Subsequently, speech features, including pitch, frequency, rhythm, and other parameters, are extracted. The time annotations, sound events, and sentiment analysis results are integrated into the speech features to form a comprehensive audio feature dataset. Finally, multimodal feature processing is performed to comprehensively analyze the audio features, providing more accurate and comprehensive audio data analysis results. This method not only improves the accuracy and efficiency of audio feature extraction but also provides strong technical support for audio data analysis in multiple fields.
[0080] See Figure 2 The specific steps for optimizing the audio feature extractor process include:
[0081] S1. Sentiment analysis, specifically including:
[0082] Audio data preprocessing: including noise reduction, normalization, and other operations to ensure the quality of audio data;
[0083] Emotional state recognition: Using sentiment analysis algorithms, the preprocessed audio data is used to identify the emotional state, such as happiness, sadness, anger, etc.
[0084] S2, Sound Event Detection, specifically including:
[0085] Sound event recognition: Detects sound events in audio data and identifies various sound events in the audio, such as clapping, laughter, crying, etc.
[0086] Event labeling: Detected sound events are labeled in detail, including the start time, duration, and type of the event, for subsequent processing;
[0087] S3, timbre recognition, specifically includes:
[0088] Phonogram identification: Using phonogram recognition technology, it determines whether audio segments have the same phonogram. If different phonograms are detected, a 3-second original speech segment is sent so that the phonogram can be restored later based on the original speech.
[0089] Voice consistency processing: If the voices are consistent, then speech recognition is performed to extract the speech content;
[0090] Time annotation: Time annotation is performed on the speech to ensure the accuracy of the time information for each speech segment;
[0091] S4, Audio Feature Extraction
[0092] S41. Feature parameter extraction, including:
[0093] Frequency feature extraction: Extracting frequency features from audio signals and analyzing the frequency distribution of the audio;
[0094] Mel frequency extraction: Extract Mel frequencies and perform audio feature analysis using Mel frequency cepstral coefficients (MFCC);
[0095] Phoneme feature extraction: Identifying phoneme features in audio and analyzing the distribution and changes of phonemes;
[0096] Energy Feature Extraction: Extracting the energy features of audio signals and analyzing the energy changes in audio;
[0097] Pitch feature extraction: Extracting pitch features from audio and analyzing pitch variations;
[0098] Specific sound pattern recognition: Identifying specific sound patterns in audio, such as laughter, crying, etc.
[0099] Volume feature extraction: Extracting the volume features of audio and analyzing volume changes;
[0100] Speech rate feature extraction: Extract speech rate features from audio and analyze changes in speech rate;
[0101] S42. Feature Integration: Integrating time stamping, sound events, and sentiment analysis into speech features to form a comprehensive audio feature dataset;
[0102] S5. Data transmission: Send the 3-second original speech segment and the audio feature dataset after feature fusion.
[0103] Through the above steps, the audio feature extraction method of the present invention not only improves the accuracy and efficiency of audio feature extraction, but also provides strong technical support for audio data analysis in multiple fields.
[0104] The video feature extractor process includes the following steps: First, for the new Nth frame, the video encoder compares it with the previous frame to determine if a scene change has occurred. If a scene change is detected, the frame is set as a keyframe, and feature extraction and fusion are performed to obtain the frame's features. Setting keyframes helps maintain video continuity and consistency during scene changes, ensuring video quality. If no scene change has occurred, the video encoder increments the frame number and checks if it is an integer multiple of the Group of Pictures (GOP). If the frame number is an integer multiple of the GOP, the frame is set as a keyframe, and feature extraction and fusion are performed. This method ensures a reasonable distribution of keyframes during video encoding by periodically setting keyframes, improving video compression efficiency and decoding quality. For cases where the frame number is not an integer multiple of the GOP, the video encoder uses optical flow to calculate the inter-frame motion vector between the current frame and the previous frame to obtain a motion description. Then, combined with the features of the previous frame, the features of the Nth frame are generated. Optical flow, by analyzing inter-frame motion information, can accurately describe inter-frame changes, thereby improving the accuracy and efficiency of video encoding. Through the process optimization described above, the video encoder can improve encoding efficiency and compression ratio while ensuring video quality. This method is not only applicable to conventional video encoding but also to high frame rate and high resolution video encoding scenarios, demonstrating broad application prospects and market value.
[0105] See Figure 3 The video feature extractor process specifically includes the following steps:
[0106] S6, Scene Switching Detection
[0107] S61. For the Nth frame extracted from the video, the scene discriminator compares it with the previous frame, analyzes the content of the Nth frame and the (N-1)th frame, and determines whether there is a scene switch. By comparing the features such as color, texture and edge between frames, the scene discriminator can accurately identify the scene switch.
[0108] S62. If a scene change is detected, set the frame as a key frame, mark the Nth frame as a key frame, and initialize the frame number to 1 to ensure the continuity and consistency of the video during scene changes. Setting key frames helps to maintain the smoothness and visual coherence of the video during scene changes.
[0109] S63. Perform feature extraction and feature fusion on keyframes to obtain the features of the frame. Extract various video features from the keyframes, such as color, texture, and edges. Preprocess the extracted features to generate processed features. These features will be used in subsequent encoding and compression processes to improve video quality. The extracted video features include:
[0110] Text feature extraction: Extracting text features from videos and analyzing text information in videos;
[0111] Facial and lip information extraction: Extract facial and lip information from the video and analyze facial expressions and lip changes;
[0112] Scene understanding: Understanding the scenes in the video and analyzing the changes in the scenes;
[0113] Edge detection: Perform edge detection to extract edge features from the video;
[0114] Contextual information understanding: Extracting contextual information from videos and analyzing the contextual relationships within the videos;
[0115] Action and pose recognition: Identify actions and poses in videos and analyze changes in actions and poses;
[0116] Light and shadow information extraction: Extract light and shadow information from the video and analyze the changes in light and shadow;
[0117] Object detection: Perform object detection to identify objects in the video;
[0118] Color and texture feature extraction: Extract color and texture features from videos and analyze changes in color and texture;
[0119] S7. Set keyframes periodically.
[0120] S71. If no scene change is detected, the video encoder increments the frame count by 1, increments the current frame count N by 1, and updates the frame count. This ensures that each frame is correctly counted and processed.
[0121] S72, Multiple Check: Determines whether the current frame number is an integer multiple of the GOP; if the frame number is an integer multiple of the GOP, then sets the frame as a keyframe and resets the frame number to 1, ensuring a reasonable distribution of keyframes. Regularly setting keyframes helps maintain video quality and consistency during long-term video encoding.
[0122] S73. Perform feature extraction and feature preprocessing on keyframes. Extract various video features from keyframes, preprocess the extracted features to generate processed features, which will be used in subsequent encoding and compression processes to improve video quality.
[0123] S8, Inter-frame motion vector calculation
[0124] S81. For cases where the number of frames is not an integer multiple of the GOP, the video feature extractor uses optical flow to calculate the inter-frame motion vector between the current frame and the previous frame. Optical flow calculation: By analyzing the motion information between the Nth frame and the (N-1)th frame using optical flow, the motion vector is obtained. Optical flow can accurately describe the motion changes between frames, thereby improving the accuracy and efficiency of video coding.
[0125] S82. Combine the features of the previous frame to generate the features of the Nth frame. Using the features of the previous frame and the calculated motion vectors, generate the features of the Nth frame. These features will be used in the subsequent encoding and compression process to improve video quality.
[0126] Through the above process optimization, the video encoder can improve encoding efficiency and compression ratio while ensuring video quality. Encoding efficiency: Optimizing the encoding process improves the efficiency of video encoding. Compression ratio: By setting keyframes reasonably and accurately describing inter-frame motion, the compression ratio of the video is improved. The optimized encoding process can significantly reduce the size of video files while maintaining high-quality visual effects.
[0127] This method is applicable to conventional video encoding, high frame rate and high resolution video encoding scenarios. Application scenarios: This method has broad application prospects and market value. It is suitable for a variety of video encoding scenarios. Whether it is conventional video encoding or high frame rate and high resolution video encoding, this method can provide an efficient and high-quality encoding solution.
[0128] The audio-video joint coding method includes: First, extracting various modal features from audio signals and video signals. These modal features are then learned by a residual network and input into a cross-modal attention mechanism. This mechanism adaptively focuses on salient features, improving the efficiency of the overall fusion process and making the model more robust and flexible. In the feature fusion stage, each modal feature and the cross-modal attention feature are fused, fully utilizing information from each modality and improving the overall model performance. The fusion strategy ensures that features from different modalities are complementary, thereby enhancing the model's predictive ability. For each fused feature, dynamic adaptive weighting is applied according to different tasks. The model can dynamically select relevant features to emphasize based on the needs of different tasks to optimize model performance and improve prediction accuracy. This dynamic weighting strategy adapts to different task scenarios and enhances the model's generalization ability. In the modality-level feature weighting stage, features at different modality levels are weighted to ensure that the importance of each modality is fully considered. Finally, the weighted features are fused. This fusion strategy improves the model's robustness and stability, avoiding information loss.
[0129] See Figure 4 The audio and video joint coding method specifically includes the following steps:
[0130] S9, Audio and Video Feature Extraction
[0131] S91. Extract features from audio: frequency features, Mel frequency, phoneme features, energy features, pitch features, specific sound patterns, and volume features;
[0132] S92. Extract features from video: text features, facial and lip information, scene understanding, edge detection, contextual information understanding, action and pose recognition, light and shadow information extraction, object detection, color and texture features;
[0133] S10. Fusion in cross-modal attention
[0134] S101. First, the extracted video features are input into the cross-modal attention;
[0135] S102. The audio features are fed into the cross-modal attention through an auxiliary residual network, and the features of different modalities are analyzed through the cross-modal attention mechanism.
[0136] S103, the cross-modal attention mechanism for audio features is the same as S101 and S102; it connects features from different modalities with the output results of cross-modal attention. This concatenation not only preserves the original features but also enhances the model's ability to focus on salient features, thereby improving the overall fusion efficiency.
[0137] Among them, the residual auxiliary network extracts deep features from the features extracted from audio and video. Because the residual auxiliary network has a deep network structure and residual connections, it can effectively capture and represent complex patterns and features.
[0138] S11, Task Recognition
[0139] S111, Task Type Determination: First, the system needs to identify the current task type. These task types may include emotion expression and reproduction, action recognition and reconstruction, scene understanding and reconstruction, subtitle generation and synchronization, speech separation, abnormal behavior detection, video summarization generation, etc. Through the task identification module, the system can automatically determine the specific type of the current task.
[0140] S112. Task Requirements Analysis: Once the task type is determined, the system will analyze the specific requirements of the task. For example, for the emotion expression and reproduction task, the system needs to focus on the emotional features in the audio and the facial expression features in the video; for the action recognition and reconstruction task, the system needs to focus on the action and posture features in the video. Through task requirements analysis, the system can determine the types of features that need to be emphasized.
[0141] S12, Dynamic Adaptive Weight Allocation
[0142] S121. Feature Importance Assessment: Based on task requirements, the system assesses the importance of audio and video features. By calculating the contribution of each feature to the current task, the system can determine the importance weight of each feature. For example, in an emotion expression task, the emotion features in audio may be more important than the color features in video.
[0143] S122. Weight Adjustment Strategy: The system dynamically adjusts the weights of audio and video features based on the feature importance evaluation results. Specific steps include: first, calculating the initial weight of each feature; then, adjusting the weights of each feature according to task requirements and feature importance, so that important features occupy a larger proportion in the fusion process.
[0144] S123. Adaptive weight allocation: During the execution of each task, the system monitors the progress of the task and the performance of the features in real time and dynamically adjusts the weight allocation. For example, if the contribution of certain features changes during the execution of a task, the system will adaptively adjust the weights of these features to ensure that the model can always use the most relevant features for prediction.
[0145] S124. Modal Data Loss Handling: When data for a certain modality is lost or does not exist, the system can generate data for the missing modality using existing modal data. The specific steps include: First, identifying the data type of the lost modality; then, using the features of other modalities, generating data for the missing modality through a generative model (such as a generative adversarial network or variational autoencoder). This method ensures the integrity and consistency of the data and improves the robustness of the model in the case of incomplete data.
[0146] By dynamically and adaptively allocating weights, the system can selectively emphasize relevant features and optimize model performance. This adaptability not only improves the model's prediction accuracy but also enhances its robustness and generalization ability across different task scenarios.
[0147] S13, Fusion Features
[0148] Multimodal features with different weights are fused together, and the fusion method of each modality feature is flexibly adjusted according to task requirements and data characteristics.
[0149] The audio-video joint decoder includes: First, decoding the multimodal features transmitted from the sending end (audio-video joint encoder) and performing normalization, denoising, and interpolation to improve data quality; next, ensuring temporal and spatial alignment of audio and video features to avoid synchronization issues; then, inputting the decoded audio and video multimodal features into the joint decoding module. The video generation module generates video through word embedding, positional encoding, and multi-head self-attention mechanisms, while the audio generation module fuses the weighted multimodal features with the timbre features extracted from the original audio by the timbre encoder, and then performs stream matching through an autoregressive Transformer. High-quality audio is generated using HiFNet. Finally, the audio and video are aligned on the timeline using timestamps or other synchronization techniques. The generated audio and video content needs to be objectively evaluated using standard evaluation metrics (such as PSNR, SSIM, etc.) to assess the quality of the generated audio and video. If the objective evaluation result is satisfactory, the video is directly output. If it is not satisfactory, it enters the weight adjustment module to dynamically adjust the weight allocation of multimodal features until a satisfactory video is generated. Through multiple iterations, the feature weights and generation process are continuously adjusted to gradually improve the quality of the generated content. After each iteration, the weight adjustment and generation results are recorded for analysis and optimization of the generation strategy.
[0150] See Figure 5 The audio and video joint decoder specifically includes the following steps:
[0151] S15, Feature Decoding and Preprocessing
[0152] The receiving end first needs to decode the multimodal features transmitted from the sending end. These features may be transmitted in compressed or encrypted form, thus requiring decoding. During the transmission of features from the sending end to the receiving end, they may be affected by various factors, such as noise, signal attenuation, and compression loss. These factors can lead to a decrease in the quality of the feature data, thereby affecting the accuracy of subsequent processing. Therefore, preprocessing at the receiving end is necessary. The preprocessing steps include:
[0153] Normalization: Scaling feature data to a standard range to eliminate dimensional differences between different features;
[0154] Denoising: Using filters or other techniques to remove noise from feature data and improve signal purity;
[0155] Interpolation: Interpolating missing data points to ensure data integrity;
[0156] Feature alignment and synchronization, including temporal alignment and spatial alignment:
[0157] Timing alignment: Features from different modalities (such as audio and video) may be out of sync in time. This asynchrony can cause problems during multimodal fusion, affecting the final processing results. For example, in subtitle generation and synchronization tasks, timing alignment of audio and video is a critical step to ensure accurate subtitle synchronization.
[0158] Spatial alignment: For image and video features, spatial alignment refers to ensuring that features of different modalities correspond in space. For example, in the fusion of image and depth information, it is necessary to ensure that each pixel in the image can find a corresponding depth value. Spatial alignment can be achieved through geometric transformation, registration and other techniques to ensure the consistency of multimodal features in space.
[0159] S16, Joint Decoding
[0160] In the joint decoding stage, AI-driven multimodal deep learning models process the decoded audio and video features. This module, combined with the audio and video generation modules, generates the final multimodal content through feature fusion and deep networks: The video generation module uses a deep learning-based encoder-decoder architecture to generate video frames from video features. This module employs AI video generation technology, particularly a Transformer encoder, to extract key video features from the input features, helping to generate more natural and fluid visual content. The audio generation module uses an AI-powered timbre encoder to extract and fuse weighted multimodal features with the timbre features of the original audio. The fused features are converted into a semantic token sequence using an autoregressive Transformer, progressively generating audio content with high consistency and clarity. Finally, the audio is synthesized using the AI HiFNet algorithm to ensure that the generated audio meets the expected sound quality and semantics.
[0161] S161, the video decoding module uses the Transformer architecture in AI deep learning models to convert input features into high-bit vectors. The encoder mainly includes:
[0162] Word embedding: Embedding video features into a high-dimensional vector space so that deep models can better process the semantic information in the video;
[0163] Positional Encoding: Adds positional information to video features. Since the Transformer structure itself does not contain temporal information, positional encoding enables the model to understand the position of the features on the timeline.
[0164] Multi-Head Self-Attention: In this stage, the AI model extracts the correlation between video features through multiple attention heads, helping to generate a high-quality attention matrix. Each attention head independently calculates the feature correlation under different viewpoints, and the results are finally concatenated for subsequent video frame generation.
[0165] Feed-Forward Neural Network: Performs non-linear transformations on the features at each location to enhance the model's expressive power;
[0166] Residual Connection and Layer Normalization: Add residual connections after each sub-layer and perform layer normalization to prevent gradient vanishing and exploding.
[0167] The decoder samples from the latent space and generates video frames. A decoder typically consists of multiple deconvolutional layers. During decoding, in addition to video features, audio features are used as conditional input to guide the generation of video frames. The decoder mainly includes a masked multi-head self-attention mechanism, encoder-decoder attention, a feedforward neural network, and residual connections. The masked multi-head self-attention mechanism is similar to the encoder's self-attention mechanism, but it masks future information during computation, ensuring the decoder only sees previously generated content. Encoder-decoder attention calculates the correlation between the decoder's current state and the encoder's output, generating an attention matrix. The feedforward neural network and residual connections perform nonlinear transformations and layer normalization, similar to the encoder.
[0168] S162. The audio decoding module, based on AI-driven audio generation technology, transforms the fused features into semantic tokens and generates the final audio content. The specific steps are as follows:
[0169] Timbre feature fusion: Through an AI timbre encoder, multimodal features are fused with the timbre features of the original audio to ensure the consistency and richness of the generated audio.
[0170] Autoregressive Transformer generates semantic token sequences: The autoregressive Transformer progressively generates each semantic token sequence during the generation process, ensuring that the generated audio semantics conform to the input features. At this stage, the AI model guarantees the coherence and high quality of the audio content through a recursive generation process.
[0171] The stream matching model generates Mel spectrograms: The semantic tokens generated by the Transformer are input into the stream matching model, and the AI processes the generation tasks of multiple languages and styles to convert the token sequence into a Mel spectrogram.
[0172] HiFNet vocoder generates speech waveforms: HiFNet uses AI spectrum conversion technology to convert Mel spectrograms into high-fidelity audio waveforms. The model uses optimized algorithms for the generation process of spectral and waveform features to ensure that the output audio conforms to the effect of natural speech.
[0173] S17. During the merging of audio and video, AI-driven synchronization technology is used to ensure that the audio and video are aligned on the timeline.
[0174] AI synchronization algorithm: Utilizing AI-based timestamp and frame synchronization technology, it ensures the temporal consistency of audio and video. By comparing the characteristic time points of each frame, it ensures precise synchronization between visual and audio content, avoiding desynchronization issues.
[0175] S18. Objective Evaluation: The AI-generated content evaluation model is used to assess audio and video quality, including multiple metrics for both video and audio.
[0176] PSNR (Peak Signal-to-Noise Ratio): Calculated by an AI model, the peak signal-to-noise ratio of the generated video and a reference video to evaluate video sharpness. A higher PSNR value indicates better video generation quality.
[0177] SSIM (Structural Similarity Index): AI-based SSIM evaluates the structural similarity of videos to ensure that the generated video is structurally consistent with the reference video. The closer the SSIM value is to 1, the higher the structural similarity.
[0178] Audio quality assessment: Use metrics such as PESQ (Perceptual Speech Quality Assessment) or STOI (Speech Transmission Index) to assess the quality of audio generated by the AI model;
[0179] S19, Dynamic Adjustment
[0180] The AI-based optimization module dynamically adjusts feature weights based on evaluation results, generating compliant audio and video content through multiple iterations.
[0181] S191, Weight Adjustment Module: The AI model uses optimization algorithms (such as gradient descent or genetic algorithms) to dynamically adjust the weight allocation of audio and video features to improve content quality;
[0182] S192. Regenerating Content: After each weight adjustment, the new multimodal features regenerate audio and video content through the joint decoding module. The AI model then evaluates the output quality again until it meets the expected standards.
[0183] Through multiple iterations, the feature weights and generation process are continuously adjusted to gradually improve the quality of the generated content. After each iteration, the weight adjustment and generation results are recorded in order to analyze and optimize the generation strategy.
[0184] Generative AI can establish deep connections between audio and video data, significantly improving overall coding efficiency by sharing information and features, and generating final audio and video content on demand. This joint coding method leverages the capabilities of generative models to more effectively capture redundant information in audio and video data, thereby achieving higher compression ratios. Generative AI also supports the customization of video content, meaning that users can flexibly adjust the style, effects, and quality of audio and video content according to their individual needs and preferences. This customization capability provides users with greater control, enabling each user to obtain a personalized viewing experience. Simultaneously, the adaptive capabilities of generative AI allow it to flexibly adjust coding strategies based on different content types and scenarios, ensuring that audio and video synchronization is maintained while minimizing quality loss. In conclusion, generative AI-based joint audio and video coding not only improves the efficiency and quality of audio and video coding but also opens up new possibilities for future multimedia applications, especially demonstrating great potential in meeting personalized needs.
[0185] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An audio and video encoding / decoding method based on generative artificial intelligence, characterized in that: Includes the following steps: Audio feature extraction; Video feature extraction; Audio and video joint encoding includes the following steps: S9. Extract various modal features from audio and video signals; S10. Fusion in cross-modal attention, including: S101. First, the extracted video features are input into the cross-modal attention; S102. The video features are fed into a cross-modal attention network through an auxiliary residual network, and the features of different modalities are analyzed through the cross-modal attention mechanism. S103, the cross-modal attention mechanism for audio features is the same as S101 and S102; it connects features from different modalities with the output results obtained through cross-modal attention. S11. Task identification: First, the current task type is identified. Once the task type is determined, the system will analyze the specific requirements of the task. S12, Dynamic Adaptive Weight Allocation, specifically including: S121. Feature Importance Assessment: Based on task requirements, the system assesses the importance of audio and video features and determines the importance weight of each feature by calculating its contribution to the current task. S122. Weight Adjustment Strategy: The system dynamically adjusts the weights of audio and video features based on the feature importance evaluation results. S123, Adaptive weight allocation: During the execution of each task, the system monitors the progress and performance of the task in real time and dynamically adjusts the weight allocation. S124. Modal data loss handling: In the event that the data of a certain modality is lost or does not exist, the system uses the data of the existing modalities to generate the data of the missing modality; S13, Feature Fusion: This involves fusing multimodal features that have been assigned different weights. Joint audio and video decoding, including: S15, Feature Decoding and Preprocessing The receiving end first needs to preprocess the multimodal features transmitted from the sending end after audio and video joint encoding. The preprocessing steps include normalization, noise reduction, interpolation, and feature alignment and synchronization. Feature alignment includes temporal alignment and spatial alignment. Then, the preprocessed multimodal features are decoded. S16, Joint Decoding The multimodal features of the audio and video decoded in step S15 are input into the joint decoding module. This joint decoding uses an AI-driven multimodal deep learning model to process the decoded audio and video features. This module, combined with the audio and video generation module, generates the final multimodal content through feature fusion and deep networks: Video generation module: Based on a deep learning encoder-decoder architecture, it is used to generate video frames from video features. This module uses a Transformer encoder to extract key video features from the input features. Audio generation module: Uses an AI timbre encoder to extract and fuse the weighted multimodal features and the timbre features of the original audio. The fused features are converted into a semantic token sequence through an autoregressive Transformer, gradually generating audio content with high consistency and clarity. Finally, the audio is synthesized using the AI HiFNet algorithm to ensure that the generated audio meets the expected sound quality and semantics. S17. Merge audio and video, using AI-driven synchronization technology to ensure that audio and video are aligned on the timeline; S18. Objective Evaluation: The AI-generated content evaluation model is used to assess the quality of audio and video. Evaluation metrics include: Peak Signal-to-Noise Ratio (PSNR), where a higher PSNR indicates better quality; Structural Similarity Index, where a value closer to 1 indicates a more similar structure between the generated video and the reference video; Audio Quality Evaluation: The quality of the generated audio is evaluated using perceptual speech quality assessment or speech transmission index. If the objective evaluation result is satisfactory, the video is directly output. If it is unsatisfactory, the weight adjustment module dynamically adjusts the weight allocation of multimodal features until a satisfactory video is generated. S19. Dynamic adjustment, including: S191, Weight Adjustment Module: Based on the objective evaluation results, adjust the weight allocation of audio and video features. The AI model uses gradient descent or genetic algorithm to dynamically adjust the weights. S192. Regenerating Content: After each weight adjustment, the new multimodal features regenerate audio and video content through the joint decoding module. The AI model then re-evaluates the output quality until it meets the expected standards.
2. The audio and video encoding and decoding method based on generative artificial intelligence as described in claim 1, characterized in that: Extracting various modal features from audio signals includes: S1. Sentiment analysis, including: Audio data preprocessing: including noise reduction and normalization; Emotional state recognition: using sentiment analysis algorithms to identify the emotional state of the preprocessed audio data and classify the emotional state in the audio. S2, sound event detection, including: Sound event recognition: Detects sound events in audio data and identifies various sound events in the audio. Event labeling: Provides detailed labels for detected sound events, including the start time, duration, and type of the event; S3, timbre recognition, including: Phonogram identification: Using phonogram recognition technology, determine whether audio segments have the same phonogram. If different phonograms are detected, send the original audio segment at a predetermined time. Voice consistency processing: If the voices are consistent, then speech recognition is performed to extract the speech content; Time annotation: Time annotation of speech; S4. Audio feature extraction, including feature parameter extraction and feature integration, integrates time annotation, sound events and sentiment analysis into speech features to form a comprehensive audio feature dataset; S5. Data transmission: Send the original speech segments and the audio feature dataset after feature fusion.
3. The audio and video encoding and decoding method based on generative artificial intelligence as described in claim 1, characterized in that: Extracting various modal features from video signals includes: S6, Scene Switching Detection S61. For the Nth frame extracted from the video, the scene discriminator compares it with the previous frame, analyzes the content of the Nth frame and the (N-1)th frame, and determines whether there is a scene switch. S62. If a scene change is detected, set the frame as a key frame, mark the Nth frame as a key frame, and initialize the frame number to 1. S63. Perform feature extraction and feature fusion on key frames to obtain the features of the frame and extract various video features from the key frames; preprocess the extracted features to generate processed features. S7. Set keyframes periodically. S71. If no scene change is detected, the video encoder increments the frame number by 1 and the current frame number N by 1. S72, Multiple Check: Determine if the current frame number is an integer multiple of the GOP; if the frame number is an integer multiple of the GOP, set the frame as a keyframe and reset the frame number to 1. S73. Perform feature extraction and feature preprocessing on keyframes. Extract various video features from keyframes, preprocess the extracted features, and generate processed features. S8, Inter-frame motion vector calculation S81. For cases where the number of frames is not an integer multiple of the GOP, the video feature extractor uses optical flow to calculate the inter-frame motion vector between the current frame and the previous frame. S82. Combine the features of the previous frame to generate the features of the Nth frame. Use the features of the previous frame and the calculated motion vector to generate the features of the Nth frame.
4. The audio and video encoding and decoding method based on generative artificial intelligence as described in claim 1, characterized in that: Step S16, joint decoding, specifically includes: S161. The video decoding module first performs word embedding, converting the input video features into vector representations, and then performs position encoding on the vectors. S162, the audio decoding module, based on AI-driven audio generation technology, transforms the fused features into semantic tokens and generates the final audio content. This includes: generating timbre features; using an AI timbre encoder to fuse multimodal features with the timbre features of the original audio to ensure the consistency and richness of the generated audio; an autoregressive Transformer is used to convert the input features into a sequence of semantic tokens, and the AI model ensures the coherence and high quality of the audio content through a recursive generation process at this stage; the semantic token sequence generated by the autoregressive Transformer is input into a stream matching model, which converts the generated semantic token sequence into a Mel spectrogram; HiFNet is based on AI spectrum conversion technology to convert the Mel spectrogram into a high-fidelity audio waveform.
5. An audio and video encoding / decoding system based on generative artificial intelligence, characterized in that: This includes an audio feature extractor, a video feature extractor, an audio-video joint encoder, and an audio-video joint decoder. The audio-video joint encoder uses an audio-video joint encoding method to jointly encode the extracted audio and video features. The audio-video joint decoder uses an audio-video joint decoding method to jointly decode the received audio and video features. The audio and video joint coding method includes the following steps: S9. Extract various modal features from audio and video signals; S10. Fusion in cross-modal attention, including: S101. First, the extracted video features are input into the cross-modal attention; S102. The video features are fed into a cross-modal attention network through an auxiliary residual network, and the features of different modalities are analyzed through the cross-modal attention mechanism. S103, the cross-modal attention mechanism for audio features is the same as S101 and S102; it connects features from different modalities with the output results obtained through cross-modal attention. S11. Task identification: First, the current task type is identified. Once the task type is determined, the system will analyze the specific requirements of the task. S12, Dynamic Adaptive Weight Allocation, specifically including: S121. Feature Importance Assessment: Based on task requirements, the system assesses the importance of audio and video features and determines the importance weight of each feature by calculating its contribution to the current task. S122. Weight Adjustment Strategy: The system dynamically adjusts the weights of audio and video features based on the feature importance evaluation results. S123, Adaptive weight allocation: During the execution of each task, the system monitors the progress and performance of the task in real time and dynamically adjusts the weight allocation. S124. Modal data loss handling: In the event that the data of a certain modality is lost or does not exist, the system uses the data of the existing modalities to generate the data of the missing modality; S13, Feature Fusion: This involves fusing multimodal features that have been assigned different weights. Audio and video joint decoding methods include: S15, Feature Decoding and Preprocessing The receiving end first needs to preprocess the multimodal features transmitted from the sending end after audio and video joint encoding. The preprocessing steps include normalization, noise reduction, interpolation, and feature alignment and synchronization. Feature alignment includes temporal alignment and spatial alignment. Then, the preprocessed multimodal features are decoded. S16, Joint Decoding The multimodal features of the audio and video decoded in step S15 are input into the joint decoding module. This joint decoding uses an AI-driven multimodal deep learning model to process the decoded audio and video features. This module, combined with the audio and video generation module, generates the final multimodal content through feature fusion and deep networks: Video generation module: Based on a deep learning encoder-decoder architecture, it is used to generate video frames from video features. This module uses a Transformer encoder to extract key video features from the input features. Audio generation module: Uses an AI timbre encoder to extract and fuse the weighted multimodal features and the timbre features of the original audio. The fused features are converted into a semantic token sequence through an autoregressive Transformer, gradually generating audio content with high consistency and clarity. Finally, the audio is synthesized using the AI HiFNet algorithm to ensure that the generated audio meets the expected sound quality and semantics. S17. Merge audio and video, using AI-driven synchronization technology to ensure that audio and video are aligned on the timeline; S18. Objective Evaluation: The AI-generated content evaluation model is used to assess the quality of audio and video. Evaluation metrics include: Peak Signal-to-Noise Ratio (PSNR), where a higher PSNR indicates better quality; Structural Similarity Index, where a value closer to 1 indicates a more similar structure between the generated video and the reference video; Audio Quality Evaluation: The quality of the generated audio is evaluated using perceptual speech quality assessment or speech transmission index. If the objective evaluation result is satisfactory, the video is directly output. If it is unsatisfactory, the weight adjustment module dynamically adjusts the weight allocation of multimodal features until a satisfactory video is generated. S19. Dynamic adjustment, including: S191, Weight Adjustment Module: Based on the objective evaluation results, adjust the weight allocation of audio and video features. The AI model uses gradient descent or genetic algorithm to dynamically adjust the weights. S192. Regenerating Content: After each weight adjustment, the new multimodal features regenerate audio and video content through the joint decoding module. The AI model then re-evaluates the output quality until it meets the expected standards.
6. The audio and video encoding and decoding system based on generative artificial intelligence as described in claim 5, characterized in that: Extracting various modal features from audio signals includes: S1. Sentiment analysis, including: Audio data preprocessing: including noise reduction and normalization; Emotional state recognition: using sentiment analysis algorithms to identify the emotional state of the preprocessed audio data and classify the emotional state in the audio. S2, sound event detection, including: Sound event recognition: Detects sound events in audio data and identifies various sound events in the audio. Event labeling: Provides detailed labels for detected sound events, including the start time, duration, and type of the event; S3, timbre recognition, including: Phonogram identification: Using phonogram recognition technology, determine whether audio segments have the same phonogram. If different phonograms are detected, send the original audio segment at a predetermined time. Voice consistency processing: If the voices are consistent, then speech recognition is performed to extract the speech content; Time annotation: Time annotation of speech; S4. Audio feature extraction, including feature parameter extraction and feature integration, integrates time annotation, sound events and sentiment analysis into speech features to form a comprehensive audio feature dataset; S5. Data transmission: Send the original speech segments and the audio feature dataset after feature fusion.
7. The audio and video encoding and decoding system based on generative artificial intelligence as described in claim 5, characterized in that: Extracting various modal features from video signals includes: S6, Scene Switching Detection S61. For the Nth frame extracted from the video, the scene discriminator compares it with the previous frame, analyzes the content of the Nth frame and the (N-1)th frame, and determines whether there is a scene switch. S62. If a scene change is detected, set the frame as a key frame, mark the Nth frame as a key frame, and initialize the frame number to 1. S63. Perform feature extraction and feature fusion on key frames to obtain the features of the frame and extract various video features from the key frames; preprocess the extracted features to generate processed features. S7. Set keyframes periodically. S71. If no scene change is detected, the video encoder increments the frame number by 1 and the current frame number N by 1. S72, Multiple Check: Determine if the current frame number is an integer multiple of the GOP; if the frame number is an integer multiple of the GOP, set the frame as a keyframe and reset the frame number to 1. S73. Perform feature extraction and feature preprocessing on keyframes. Extract various video features from keyframes, preprocess the extracted features, and generate processed features. S8, Inter-frame motion vector calculation S81. For cases where the number of frames is not an integer multiple of the GOP, the video feature extractor uses optical flow to calculate the inter-frame motion vector between the current frame and the previous frame. S82. Combine the features of the previous frame to generate the features of the Nth frame. Use the features of the previous frame and the calculated motion vector to generate the features of the Nth frame.
8. The audio and video encoding and decoding system based on generative artificial intelligence as described in claim 5, characterized in that: Step S16, joint decoding, specifically includes: S161. The video decoding module first performs word embedding, converting the input video features into vector representations, and then performs position encoding on the vectors. S162, the audio decoding module, based on AI-driven audio generation technology, transforms the fused features into semantic tokens and generates the final audio content. This includes: generating timbre features; using an AI timbre encoder to fuse multimodal features with the timbre features of the original audio to ensure the consistency and richness of the generated audio; an autoregressive Transformer is used to convert the input features into a sequence of semantic tokens, and the AI model ensures the coherence and high quality of the audio content through a recursive generation process at this stage; the semantic token sequence generated by the autoregressive Transformer is input into a stream matching model, which converts the generated semantic token sequence into a Mel spectrogram; HiFNet is based on AI spectrum conversion technology to convert the Mel spectrogram into a high-fidelity audio waveform.