Method, apparatus, device, and storage medium for processing audio and video data

By performing multi-scale wavelet transformation and nonlinear mapping reconstruction on the audio data, local contrast enhancement and super-resolution fusion reconstruction are carried out on the video data, and optical flow displacement tracking and audio feature matching are used to solve the problem of time synchronization deviation of audio and video data, and efficient matching and collaborative operation of audio and video data is achieved.

CN119277122BActive Publication Date: 2025-06-20ORIGJOY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411242227.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-05
Publication Date
2025-06-20
Estimated Expiration
2044-09-05

AI Technical Summary

Technical Problem

There are often deviations and inconsistencies in time synchronization during the application process of audio and video data, which affects the user experience and brings challenges to subsequent analysis and application.

Method used

Audio data is reconstructed through multi-scale wavelet transformation and nonlinear mapping to construct distortion-corrected audio data; audio data is subject to multi-time window segmentation processing and timing feature evolution fitting to construct audio timing semantic feature maps; local contrast enhancement and super-resolution fusion reconstruction of video data are carried out to generate ultra-definition reconstructed videos; dynamic video frame sequences are constructed using optical flow displacement tracking, and videos are positioned and matched frame by frame and fine-positioned and embedded based on audio features, and audio time offset is calculated and global delay fine-tuned.

Benefits of technology

It realizes accurate matching and alignment of audio and video data, solves the problem of out-of-synchronization of audio and video time, and improves the quality and coordinated operation effect of audio and video data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119277122B_ABST
    Figure CN119277122B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of data processing, and particularly to a method, apparatus, device and storage medium for processing audio and video data. The method includes the following steps: obtaining the audio data to be processed and the video data to be processed; performing multi-scale wavelet transform decomposition on the audio data to be processed, and performing non-linear mapping reconstruction to construct distorted corrected audio data; performing multi-time window segmentation processing on the distorted corrected audio data, and performing temporal feature evolution fitting to construct an audio temporal semantic feature map; performing local contrast enhancement processing on the video data to be processed to generate a locally enhanced video; performing super-resolution fusion reconstruction on the locally enhanced video to construct an ultra-clear reconstructed video; performing frame-by-frame temporal decomposition on the ultra-clear reconstructed video, and performing optical flow displacement tracking between adjacent frames to construct a dynamic video frame sequence. The present invention improves the time synchronization of audio and video data and realizes the optimization of the time deviation delay of audio and video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and particularly to a method, device, equipment and storage medium for processing audio and video data. Background Art

[0002] With the rapid development of multimedia technology, audio and video data have been widely used in various application scenarios, such as smart home, intelligent security, virtual reality, etc. However, in the actual application process, there are often deviations and inconsistencies in time synchronization between audio and video data, which are caused by reasons such as mismatched parameters of audio and video acquisition devices, differences in transmission channels, and imperfect processing processes. This situation of out-of-sync audio and video will seriously affect the user experience and pose challenges to subsequent audio and video analysis and applications.

[0003] Traditional audio and video synchronization correction methods usually rely on manual observation and manual adjustment. This method is inefficient and it is difficult to accurately eliminate time deviation. Therefore, there is an urgent need for an intelligent method for processing audio and video data to achieve automation and optimization of audio and video time synchronization. Summary of the Invention

[0004] To solve the above technical problems, the present invention proposes a method, device, equipment and storage medium for processing audio and video data to solve at least one of the above technical problems.

[0005] To achieve the above object, the present invention provides a method for processing audio and video data, including the following steps:

[0006] Step S1: Obtain the audio data to be processed and the video data to be processed; perform multi-scale wavelet transform decomposition on the audio data to be processed, and perform non-linear mapping reconstruction to construct distortion-corrected audio data;

[0007] Step S2: Perform multi-time window segmentation processing on the distortion-corrected audio data, and perform time series feature evolution fitting to construct an audio time series semantic feature map;

[0008] Step S3: Perform local contrast enhancement processing on the video data to be processed to generate a locally enhanced video; perform super-resolution fusion reconstruction on the locally enhanced video to construct an ultra-clear reconstructed video;

[0009] Step S4: Perform frame-by-frame time series decomposition on the ultra-clear reconstructed video, and perform optical flow displacement tracking between adjacent frames to construct a dynamic video frame sequence;

[0010] Step S5: Perform frame-by-frame positioning and matching on the dynamic video frame sequence based on the audio time series semantic feature map, and perform fine positioning embedding to construct an initial audio-embedded video;

[0011] Step S6: Calculate the audio time offset for the initial audio-embedded video to obtain the audio time delay value; perform global time delay fine-tuning based on the audio time delay value to construct a globally time delay optimized video, thereby completing the processing task of audio-visual data.

[0012] The present invention also provides a processing device for audio-visual data, comprising:

[0013] A non-linear reconstruction module, configured to obtain the audio data to be processed and the video data to be processed; perform multi-scale wavelet transform decomposition on the audio data to be processed, and perform non-linear mapping reconstruction to construct distortion-corrected audio data;

[0014] A semantic feature module, configured to perform multi-time window segmentation processing on the distortion-corrected audio data, and perform temporal feature evolution fitting to construct an audio temporal semantic feature map;

[0015] An image enhancement module, configured to perform local contrast enhancement processing on the video data to be processed to generate a locally enhanced video; perform super-resolution fusion reconstruction on the locally enhanced video to construct an ultra-high definition reconstructed video;

[0016] An optical flow displacement tracking module, configured to perform frame-by-frame temporal decomposition on the ultra-high definition reconstructed video, and perform optical flow displacement tracking between adjacent frames to construct a dynamic video frame sequence;

[0017] A fine positioning embedding module, configured to perform frame-by-frame positioning matching on the dynamic video frame sequence based on the audio temporal semantic feature map, and perform fine positioning embedding to construct an initial audio-embedded video;

[0018] A time delay optimization module, configured to calculate the audio time offset for the initial audio-embedded video to obtain the audio time delay value; perform global time delay fine-tuning based on the audio time delay value to construct a globally time delay optimized video, thereby completing the processing task of audio-visual data.

[0019] The present invention also provides a computer device, comprising a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the steps of the audio-visual data processing method described in any one of the above are implemented.

[0020] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the audio-visual data processing method described in any one of the above are implemented.

[0021] The beneficial effects of the present invention are specifically as follows: By performing multi-scale wavelet transform on audio data to capture signal features at different time scales, through non-linear mapping reconstruction to correct the distortion in audio data, improve the quality and accuracy of audio data, perform multi-time window segmentation processing on the distortion-corrected audio data to capture audio features within different time periods, through temporal feature evolution fitting to construct an audio temporal semantic feature map, for a deeper understanding and analysis of audio data, enhance the local contrast of the video to improve the clarity and detail sense of the video image, through super-resolution fusion reconstruction to improve the clarity and detail restoration ability of the video, generate a high-quality ultra-clear reconstructed video, perform frame-by-frame temporal decomposition on the ultra-clear reconstructed video to capture the dynamic changes between video frames, through adjacent frame optical flow displacement tracking to reveal the movement trajectories and dynamic change information of objects in the video, perform frame-by-frame positioning and matching on the dynamic video frame sequence based on the audio temporal semantic feature map to achieve accurate matching and alignment of audio-visual data, through fine positioning embedding to embed audio information into video data to achieve a closer association of audio-visual data, calculate the audio time offset value to solve the problem of temporal asynchrony of audio-visual data, and through global delay fine-tuning to optimize the temporal relationship of audio-visual data to achieve better matching and coordinated operation of the overall audio-visual data. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 is a schematic flow chart of the steps of a method for processing audio-visual data according to the present invention;

[0023] Figure 2 is a schematic flow chart of the detailed implementation steps of step S1;

[0024] Figure 3 is a schematic flow chart of the detailed implementation steps of step S2;

[0025] Figure 4 is a schematic flow chart of the detailed implementation steps of step S3. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0027] The embodiments of the present application provide a method, device, equipment and storage medium for processing audio-visual data. The execution subjects of the method, device, equipment and storage medium for processing audio-visual data include but are not limited to: mechanical equipment, data processing platforms, cloud server nodes, network upload devices, etc. that carry this system, which can be regarded as general computing nodes of the present application. The data processing platform includes but is not limited to: at least one of an audio image management system, an information management system, and a cloud data management system.

[0028] Please refer to Figures 1 to 4, the present invention provides a method for processing audio - video data, and the method for processing audio - video data includes the following steps:

[0029] Step S1: Obtain the audio data to be processed and the video data to be processed; perform multi - scale wavelet transform decomposition on the audio data to be processed, and perform non - linear mapping reconstruction to construct distortion - corrected audio data;

[0030] Step S2: Perform multi - time - window segmentation processing on the distortion - corrected audio data, and perform temporal feature evolution fitting to construct an audio temporal semantic feature map;

[0031] Step S3: Perform local contrast enhancement processing on the video data to be processed to generate a locally enhanced video; perform super - resolution fusion reconstruction on the locally enhanced video to construct an ultra - high - definition reconstructed video;

[0032] Step S4: Perform frame - by - frame temporal decomposition on the ultra - high - definition reconstructed video, and perform optical flow displacement tracking between adjacent frames to construct a dynamic video frame sequence;

[0033] Step S5: Perform frame - by - frame localization matching on the dynamic video frame sequence based on the audio temporal semantic feature map, and perform fine - localization embedding to construct an initial audio - embedded video;

[0034] Step S6: Calculate the audio time offset of the initial audio - embedded video to obtain an audio delay value; perform global delay fine - tuning based on the audio delay value to construct a globally delay - optimized video to complete the processing operation of the audio - video data.

[0035] Decompose the audio data through multi - scale wavelet transform to capture signal features at different time scales. Through non - linear mapping reconstruction, correct the distortion in the audio data, improve the quality and accuracy of the audio data. Perform multi - time - window segmentation processing on the distortion - corrected audio data to capture audio features in different time periods. Through temporal feature evolution fitting, construct an audio temporal semantic feature map to gain a deeper understanding and analysis of the audio data. Enhance the local contrast of the video to improve the clarity and detail sense of the video image. Through super - resolution fusion reconstruction, improve the clarity and detail restoration ability of the video to generate a high - quality ultra - high - definition reconstructed video. Perform frame - by - frame temporal decomposition on the ultra - high - definition reconstructed video to capture the dynamic changes between video frames. Through optical flow displacement tracking between adjacent frames, reveal the motion trajectories and dynamic change information of objects in the video. Perform frame - by - frame localization matching on the dynamic video frame sequence based on the audio temporal semantic feature map to achieve accurate matching and alignment of audio - video data. Through fine - localization embedding, embed audio information into the video data to achieve a closer association of audio - video data. Calculate the audio time offset value to solve the problem of temporal asynchrony of audio - video data. Through global delay fine - tuning, optimize the temporal relationship of audio - video data to achieve better matching and coordinated operation of the overall audio - video data.

[0036] In an embodiment of the present invention, refer to Figure 1 , which is a schematic flowchart of the steps of a method for processing audio - video data according to the present invention. In this example, the steps of the method for processing audio - video data include:

[0037] Step S1: Obtain the audio data to be processed and the video data to be processed; perform multi - scale wavelet transform decomposition on the audio data to be processed, and perform non - linear mapping reconstruction to construct distortion - corrected audio data;

[0038] In this embodiment, the original audio data and video data to be processed are obtained. Using multi - scale wavelet transform, such as discrete wavelet transform (DWT), the audio signal is decomposed into sub - bands of different frequency scales to obtain wavelet coefficients at each scale. A non - linear mapping function, such as a power function, hyperbolic tangent function, etc., is designed to perform non - linear transformation and reconstruction on the wavelet coefficients, thereby suppressing the distortion and noise components in the original audio.

[0039] Step S2: Perform multi - time - window segmentation processing on the distortion - corrected audio data, and perform temporal feature evolution fitting to construct an audio temporal semantic feature map;

[0040] In this embodiment, according to the characteristics and analysis requirements of the audio signal, the length of the segmentation window (such as 0.5 seconds, 1 second) and the overlap rate (such as 50%) are determined. An audio processing library (such as librosa or scipy) is used to segment the distortion - corrected audio data. The audio signal is segmented according to the set window length and overlap rate to obtain multiple time - window segments. The segmented time - window segments are organized into a list or array for subsequent processing and analysis. A suitable feature extraction method is selected, such as short - time Fourier transform (STFT), Mel - frequency cepstral coefficients (MFCC), or time - domain features (such as zero - crossing rate, energy, etc.). The selected feature extraction method is applied to each time - window segment to calculate the feature vector of each segment. If MFCC is used, the librosa.feature.mfcc function is used to extract features. The feature vectors of all time - window segments are organized into a matrix, where the rows represent time windows and the columns represent feature dimensions. A suitable temporal fitting model is selected, such as linear regression, support vector machine (SVM), neural network, etc. A suitable model is selected according to the complexity of the features and the amount of data. The selected model is trained using the extracted temporal feature data to fit the evolution curve of the audio features. An audio temporal semantic feature map is generated according to the fitting result, which shows the relationship between time and features in the figure. A visualization tool (such as Matplotlib) is used to draw the feature map for easy understanding and analysis. The generated audio temporal semantic feature map is saved as an image file (such as PNG or SVG) or stored in a data structure for subsequent use.

[0041] Step S3: Perform local contrast enhancement processing on the video data to be processed to generate a locally enhanced video; perform super-resolution fusion reconstruction on the locally enhanced video to construct an ultra-high-definition reconstructed video;

[0042] In this embodiment, local contrast enhancement techniques such as adaptive histogram equalization (CLAHE) are used to perform local contrast enhancement on each frame of the video to improve the clarity and saliency of image details. Deep learning super-resolution fusion techniques such as the residual convolutional neural network (SRCNN), ESRGAN, etc. are used to upscale the low-resolution video frames to high resolution, and higher clarity reconstruction is achieved through feature fusion. Finally, ultra-high-clarity reconstructed video data is constructed.

[0043] Step S4: Perform frame-by-frame temporal decomposition on the ultra-high-definition reconstructed video, and perform optical flow displacement tracking between adjacent frames to construct a dynamic video frame sequence;

[0044] In this embodiment, video frame segmentation technology is used to decompose the video data frame by frame along the time dimension to extract the independent video frame image data of each frame. Optical flow estimation algorithms such as the Lucas-Kanade algorithm and the Farneback algorithm are used to perform optical flow displacement tracking between adjacent video frames, calculate the optical flow displacement vector of each pixel point, generate displacement vector data describing the dynamic optical flow change between adjacent frames, and serially connect each frame of video frame according to the optical flow motion relationship to construct a dynamic video frame sequence reflecting continuous motion.

[0045] Step S5: Perform frame-by-frame localization matching on the dynamic video frame sequence based on the audio temporal semantic feature map, and perform fine localization embedding to construct an initial audio-embedded video;

[0046] In this embodiment, the constructed audio temporal semantic feature map is aligned with the dynamic video frame sequence to establish a time synchronization relationship between the two. After alignment, the similarity between each frame of the video frame and the audio temporal semantic feature is calculated to obtain the semantic feature similarity data of each frame of the video frame. Combining optimization algorithms such as dynamic programming and hidden Markov models, fine localization matching is performed on the audio temporal semantic feature and the video frame sequence to obtain the specific localization data of the audio semantic feature in the video frame sequence. According to the localization data of the semantic feature, it is accurately embedded into the corresponding video frame to construct a preliminary audio semantic-embedded video.

[0047] Step S6: Calculate the audio time offset of the initial audio-embedded video to obtain the audio delay value; perform global delay fine-tuning based on the audio delay value to construct a globally delay-optimized video to complete the processing task of the audio-visual data.

[0048] In this embodiment, an audio synchronization detection algorithm, such as the correlation peak method, the phase difference method, etc., is adopted to analyze the time difference between the audio and the video within the video frame, calculate the specific audio time offset of each video frame, obtain the numerical data describing the audio delay state, and perform global time calibration on the initial audio embedded in the video. Technologies such as dynamic time warping (DTW) are used to dynamically insert or delete the audio signal, so that the time between the audio and the video is completely coordinated, and the final video output with optimized global delay is constructed.

[0049] In this embodiment, refer to Figure 2 , which is a schematic diagram of the detailed implementation steps of step S1. In this embodiment, the detailed implementation steps of the said step S1 include:

[0050] Step S11: Obtain the audio-visual data to be processed;

[0051] Step S12: Perform deep semantic segmentation processing on the audio-visual data to be processed, so as to extract the audio data to be processed and the video data to be processed;

[0052] Step S13: Perform multi-scale wavelet transform decomposition on the audio data to be processed to obtain audio information in different frequency bands;

[0053] Step S14: Perform signal-to-noise ratio prediction calculation on the audio information in different frequency bands to generate the signal-to-noise ratio of each frequency band;

[0054] Step S15: Perform dynamic filtering optimization processing on the audio information in different frequency bands based on the signal-to-noise ratio of each frequency band, so as to obtain multiple segments of noise-reduced and optimized audio signals;

[0055] Step S16: Perform non-linear mapping reconstruction on the multiple segments of noise-reduced and optimized audio signals to construct the audio data with distortion correction.

[0056] In this embodiment, the audio-visual data to be analyzed is obtained through a collection device or a database. Using a pre-trained deep learning semantic segmentation model, such as YOLO, Mask R-CNN, etc., the input audio-visual data is input into the segmentation model for forward inference. The model will perform semantic-level segmentation and extraction on the input audio-visual data, thereby successfully separating and extracting the audio data and video data. The segmentation result will output two independent datasets: the audio data to be processed and the video data to be processed. Technologies such as discrete wavelet transform (DWT) or wavelet packet analysis are used to perform multi-scale decomposition on the audio signal along the time dimension to obtain audio sub-signals in different frequency bands. These sub-signals contain frequency feature information at different granularities. Using statistical analysis or machine learning methods, such as Gaussian process regression, the signal-to-noise ratio of the audio signal in each frequency band is predicted and calculated. The signal-to-noise ratio reflects the severity of the noise in the audio signal in that frequency band. These signal-to-noise ratio data will provide a basis for subsequent dynamic filtering optimization. An adaptive filtering algorithm, such as Wiener filtering, Kalman filtering, etc., is used to perform dynamic filtering on the audio sub-signals in each frequency band. According to the signal-to-noise ratio in different frequency bands, the noise interference is reduced targeted. Finally, multiple segments of noise-reduced and optimized audio signals are obtained. Nonlinear mapping technologies such as neural networks and Fourier transform are used to perform distortion correction and signal reconstruction on the optimized audio signals to generate the final high-quality noise-reduced and optimized audio data.

[0057] In this embodiment, refer to Figure 3 , which is a schematic diagram of the detailed implementation steps of step S2. In this embodiment, the detailed implementation steps of the said step S2 include:

[0058] Step S21: Perform multi-time window segmentation on the distortion-corrected audio data to obtain audio data in multiple time windows;

[0059] Step S22: Perform audio frequency analysis on the audio data in multiple time windows one by one to extract the frequency features of each time window;

[0060] Step S23: Perform key semantic feature recognition on the audio data in multiple time windows to obtain key semantic features in multiple time periods;

[0061] Step S24: Perform time-series change trend analysis on the frequency features of each time window to generate audio time-series change trend data;

[0062] Step S25: Perform time-series feature evolution fitting on the audio time-series change trend data to construct an audio time-series trend curve;

[0063] Step S26: Perform dynamic semantic mapping on the audio time-series trend curve based on the key semantic features in multiple time periods to construct an audio time-series semantic feature map.

[0064] In this embodiment, the sliding window technique is used to divide the audio signal into multiple time windows, and each time window contains a segment of continuous audio data. In this way, the audio data is segmented in the time dimension. According to the audio data of each time window, the frequency characteristics of the audio signal in each time window are extracted by using Fourier transform or other spectrum analysis methods, such as indicators like frequency distribution and band energy. By using speech recognition and natural language processing technologies, the key semantic features in each time window are identified, such as keywords and sentiment tendencies, which provide input for the construction of the temporal semantic feature map. Based on the frequency characteristics of each extracted time window and combined with the time series information of the window, the changing trends of various frequency characteristics over time are analyzed to generate trend data reflecting the temporal variation law of the audio. By using technologies such as time series analysis and curve fitting, a trend curve model that can describe the evolution law of the audio temporal characteristics is constructed for the obtained temporal variation trend data. Using the key semantic features, the constructed temporal trend curves are dynamically associated to construct a temporal semantic feature map of the audio signal in the time dimension, and this feature map can intuitively reflect the temporal variation of the semantic content in the audio.

[0065] In this embodiment, refer to Figure 4 , which is a schematic diagram of the detailed implementation steps of step S3. In this embodiment, the detailed implementation steps of the said step S3 include:

[0066] Step S31: Perform edge contour visual recognition on the video data to be processed and extract the video edge contour features;

[0067] Step S32: Perform local contrast enhancement processing based on the video edge contour features to generate a locally enhanced video;

[0068] Step S33: Perform multi-sampling convolution operation on the locally enhanced video to generate a multi-scale feature map;

[0069] Step S34: Perform texture detail generalization learning on the multi-scale feature map to obtain the texture details of each scale;

[0070] Step S35: Perform fully connected detail mapping on the multi-scale feature map according to the texture details of each scale to generate multiple texture detail feature maps;

[0071] Step S36: Perform super-resolution fusion reconstruction on the multiple texture detail feature maps to construct a super-clear reconstructed video.

[0072] In this embodiment, a suitable edge detection algorithm is selected, such as Canny edge detection, Sobel operator or Laplacian operator. The selected edge detection algorithm is applied to each frame of the video to extract the edge contour features of the video. Combining with a contrast enhancement algorithm, such as histogram equalization, local contrast enhancement processing is performed on the original video data to highlight the local details and texture information in the video, generating video data with enhanced local contrast. Convolution kernels of various sizes are designed to capture features at different scales. Multiple convolution kernels are applied to each frame of the locally enhanced video for multi-sampling convolution operations to generate multi-scale feature maps. A suitable texture detail learning method is selected, such as a convolutional neural network (CNN) or GAN. The multi-scale feature maps are input into the selected learning model for training. The multi-scale feature maps are trained to extract the texture details at each scale. A fully connected layer is designed to connect the texture details at each scale with the corresponding multi-scale feature maps, and full connection detail mapping is performed on the multi-scale feature maps to generate multiple texture detail feature maps. A suitable super-resolution reconstruction algorithm, such as SRCNN, ESPCN or GAN, is selected. The multiple texture detail feature maps are input into the super-resolution model for fusion reconstruction to generate a super-clear reconstructed video. Visualization processing is performed on the reconstructed video to ensure that its quality and details meet the requirements.

[0073] In this embodiment, step S4 includes the following steps:

[0074] Step S41: Perform frame-by-frame temporal decomposition on the super-clear reconstructed video to extract each frame of video frame image;

[0075] Step S42: Perform inter-frame optical flow displacement tracking on each frame of video frame image to generate an inter-frame dynamic optical flow displacement vector;

[0076] Step S43: Perform continuous frame correlation mining on the inter-frame dynamic optical flow displacement vector to generate a continuous inter-frame optical flow motion relationship;

[0077] Step S44: Based on the continuous inter-frame optical flow motion relationship, perform dynamic temporal concatenation processing on each frame of video frame image to construct a dynamic video frame sequence.

[0078] In this embodiment, the video frame segmentation technology is adopted to decompose the video data frame by frame along the time dimension, and the independent video frame image data of each frame is extracted. The optical flow estimation algorithms, such as Lucas-Kanade algorithm, Farneback algorithm, etc., are used to track the optical flow displacement between adjacent video frames, calculate the optical flow displacement vector of each pixel point, and generate the displacement vector data describing the dynamic optical flow change between adjacent frames. Combining with the time series analysis technology, the optical flow motion correlation between consecutive video frames is mined, and a model describing the optical flow change relationship between consecutive frames is established. Using the obtained optical flow motion relationship between consecutive frames, each extracted video frame image is processed by dynamic time series concatenation according to the optical flow change relationship, and a dynamic video frame sequence reflecting the continuous motion change is generated, which can better describe the dynamic detail information in the video.

[0079] In this embodiment, the specific steps of step S5 are as follows:

[0080] Step S51: Align the dynamic video frame sequence based on the audio temporal semantic feature map, and calculate the audio semantic feature similarity to obtain the semantic feature similarity of each video frame.

[0081] Step S52: Perform frame-by-frame localization matching based on the semantic feature similarity of each video frame to obtain the localization data of the audio semantic features.

[0082] Step S53: Extract multiple segments of audio semantic features based on the audio temporal semantic feature map.

[0083] Step S54: Fine-locate and embed the multiple segments of audio semantic features into the dynamic video frame sequence according to the localization data of the audio semantic features to construct the initial audio-embedded video.

[0084] In this embodiment, temporal semantic feature maps are extracted from the audio signal. Methods such as MFCC and audio spectrograms are used to synchronize and align the dynamic video frame sequence with the audio temporal features, ensuring that each video frame corresponds to the correct audio segment. Similarity calculation methods (such as cosine similarity and Euclidean distance) are used to calculate the similarity between each video frame and the corresponding audio features. A suitable positioning and matching algorithm is selected, such as dynamic time warping (DTW) or simple threshold matching. According to the similarity calculation results, each video frame is frame-by-frame positioned and matched to determine the audio feature segment that is most similar to it. According to the temporal diagram of the audio features, the audio signal is segmented, and multiple segments of audio semantic features corresponding to the video frames are extracted. The extracted multiple segments of audio semantic features are sorted out to ensure that each segment of features is consistent with the positioning data of the video frame. A strategy for embedding audio into video is designed to determine how to embed the audio features into the video frames (such as audio overlay and audio clip). According to the positioning data, the multiple segments of audio semantic features are finely embedded into the corresponding dynamic video frame sequence to generate the initial audio-embedded video. The generated initial audio-embedded video is played and quality-checked to ensure the synchronization and coordination of the audio and video.

[0085] In this embodiment, the specific steps of step S6 are as follows:

[0086] Step S61: Identify the audio time deviation frame by frame for the initial audio-embedded video and extract the audio time deviation video frames;

[0087] Step S62: Calculate the audio time offset for the audio time deviation video frames to obtain the audio time delay value;

[0088] Step S63: Analyze the motion feature distribution of the audio time deviation video frames to obtain the video frame motion distribution features;

[0089] Step S64: Identify the optimal time optimization point based on the video frame motion distribution features to obtain the optimal time delay insertion optimization point in the frame;

[0090] Step S65: Dynamically correct the audio offset for the optimal time delay insertion optimization point in the frame according to the audio time delay value to obtain the audio offset corrected frame;

[0091] Step S66: Perform global time delay fine-tuning on the initial audio-embedded video based on the audio offset corrected frame to construct the global time delay optimized video, so as to complete the processing operation of the audio and video data.

[0092] In this embodiment, a suitable time deviation detection method is selected, such as time delay estimation based on correlation analysis or signal processing, to identify the audio time deviation for each frame, mark the video frames with audio time deviation, analyze the time difference between audio and video within the video frames, calculate the specific audio time offset for each target frame, obtain the numerical data describing the audio time delay state, use methods such as optical flow analysis to extract the motion feature distribution information within each target frame, such as indicators like motion direction and speed distribution, use optimization algorithms such as genetic algorithms and particle swarm optimization to identify the optimal position in each target video frame that is most suitable for inserting the audio time delay, obtain the optimal time delay insertion point that minimizes video distortion, select a suitable audio offset correction method, such as linear interpolation or spline interpolation, perform dynamic audio offset correction on the optimal time delay insertion optimization point according to the calculated audio time delay value, generate audio offset correction frames, determine the strategy for global time delay fine-tuning, ensure the synchronization of audio and video, selectively replace the corresponding frames in the initial audio embedded in the video, achieve the optimal fine-tuning of the global time delay, and finally construct a globally time delay optimized video with completely coordinated audio and video time.

[0093] In this embodiment, the present invention also provides an audio-visual data processing device, including:

[0094] A non-linear reconstruction module, configured to obtain the audio data to be processed and the video data to be processed; perform multi-scale wavelet transform decomposition on the audio data to be processed, and perform non-linear mapping reconstruction to construct distortion-corrected audio data;

[0095] A semantic feature module, configured to perform multi-time window segmentation processing on the distortion-corrected audio data, and perform time series feature evolution fitting to construct an audio time series semantic feature map;

[0096] An image enhancement module, configured to perform local contrast enhancement processing on the video data to be processed to generate a locally enhanced video; perform super-resolution fusion reconstruction on the locally enhanced video to construct an ultra-high definition reconstructed video;

[0097] An optical flow displacement tracking module, configured to perform frame-by-frame time series decomposition on the ultra-high definition reconstructed video, and perform optical flow displacement tracking between adjacent frames to construct a dynamic video frame sequence;

[0098] A fine positioning and embedding module, configured to perform frame-by-frame positioning and matching on the dynamic video frame sequence based on the audio time series semantic feature map, and perform fine positioning and embedding to construct an initial audio embedded video;

[0099] A time delay optimization module, configured to calculate the audio time offset for the initial audio embedded video to obtain the audio time delay value; perform global time delay fine-tuning based on the audio time delay value to construct a globally time delay optimized video, so as to complete the processing operation of the audio-visual data.

[0100] The present invention also provides a computer device, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the audio and video data processing method described in any one of the above is implemented.

[0101] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the audio and video data processing method described in any one of the above are implemented.

[0102] Those skilled in the art clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, systems and units are referred to the corresponding processes in the foregoing method embodiments and will not be described herein again.

[0103] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it is stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application essentially, or the part that contributes to the prior art, or all or part of the technical solution, is embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc for storing program codes.

[0104] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the application documents are intended to be included in the present invention.

[0105] As described above, this is only the specific implementation manner of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for processing audio and video data, characterized in that: The following steps are involved: Step S1: obtaining audio data to be processed and video data to be processed; performing multi-scale wavelet transform decomposition on the audio data to be processed, and performing nonlinear mapping reconstruction to construct distortion-corrected audio data; Step S2: performing multi-time window segmentation processing on the distortion-corrected audio data, and performing time series feature evolution fitting to construct an audio time series semantic feature graph; Step S3: performing local contrast enhancement processing on the video data to be processed to generate a local enhanced video; performing super-resolution fusion reconstruction on the local enhanced video to construct an ultra-clear reconstructed video; Step S4: Decompose the ultra-high-definition reconstructed video frame by frame, and track the optical flow displacement between adjacent frames to construct a dynamic video frame sequence; the specific steps of step S4 are: Step S41: decomposing the ultra-high-definition reconstructed video frame by frame, and extracting each video frame image; Step S42: tracking the optical flow displacement between adjacent frames of each video frame image to generate an inter-frame dynamic optical flow displacement vector; Step S43: performing continuous frame association mining on the inter-frame dynamic optical flow displacement vector to generate the continuous inter-frame optical flow motion relationship; Step S44: performing dynamic time series processing on each video frame image based on the optical flow motion relationship between consecutive frames to construct a dynamic video frame sequence; Step S5: Based on the audio temporal semantic feature graph, the dynamic video frame sequence is positioned and matched frame by frame, and fine positioning and embedding are performed to construct an initial audio embedded video; the specific steps of step S5 are: Step S51: aligning the dynamic video frame sequence based on the audio time-series semantic feature graph, and calculating the audio semantic feature similarity to obtain the semantic feature similarity of each video frame; Step S52: performing frame-by-frame positioning matching based on the semantic feature similarity of each video frame, thereby obtaining positioning data of the audio semantic feature; Step S53: extracting semantic features of multiple audio segments based on the audio time-series semantic feature graph; Step S54: finely locate and embed multiple audio semantic features into the dynamic video frame sequence according to the positioning data of the audio semantic features, and construct an initial audio-embedded video; The fine positioning embedding specifically includes: finely embedding multiple audio semantic features into corresponding dynamic video frame sequences according to the positioning data; Step S6: Calculate the audio time offset of the initial audio embedded in the video to obtain an audio delay value; Fine-tune the global delay based on the audio delay value and build a global delay optimization video to complete the processing of audio and video data.

2. The method for processing audio and video data according to claim 1, characterized in that: The specific steps of step S1 are: Step S11: Obtaining audio and video data to be processed; Step S12: performing deep semantic segmentation processing on the audio and video data to be processed, thereby extracting the audio data to be processed and the video data to be processed; Step S13: performing multi-scale wavelet transform decomposition on the audio data to be processed to obtain audio information of different frequency bands; Step S14: performing signal-to-noise ratio prediction calculation on audio information of different frequency bands to generate a signal-to-noise ratio for each frequency band; Step S15: performing dynamic filtering optimization processing on audio information of different frequency bands based on the signal-to-noise ratio of each frequency band, thereby obtaining multiple noise reduction optimized audio signals; Step S16: Perform nonlinear mapping reconstruction on multiple noise reduction optimized audio signals to construct distortion corrected audio data.

3. The method for processing audio and video data according to claim 1, characterized in that: The specific steps of step S2 are: Step S21: performing multi-time window segmentation processing on the distortion-corrected audio data to obtain audio data of multiple time windows; Step S22: performing audio frequency analysis on the audio data of multiple time windows one by one, and extracting the frequency features of each time window; Step S23: performing key semantic feature recognition on the audio data of multiple time windows to obtain key semantic features of multiple time periods; Step S24: performing time series change trend analysis on the frequency characteristics of each time window to generate audio time series change trend data; Step S25: performing time series feature evolution fitting on the audio time series change trend data to construct an audio time series trend curve; Step S26: Perform dynamic semantic mapping on the audio time series trend curve based on multiple time period key semantic features to construct an audio time series semantic feature graph.

4. The method for processing audio and video data according to claim 1, characterized in that: The specific steps of step S3 are: Step S31: Performing edge contour visual recognition on the video data to be processed to extract video edge contour features; Step S32: performing local contrast enhancement processing based on the edge contour features of the video to generate a local enhanced video; Step S33: performing a multi-sampling convolution operation on the local enhanced video to generate a multi-scale feature map; Step S34: performing texture detail generalization learning on the multi-scale feature map to obtain texture details at each scale; Step S35: performing fully connected detail mapping on the multi-scale feature map according to the texture details of each scale to generate multiple texture detail feature maps; Step S36: Perform super-resolution fusion reconstruction on multiple texture detail feature maps to construct an ultra-clear reconstructed video.

5. The method for processing audio and video data according to claim 1, characterized in that: The specific steps of step S6 are: Step S61: performing frame-by-frame audio time deviation recognition on the initial audio-embedded video, and extracting audio time deviation video frames; Step S62: performing audio time offset calculation on the audio time deviation video frame to obtain an audio delay value; Step S63: performing motion feature distribution analysis on the audio time deviation video frame to obtain the video frame motion distribution feature; Step S64: identifying the optimal time optimization point based on the video frame motion distribution characteristics to obtain the optimal delay insertion optimization point in the frame; Step S65: Perform dynamic audio offset correction on the optimal delay insertion optimization point in the frame according to the audio delay value to obtain an audio offset correction frame; Step S66: fine-tune the global delay of the initial audio-embedded video based on the audio offset correction frame, and construct a global delay optimized video to complete the processing of audio and video data.

6. A device for processing audio and video data, characterized in that: The method for processing audio and video data according to claim 1 comprises: A nonlinear reconstruction module is used to obtain the audio data to be processed and the video data to be processed; multi-scale wavelet transform is performed on the audio data to be processed, and nonlinear mapping reconstruction is performed to construct distortion-corrected audio data; The semantic feature module is used to perform multi-time window segmentation processing on the distortion-corrected audio data, perform time series feature evolution fitting, and construct an audio time series semantic feature graph; The image enhancement module is used to perform local contrast enhancement processing on the video data to be processed to generate a local enhanced video; and to perform super-resolution fusion reconstruction on the local enhanced video to construct an ultra-clear reconstructed video; The optical flow displacement tracking module is used to decompose the ultra-high-definition reconstructed video frame by frame, and to track the optical flow displacement between adjacent frames to construct a dynamic video frame sequence; A fine positioning embedding module is used to perform frame-by-frame positioning matching of dynamic video frame sequences based on audio temporal semantic feature graphs, and to perform fine positioning embedding to construct an initial audio embedded video; The delay optimization module is used to calculate the audio time offset of the initial audio embedded in the video to obtain the audio delay value; based on the audio delay value, the global delay is fine-tuned to construct a global delay optimized video to complete the processing of audio and video data.

7. A computer device comprising a memory and a processor, wherein a computer program is stored in the memory, wherein: When the processor executes the computer program, the steps of the method for processing audio and video data according to any one of claims 1 to 5 are implemented.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for processing audio and video data according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Storage and playing method and storage and playing device for monitoring video

    CN112969068A

  • Audio and video synchronization method and device, medium, equipment and program product

    CN114422825A