An ultra-high-definition video AI repair method based on multi-modal feature fusion

By using multimodal feature fusion and AI image restoration technology, distorted regions in ultra-high-definition videos are identified and reconstructed, solving the structural drift and temporal stability problems of video coding reference relationships in existing technologies, and improving the stability of video temporal structure and visual continuity.

CN122457776APending Publication Date: 2026-07-24SHAANXI GUANGXIN NEW MEDIA CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHAANXI GUANGXIN NEW MEDIA CO LTD
Filing Date
2026-06-26
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing technologies lack the ability to analyze distortion propagation paths, cross-temporal structural correlation features, and multimodal consistency between audio and visual data in ultra-high-definition video coding reference relationships. This leads to structural drift, motion trajectory breakage, and decreased temporal stability after reference frame contamination.

Method used

By identifying structural distortion states, constructing frame-level dependency graphs, evaluating motion vector magnitudes, and fusing multimodal features, candidate distortion regions are identified and differentiated reconstruction and restoration are performed using an AI image restoration model.

Benefits of technology

It effectively avoids the structural misalignment problem caused by traditional local enhancement, and improves the temporal structural stability and visual continuity of ultra-high-definition video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122457776A_ABST
    Figure CN122457776A_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on multimodal feature fusion's ultra-high-definition video AI repair method, it is related to video repair technical field, for solving the structure drift after reference frame pollution cannot be accurately identified, and the structure misplacement of repair area appears in continuous playing process Problem, by executing code rate fluctuation detection and adjacent frame jitter analysis to the video stream to be measured, to identify structure distortion state, video stream is decoded and extracted coding structure information, the frame level dependency graph corresponding to video frame is constructed, motion vector amplitude is collected and distortion intensity is generated in combination with frame distortion analysis result, the initial frame weight is dynamically regulated, to evaluate the propagation trend of distortion in reference frame link, fusion audio feature data and visual feature data execute consistency determination, in combination with distortion propagation trend mark candidate distortion area, and utilize AI image repair model to execute different reconstruction repair to candidate distortion area, avoid the structure misplacement problem caused by traditional local enhancement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video restoration technology, and more specifically, to an AI-based method for restoring ultra-high-definition videos based on multimodal feature fusion. Background Technology

[0002] With the rapid development of ultra-high-definition video services, ultra-high-definition video has been widely used in scenarios such as remote conferencing, film and television production, intelligent security, industrial inspection, and live streaming. Due to the characteristics of ultra-high-definition video, such as high resolution, large bit rate, and strong inter-frame correlation, it is easily affected by factors such as transmission congestion, abnormal reference frames, coding link drift, and complex motion scene switching during video acquisition, encoding compression, network transmission, and terminal decoding. This can lead to problems such as frame structure damage, local texture disorder, abnormal temporal continuity, and cross-frame content drift.

[0003] The existing technology has the following shortcomings: Currently, existing technologies typically use single-frame enhancement, local texture compensation, or fixed-weight repair to process distorted areas in ultra-high-definition videos. However, they lack joint analysis capabilities for distortion propagation paths in video coding reference relationships, cross-temporal structural correlation features, and multimodal consistency between audio and vision. This can easily lead to the inaccurate identification of structural drift after reference frame contamination, and result in structural misalignment, motion trajectory breakage, and decreased temporal stability in the repaired area during continuous playback. Therefore, this paper proposes an AI-based repair method for ultra-high-definition videos based on multimodal feature fusion.

[0004] The information disclosed in the background section is only intended to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0005] To overcome the aforementioned deficiencies of the prior art, embodiments of the present invention provide an AI-based restoration method for ultra-high-definition video based on multimodal feature fusion. This method utilizes a structural distortion state recognition mechanism to perform bitrate fluctuation analysis and neighboring frame jitter detection on the video stream, and constructs a frame-level dependency graph by combining encoding structure information. It dynamically evaluates the distortion propagation trend using motion vector amplitude and distortion intensity, and simultaneously performs consistency judgment by fusing audio feature data and visual feature data. This allows for differentiated intelligent reconstruction and restoration of candidate distortion regions to address the problems mentioned in the background art.

[0006] To achieve the above objectives, the present invention provides the following technical solution: an AI-based method for repairing ultra-high-definition video based on multimodal feature fusion, comprising the following steps: Step S1: Receive the video stream to be tested, obtain the video frame bitrate data of the video stream to be tested and calculate the bitrate fluctuation coefficient, detect the jitter interval between adjacent frames of the video stream to be tested, and analyze the structural distortion state of the video stream to be tested in combination with the bitrate fluctuation coefficient. Step S2: Determine whether to decode the video stream under test based on the structural distortion state, extract the video frames and encoding structure information of the video stream under test based on the decoding results, and construct a frame-level dependency graph for the video frames using the encoding structure information; Step S3: Collect the motion vector amplitude of the video stream under test, calculate the initial frame weight based on the motion vector amplitude, perform frame distortion analysis on the video stream under test and generate distortion intensity, use the distortion intensity to adjust the initial frame weight and evaluate the distortion propagation trend; Step S4: Extract audio and visual feature data of the video stream to be tested and perform consistency judgment. Mark candidate distortion regions by combining the distortion propagation trend and consistency judgment results. Perform reconstruction and repair on candidate distortion regions through AI image restoration model.

[0007] In a preferred embodiment, in step S1, the video stream to be tested is received in real time, and the consecutive video frames are buffered and arranged according to the decoding timestamp order in the video stream to be tested. Video frame bitrate data refers to the amount of data transmitted per unit time for a corresponding video frame. The average bitrate of the video frame is obtained by calculating the average bitrate of all video frame bitrate data. Calculate the dispersion of all video frame bitrate data relative to the average video frame bitrate to generate the standard deviation of video frame bitrate. Normalize the standard deviation of video frame bitrate and the average video frame bitrate to obtain the bitrate fluctuation coefficient. Read the received timestamp information corresponding to consecutive video frames and calculate the time interval between adjacent video frames to obtain the actual received time interval; Based on the theoretical frame interval time preset in the video coding protocol, deviation analysis is performed on the actual received time interval to obtain the adjacent frame jitter interval.

[0008] In a preferred embodiment, in step S1, the jitter interval between adjacent frames of consecutive video frames is statistically analyzed and the average value is calculated to obtain the average jitter interval. The bit rate fluctuation coefficient and the average jitter interval are jointly analyzed to construct a structural distortion evaluation value. When the structural distortion assessment value is greater than or equal to the structural distortion threshold, the structural distortion state of the video stream under test is determined to be an abnormal distortion state. If the structural distortion assessment value is less than the structural distortion threshold, the structural distortion state of the video stream under test is determined to be a non-abnormal distortion state.

[0009] In a preferred embodiment, in step S2, when the structural distortion state of the video stream under test is determined to be an abnormal distortion state, the decoding analysis mechanism is triggered; otherwise, the basic playback processing flow is maintained. After the decoding and analysis mechanism is triggered, frame-by-frame decoding processing is performed on the video stream under test, and the video frames and encoding structure information corresponding to the video stream under test are extracted simultaneously. A video frame is an image data unit obtained after decoding. The encoded structure information includes frame type identifier, reference frame index relationship, and display sequence number. The frame type identifier is used to characterize the encoding prediction method of the current video frame, and the reference frame index relationship refers to the set of reference video frame numbers referenced by the current video frame during the encoding prediction process. A reference frame is a historical or future video frame that is referenced by the current video frame to generate the predicted image content during the video coding prediction process.

[0010] In a preferred embodiment, in step S2, each video frame is treated as a node in a graph structure, and the reference frame index relationship is treated as a directed edge between nodes. Extract the time span and frame type identifier between the current video frame and the reference video frame; The time span is the absolute difference between the display sequence number of the video frame and the display sequence number of the reference video frame. A frame type influence factor is generated based on the frame type identifier. The average value is calculated by summing the time span and the frame type influence factor to generate frame-level dependent weights; Based on all video frame nodes and their corresponding frame-level dependency weights, a frame-level dependency graph corresponding to the video stream under test is constructed.

[0011] In a preferred embodiment, in step S3, the macroblock motion vector corresponding to each video frame is extracted, and the macroblock motion vector includes a horizontal displacement component and a vertical displacement component. Perform amplitude calculations on the macroblock motion vectors in the video frame to generate motion vector amplitudes; A macroblock is a basic processing unit obtained by dividing a video frame during the video encoding process; Statistical processing is performed on the motion vector magnitudes of all macroblocks in the video frame and the average value is calculated to generate the average motion vector magnitude corresponding to the video frame; The initial frame weights corresponding to the current video frame are calculated by summing all frame-level dependency weights and combining them with the average motion vector magnitude.

[0012] In a preferred embodiment, in step S3, the brightness structure difference between the current video frame and the corresponding reference video frame is detected, and the edge gradient change and energy attenuation are extracted. Extract the edge gradient values ​​of the current video frame and the reference video frame respectively, calculate the corresponding difference, and generate the edge gradient change. Energy decay is used to characterize the degree of loss of texture details in the current video frame relative to the reference video frame. The edge gradient change and energy decay are jointly calculated, their cumulative value is calculated and normalized to generate the distortion intensity corresponding to the current video frame. The initial frame weights are adjusted using the distortion intensity to generate adjusted frame weights; Time series statistical analysis was performed on the control frame weights corresponding to consecutive video frames, and the growth rate of the control frame weights of consecutive video frames was calculated as the distortion propagation trend.

[0013] In a preferred embodiment, in step S4, the audio signal in the video stream to be tested is subjected to frame segmentation processing, the audio signal is segmented according to a preset audio time window, and the corresponding audio feature data is extracted. The audio feature data includes short-time energy value and audio rhythm change value. Short-time energy value characterizes the degree of energy change of an audio signal within a unit time window; The changes in short-term energy values ​​between consecutive audio time windows are statistically analyzed to generate audio rhythm variation values; Motion analysis is performed on the target region in the video frame, and visual feature data of the corresponding region is extracted. The visual feature data includes the change in region displacement and the change in region edge. The target region refers to the local image region in a video frame used to perform visual feature analysis and distortion detection.

[0014] In a preferred embodiment, in step S4, the change in regional displacement is obtained by calculating the offset distance of the center coordinates of the corresponding target region between consecutive video frames; Edge detection processing is performed on the target area, and the edge changes between consecutive video frames are counted to generate the amount of edge change in the region; The changes in regional displacement and the changes in regional edge are accumulated to generate visual change values. The deviation between the changes in audio rhythm and the visual change values ​​is then calculated. Multiply the change deviation value by the distortion propagation trend to generate the regional anomaly assessment value; When the anomaly assessment value of a region is greater than or equal to the preset anomaly threshold, the corresponding target region is marked as a candidate distorted region; otherwise, the original image state of the corresponding target region is maintained.

[0015] In a preferred embodiment, in step S4, after the candidate distortion regions are marked, the candidate distortion regions are input into the AI ​​image restoration model to perform reconstruction and restoration processing: The AI ​​image restoration model is used to reconstruct the edge structure, texture details and local geometry in the candidate distorted area, and the restoration result is constrained and corrected by combining the spatial texture continuity of the normal area around the candidate distorted area. AI image restoration models refer to image reconstruction models built based on deep learning networks.

[0016] The technical effects and advantages of this invention are as follows: This invention identifies structural distortion by performing bitrate fluctuation detection and adjacent frame jitter analysis on the video stream under test. Subsequently, the video stream is decoded and encoded structural information is extracted to construct a frame-level dependency graph corresponding to each video frame. Further, motion vector amplitudes are collected and combined with frame distortion analysis results to generate distortion intensity, dynamically adjusting the initial frame weights to assess the propagation trend of distortion in the reference frame link. Finally, consistency judgment is performed by fusing audio and visual feature data, candidate distortion regions are marked based on the distortion propagation trend, and a differentiated reconstruction and repair of these candidate distortion regions is performed using an AI image restoration model. By identifying cross-temporal structural drift caused by reference frame contamination, this invention avoids structural misalignment problems caused by traditional local enhancement, thereby improving the temporal structural stability of ultra-high-definition video. Attached Figure Description

[0017] Figure 1 This is a flowchart illustrating the implementation of an AI-based method for repairing ultra-high-definition video based on multimodal feature fusion, as described in this invention.

[0018] Figure 2 This is a schematic diagram illustrating the steps of an ultra-high-definition video AI restoration method based on multimodal feature fusion according to the present invention. Detailed Implementation

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] This invention identifies structural distortion by performing bitrate fluctuation detection and adjacent frame jitter analysis on the video stream under test. Subsequently, the video stream is decoded and encoded structural information is extracted to construct a frame-level dependency graph corresponding to each video frame. Further, motion vector amplitudes are collected and combined with frame distortion analysis results to generate distortion intensity, dynamically adjusting the initial frame weights to assess the propagation trend of distortion in the reference frame link. Finally, consistency judgment is performed by fusing audio and visual feature data, candidate distortion regions are marked based on the distortion propagation trend, and a differentiated reconstruction and repair of the candidate distortion regions is performed using an AI image restoration model. By identifying cross-temporal structural drift caused by reference frame contamination, this invention avoids the structural misalignment problem caused by traditional local enhancement.

[0021] Example 1, such as Figures 1 to 2 As shown, an AI-based method for restoring ultra-high-definition videos based on multimodal feature fusion includes the following steps: Step S1: Receive the video stream to be tested, obtain the video frame bitrate data of the video stream to be tested and calculate the bitrate fluctuation coefficient, detect the jitter interval between adjacent frames of the video stream to be tested, and analyze the structural distortion state of the video stream to be tested in combination with the bitrate fluctuation coefficient. Step S2: Determine whether to decode the video stream under test based on the structural distortion state, extract the video frames and encoding structure information of the video stream under test based on the decoding results, and construct a frame-level dependency graph for the video frames using the encoding structure information; Step S3: Collect the motion vector amplitude of the video stream under test, calculate the initial frame weight based on the motion vector amplitude, perform frame distortion analysis on the video stream under test and generate distortion intensity, use the distortion intensity to adjust the initial frame weight and evaluate the distortion propagation trend; Step S4: Extract audio and visual feature data of the video stream to be tested and perform consistency judgment. Mark candidate distortion regions by combining the distortion propagation trend and consistency judgment results. Perform reconstruction and repair on candidate distortion regions through AI image restoration model.

[0022] The specific implementation is as follows: In step S1, the video stream to be tested is received, the video frame bitrate data of the video stream to be tested is obtained and the bitrate fluctuation coefficient is calculated, the adjacent frame jitter interval of the video stream to be tested is detected, and the structural distortion state of the video stream to be tested is analyzed in combination with the bitrate fluctuation coefficient.

[0023] Specifically, the video stream to be tested is received in real time, and consecutive video frames are buffered and arranged according to the decoding timestamp order in the video stream to be tested.

[0024] It should be noted that the video stream under test is an ultra-high-definition video stream based on the H.264 or H.265 encoding protocol. After receiving the video stream under test, the video stream access module parses the video frame header information in the video stream and extracts the data packet size and corresponding timestamp information for each video frame.

[0025] It should be noted that video frame header information refers to the data control field that exists in the video bitstream and is used to describe the basic attributes of the current video frame.

[0026] Video frame rate refers to the amount of data transmitted per unit of time for a given video frame. It reflects the encoding complexity of the current video frame and the stability of network transmission. The video frame rate is calculated by counting the total number of bytes in the data packets corresponding to each video frame and combining this with the frame duration. The specific calculation method is as follows: ; in, For video frame rate data, Total number of bytes in the data packet. The duration of the frame.

[0027] Video frame bitrate data represents the density of encoded data in the current video frame. The larger the value, the more complex the texture details contained in the current video frame, or the higher the amount of data allocated to the current frame by the encoder; the smaller the value, the higher the compression degree of the current video frame or the data compression limitation during transmission.

[0028] After acquiring the video frame bitrate data corresponding to consecutive video frames, statistical analysis is performed on all video frame bitrate data to calculate the bitrate fluctuation coefficient corresponding to the video stream under test. First, the average bitrate data of all video frames is calculated to obtain the average bitrate of the video frames; Subsequently, the dispersion of all video frame bitrate data relative to the average video frame bitrate is calculated to generate the standard deviation of the video frame bitrate. The standard deviation of the video frame bitrate and the average video frame bitrate are then normalized to obtain the bitrate fluctuation coefficient, which is calculated as follows: ; in, This is the bitrate fluctuation coefficient. The standard deviation of the video frame rate. This represents the average bitrate of the video frames.

[0029] The bitrate fluctuation coefficient characterizes the stability of the bitrate between consecutive video frames. The larger the value, the more obvious the bitrate abrupt change phenomenon in the video stream, that is, there is a large fluctuation in the encoding complexity or network transmission status; the smaller the value, the more stable the overall transmission status of the video stream.

[0030] After calculating the bitrate fluctuation coefficient, the jitter interval between adjacent frames in the video stream under test is further detected. Specifically, the received timestamp information of consecutive video frames is read, and the time interval between adjacent video frames is calculated to obtain the actual received time interval. Subsequently, based on the theoretical frame interval time preset in the video coding protocol, a deviation analysis is performed on the actual received time interval to obtain the adjacent frame jitter interval, which is calculated as follows: ; in, The jitter interval between adjacent frames. This is the actual reception time interval. This is the theoretical frame interval time.

[0031] The adjacent frame jitter interval is used to characterize the degree of deviation between the arrival rhythm of video frames and the theoretical playback rhythm. The larger the value, the more obvious the timing jitter phenomenon of the video frame during transmission; the smaller the value, the more stable the timing of the video frame transmission.

[0032] Furthermore, the jitter intervals of adjacent frames in consecutive video frames are statistically analyzed and averaged to obtain the average jitter interval. The bitrate fluctuation coefficient and the average jitter interval are then jointly analyzed to construct a structural distortion evaluation value to assess the structural distortion state of the video stream under test. The calculation method is as follows: ; in, This is the structural distortion assessment value. This is the bitrate fluctuation coefficient. The jitter interval between adjacent frames. and These are the preset weighting coefficients.

[0033] The structural distortion assessment value characterizes the risk of implicit structural distortion in the current video stream under test. The larger the value, the stronger the abnormal bitrate fluctuation and unstable timing transmission phenomenon in the video stream, which makes it easier to cause reference frame contamination and cross-frame structural drift. The smaller the value, the more stable the overall encoding and transmission status of the current video stream.

[0034] The structural distortion assessment value is compared with the preset structural distortion threshold: When the structural distortion assessment value is greater than or equal to the structural distortion threshold, the structural distortion state of the video stream under test is determined to be an abnormal distortion state. If the structural distortion assessment value is less than the structural distortion threshold, the structural distortion state of the video stream under test is determined to be a non-abnormal distortion state.

[0035] It should be noted that when setting the preset structural distortion threshold, the structural distortion evaluation values ​​corresponding to historical normal video streams are collected, and the average and standard deviation of all structural distortion evaluation values ​​are calculated. The sum of the mean and standard deviation is used as the preset structural distortion threshold.

[0036] In step S2, it is determined whether to decode the video stream under test based on the structural distortion state. Based on the decoding result, the video frames and encoding structure information of the video stream under test are extracted, and the encoding structure information is used to construct a frame-level dependency graph for the video frames.

[0037] Specifically, when the structural distortion state of the video stream under test is determined to be an abnormal distortion state, the decoding analysis mechanism is triggered; otherwise, the decoding analysis mechanism is not triggered and the basic playback processing flow is maintained.

[0038] After the decoding analysis mechanism is triggered, frame-by-frame decoding processing is performed on the video stream under test. Specifically, according to the order of the decoding timestamps in the video stream, the compressed bitstream is subjected to inverse quantization, inverse transform, and motion compensation processing to recover the corresponding video frame image data.

[0039] It should be noted that inverse quantization refers to the process of restoring the frequency domain coefficients after quantization and compression during video encoding; inverse transform refers to the process of reconstructing the spatial domain coefficients after inverse quantization; and motion compensation refers to the process of reconstructing the predicted region in the current video frame using historical image block data and motion vectors from the reference frame.

[0040] During the decoding process, the video frames and encoding structure information corresponding to the video stream under test are extracted simultaneously. A video frame refers to the image data unit obtained after decoding, and each video frame corresponds to a unique frame number. The encoding structure information includes frame type identifiers, reference frame index relationships, and display sequence numbers.

[0041] The frame type identifier is used to characterize the coding prediction method of the current video frame. Specifically, it includes I-frames, P-frames, and B-frames; I-frames represent intra-coded frames, whose generation process does not depend on other video frames; P-frames represent forward-predicted frames, whose generation process depends on historical reference frames; and B-frames represent bidirectional-predicted frames, whose generation process depends on both historical and future reference frames.

[0042] It should be noted that a reference frame refers to a historical or future video frame that is referenced by the current video frame to generate the predicted image content during the video coding prediction process. It is used to provide the basis for motion prediction for the current video frame, and its reference relationship can reflect the temporal propagation structure between video frames.

[0043] Subsequently, the reference frame index relationship corresponding to each video frame is extracted. The reference frame index relationship refers to the set of reference video frame numbers referenced by the current video frame in the encoding and prediction process. The reference frame set reflects the prediction dependency relationship between the current video frame and historical video frames. The more reference frames there are, the higher the cross-frame prediction complexity of the current video frame; the larger the span of the reference frames, the stronger the possibility that the current video frame is affected by the propagation of historical structural information.

[0044] Furthermore, a frame-level dependency graph is constructed based on the reference frame index relationships between video frames. Specifically, each video frame is treated as a node in the graph structure, and the reference frame index relationships are represented as directed edges between nodes. For example, if video frames... Reference video frames As a reference frame, a system is established by nodes. Pointing to node The directed edge.

[0045] After constructing directed edges, the dependency strength of the corresponding reference frame index relationship is further calculated to generate frame-level dependency weights. Specifically, the temporal span and frame type identifier between the current video frame and the reference video frame are extracted. The temporal span is the absolute difference between the display sequence number of the current video frame and the display sequence number of the reference video frame, representing the temporal distance between the current video frame and the reference video frame. A larger value indicates a longer propagation path for the reference information and a higher risk of historical structural error accumulation; a smaller value indicates a more localized reference relationship.

[0046] Subsequently, a frame type influence factor is generated based on the frame type identifier. Specifically, different propagation risk weights are set for I-frames, P-frames, and B-frames. The propagation risk weight values ​​range from 0 to 1. When the reference relationship involves I-frames, a lower propagation risk weight is assigned; when the reference relationship involves consecutive P-frame or B-frame prediction chains, a higher propagation risk weight is assigned.

[0047] Furthermore, the average value is calculated by summing the time span and the frame type influence factor to generate a frame-level dependency weight. The frame-level dependency weight characterizes the degree of influence of the reference video frame on the formation of the current video frame structure. The larger the value, the higher the dependence of the current video frame on the reference video frame, and the easier it is for structural errors in the reference link to propagate to the current video frame. The smaller the value, the weaker the influence of the historical structure of the reference frame on the current video frame.

[0048] Finally, based on all video frame nodes and their corresponding frame-level dependency weights, a frame-level dependency graph is constructed for the video stream under test. The frame-level dependency graph is used to characterize the structural propagation path between reference frames within the video stream, thus providing a basic structural basis for subsequent distortion propagation trend assessment.

[0049] In step S3, the motion vector amplitude of the video stream under test is acquired, the initial frame weight is calculated based on the motion vector amplitude, frame distortion analysis is performed on the video stream under test and distortion intensity is generated, and the frame weight is adjusted and the distortion propagation trend is evaluated using the distortion intensity.

[0050] Specifically, after constructing the frame-level dependency graph, the motion vector magnitude corresponding to each video frame is extracted. The motion vector magnitude is the position offset generated during the inter-frame prediction process, used to characterize the motion displacement of the image patch in the current video frame relative to the reference video frame.

[0051] For macroblock motion vectors in video frames, amplitude calculation is performed to generate motion vector amplitudes. Macroblock motion vectors include horizontal and vertical displacement components. The motion vector amplitudes are calculated as follows: ; in, The magnitude of the motion vector. For horizontal displacement components, This represents the vertical displacement component.

[0052] The magnitude of the motion vector represents the spatial motion intensity of the corresponding macroblock. The larger the value, the more obvious the motion changes in the corresponding region; the smaller the value, the more static the corresponding region tends to be.

[0053] It should be noted that a macroblock refers to the basic processing unit obtained by dividing video frames during the video encoding process.

[0054] Subsequently, the motion vector amplitudes of all macroblocks in the video frame are statistically processed and the average value is calculated to generate the average motion vector amplitude corresponding to the video frame. The average motion vector amplitude represents the overall motion activity of the current video frame. The larger the value, the stronger the motion change in the current video frame; the smaller the value, the weaker the overall change in the current video frame.

[0055] Furthermore, the initial frame weights corresponding to the video frames are calculated by combining the frame-level dependency weights. Specifically, all frame-level dependency weights corresponding to the current video frame are accumulated and weighted by combining them with the average motion vector magnitude. The calculation method is as follows: ; in, As the initial frame weights, The magnitude of the motion vector. For frame-level dependent weights, For video frame index values, For reference video frames, This represents the number of reference video frames corresponding to the video frame.

[0056] The initial frame weight characterizes the degree of structural influence of the current video frame in the reference propagation link. The larger the value, the more the current video frame has both strong motion changes and a high degree of reference dependence, making it easier to form structural error propagation; the smaller the value, the weaker the influence of the current video frame on the propagation of the reference link.

[0057] After calculating the initial frame weights, frame distortion analysis is performed on the video stream under test. Specifically, the brightness structure difference between the current video frame and the corresponding reference video frame is detected, and the edge gradient change and energy attenuation are extracted.

[0058] The edge gradient change represents the degree of change in the edge structure of the current video frame relative to the reference video frame. Specifically, the edge gradient values ​​of the current video frame and the reference video frame are extracted respectively, and the corresponding differences are calculated to generate the edge gradient change.

[0059] Energy attenuation is used to characterize the degree of loss of texture details in the current video frame relative to the reference video frame. Specifically, a frequency domain transformation is performed on the current video frame, and the corresponding spectral energy value is calculated; then, the same processing is performed on the reference video frame, and the difference between the two is calculated to generate the energy attenuation. The larger the energy attenuation value, the more significant the loss of texture details in the current video frame.

[0060] Furthermore, the edge gradient change and energy attenuation are jointly calculated, their cumulative value is calculated and normalized to generate the distortion intensity corresponding to the current video frame. The distortion intensity characterizes the degree of structural distortion of the current video frame. The larger the value, the more obvious the edge structure anomaly and texture detail loss of the current video frame. The smaller the value, the more stable the structural state of the current video frame.

[0061] Subsequently, the initial frame weights are adjusted using the distortion intensity to generate adjusted frame weights, which are calculated as follows: ; in, To adjust frame weights, As the initial frame weights, For distortion intensity, This is the video frame index value.

[0062] The frame weight represents the actual influence of the current video frame in the process of structural distortion propagation. The larger the value, the higher the reference dependence of the current video frame and the more obvious the structural distortion has appeared, making it more likely to become the source of distortion propagation in the subsequent reference link.

[0063] Furthermore, time-series statistical analysis is performed on the control frame weights corresponding to consecutive video frames to assess the distortion propagation trend. Specifically, the growth rate of the control frame weights of consecutive video frames is calculated as the distortion propagation trend, and the calculation method is as follows: ; in, As a trend toward distorted dissemination, The total number of consecutive video frames. and The weight of the adjustment frame is determined by the adjacent video frames. This is the video frame index value.

[0064] The distortion propagation trend characterizes the diffusion trend of structural distortion in the reference frame link. The larger the value, the more the distortion effect is continuously increasing along the video frame reference path; the smaller the value, the more stable or gradually weakening the distortion propagation is.

[0065] In step S4, audio and visual feature data of the video stream to be tested are extracted and consistency judgment is performed. Candidate distortion regions are marked by combining the distortion propagation trend and consistency judgment results. The candidate distortion regions are then reconstructed and repaired using an AI image restoration model.

[0066] Specifically, the audio signal in the video stream under test is processed by frame segmentation. The audio signal is segmented according to a preset audio time window, and the corresponding audio feature data is extracted. The audio feature data includes short-time energy value and audio rhythm change value.

[0067] The short-time energy value characterizes the degree of energy change of the audio signal within a unit time window, and its calculation method is as follows: ; in, This is a short-term energy value. For audio sample values, This is the index value of the audio sample value. This represents the number of sampling points within the audio time window.

[0068] The short-term energy value reflects the level of audio activity within the current time window. The larger the value, the more obvious the changes in speech, explosions, or background noise in the current time period; the smaller the value, the weaker the audio changes in the current time period.

[0069] Subsequently, the short-term energy value changes between consecutive audio time windows are statistically analyzed to generate audio rhythm change values, which are calculated as follows: ; in, This represents the change in audio rhythm. and This represents the short-time energy value of adjacent audio time windows. This is the index value of the audio time window.

[0070] The audio rhythm variation value represents the amplitude of audio variation between adjacent time windows. The larger the value, the more obvious the change in audio rhythm; the smaller the value, the more stable the audio rhythm.

[0071] After extracting the audio feature data, visual feature extraction is performed on the video stream under test. Specifically, motion analysis is performed on the target regions in the video frames, and the visual feature data of the corresponding regions is extracted.

[0072] It should be noted that the target region refers to the local image region in a video frame used to perform visual feature analysis and distortion detection, which can be divided into the mouth region of a person, the edge region of a face, the fast motion region, or the high-frequency texture region.

[0073] The visual feature data includes changes in regional displacement and changes in regional edges. The changes in regional displacement are obtained by calculating the offset distance between the center coordinates of the target region and consecutive video frames. These changes characterize the degree of motion of the target region in consecutive video frames; a larger value indicates more drastic motion, while a smaller value indicates more stable motion.

[0074] Subsequently, edge detection processing is performed on the target area, and the edge changes between consecutive video frames are statistically analyzed to generate the area edge change value. The area edge change value reflects the stability of the target area edge structure. The larger the value, the more obvious the edge change of the target area; the smaller the value, the more stable the edge structure of the target area.

[0075] Furthermore, the changes in regional displacement and the changes in regional edges are accumulated to generate a visual change value. The visual change value represents the overall degree of visual change in the current target area. The larger the value, the more obvious the visual motion or structural change in the current area; the smaller the value, the more stable the visual state of the current area.

[0076] After extracting audio and visual feature data, a consistency determination is performed on the audio rhythm variation values ​​and visual variation values. Specifically, the variation deviation value between the two is calculated as follows: ; in, This represents the variation deviation value. This represents the change in audio rhythm. This represents the visual change value.

[0077] The variation deviation value characterizes the degree of synchronization between audio and visual changes. The larger the value, the more obvious the asynchrony between the visual and audio changes in the current target area, indicating that there may be structural distortion in the current area. The smaller the value, the stronger the synchronization between the visual and audio changes, and the more likely the current area belongs to the real scene change.

[0078] Subsequently, the variation deviation value and the distortion propagation trend are multiplied to generate a regional anomaly assessment value. The regional anomaly assessment value represents the possibility of structural distortion propagation in the current target area. The larger the value, the more likely the current target area is to have both obvious audio and video asynchrony and distortion propagation trend, making it easier to form reference link structure contamination. The smaller the value, the more stable the overall structural state of the current target area.

[0079] Furthermore, the regional anomaly assessment values ​​are compared with preset anomaly thresholds: When the regional anomaly assessment value is greater than or equal to the preset anomaly threshold, the corresponding target region is marked as a candidate distorted region. When the anomaly assessment value of a region is less than the preset anomaly threshold, the original image state of the corresponding target region is maintained.

[0080] It should be noted that when setting the preset anomaly threshold, the distribution of regional anomaly assessment values ​​corresponding to historical normal video areas is statistically analyzed, and the maximum stable fluctuation range of regional anomaly assessment values ​​under normal conditions is extracted as the preset anomaly threshold.

[0081] After the candidate distorted regions are labeled, they are input into the AI ​​image inpainting model for reconstruction and restoration. Specifically, the AI ​​image inpainting model is used to reconstruct the edge structure, texture details, and local geometry of the candidate distorted regions. The restoration results are then constrained and corrected by combining the spatial texture continuity of the normal regions surrounding the candidate distorted regions, thereby restoring the structural integrity of the candidate distorted regions.

[0082] It should be noted that AI image restoration models refer to image reconstruction models built on deep learning networks, used to restore texture details, edge structures, and local geometric shapes in candidate distorted regions. AI image restoration models learn the spatial texture distribution patterns in normal video images to perform structural completion and detail restoration in distorted areas.

[0083] Finally, the repaired target video stream is output to reduce the cross-temporal structural drift caused by reference frame contamination and improve the temporal structural stability and visual continuity of ultra-high-definition video.

[0084] Finally, it should be noted that in this paper, relational terms such as first and second are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations.

[0085] Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0086] In this document, the singular forms “a,” “an,” and “the” may also include the plural forms unless the context clearly indicates otherwise. It should also be understood that terms such as “comprising / including” or “having” specify the presence of the stated features, integrals, steps, operations, components, parts, or combinations thereof, but do not preclude the possibility of the presence or addition of one or more other features, integrals, steps, operations, components, parts, or combinations thereof. Meanwhile, the term “and / or” as used in this specification includes any and all combinations of the associated listed items.

[0087] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can be referred to each other.

[0088] The above description of the disclosed embodiments will enable those skilled in the art to make or use various modifications to these embodiments. It will be readily apparent to those skilled in the art that the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for AI-based restoration of ultra-high-definition video based on multimodal feature fusion, characterized in that: Includes the following steps: Step S1: Receive the video stream to be tested, obtain the video frame bitrate data of the video stream to be tested and calculate the bitrate fluctuation coefficient, detect the jitter interval between adjacent frames of the video stream to be tested, and analyze the structural distortion state of the video stream to be tested in combination with the bitrate fluctuation coefficient. Step S2: Determine whether to decode the video stream under test based on the structural distortion state, extract the video frames and encoding structure information of the video stream under test based on the decoding results, and construct a frame-level dependency graph for the video frames using the encoding structure information; Step S3: Collect the motion vector amplitude of the video stream under test, calculate the initial frame weight based on the motion vector amplitude, perform frame distortion analysis on the video stream under test and generate distortion intensity, use the distortion intensity to adjust the initial frame weight and evaluate the distortion propagation trend; Step S4: Extract audio and visual feature data of the video stream to be tested and perform consistency judgment. Mark candidate distortion regions by combining the distortion propagation trend and consistency judgment results. Perform reconstruction and repair on candidate distortion regions through AI image restoration model.

2. The ultra-high-definition video AI restoration method based on multimodal feature fusion according to claim 1, characterized in that: In step S1, the video stream to be tested is received in real time, and the consecutive video frames are buffered and arranged according to the decoding timestamp order in the video stream to be tested. Video frame bitrate data refers to the amount of data transmitted per unit time for a corresponding video frame. The average bitrate of the video frame is obtained by calculating the average bitrate of all video frame bitrate data. Calculate the dispersion of all video frame bitrate data relative to the average video frame bitrate to generate the standard deviation of video frame bitrate. Normalize the standard deviation of video frame bitrate and the average video frame bitrate to obtain the bitrate fluctuation coefficient. Read the received timestamp information corresponding to consecutive video frames and calculate the time interval between adjacent video frames to obtain the actual received time interval; Based on the theoretical frame interval time preset in the video coding protocol, deviation analysis is performed on the actual received time interval to obtain the adjacent frame jitter interval.

3. The ultra-high-definition video AI restoration method based on multimodal feature fusion according to claim 2, characterized in that: In step S1, the jitter interval between adjacent frames of consecutive video frames is statistically analyzed and the average value is calculated to obtain the average jitter interval. The bit rate fluctuation coefficient and the average jitter interval are jointly analyzed to construct the structural distortion evaluation value. When the structural distortion assessment value is greater than or equal to the structural distortion threshold, the structural distortion state of the video stream under test is determined to be an abnormal distortion state. If the structural distortion assessment value is less than the structural distortion threshold, the structural distortion state of the video stream under test is determined to be a non-abnormal distortion state.

4. The ultra-high-definition video AI restoration method based on multimodal feature fusion according to claim 1, characterized in that: In step S2, when the structural distortion state of the video stream under test is determined to be an abnormal distortion state, the decoding analysis mechanism is triggered; otherwise, the basic playback processing flow is maintained. After the decoding and analysis mechanism is triggered, frame-by-frame decoding processing is performed on the video stream under test, and the video frames and encoding structure information corresponding to the video stream under test are extracted simultaneously. A video frame is an image data unit obtained after decoding. The encoded structure information includes frame type identifier, reference frame index relationship, and display sequence number. The frame type identifier is used to characterize the encoding prediction method of the current video frame, and the reference frame index relationship refers to the set of reference video frame numbers referenced by the current video frame during the encoding prediction process. A reference frame is a historical or future video frame that is referenced by the current video frame to generate the predicted image content during the video coding prediction process.

5. The ultra-high-definition video AI restoration method based on multimodal feature fusion according to claim 4, characterized in that: In step S2, each video frame is treated as a node in the graph structure, and the reference frame index relationship is treated as a directed edge between the nodes. Extract the time span and frame type identifier between the current video frame and the reference video frame; The time span is the absolute difference between the display sequence number of the video frame and the display sequence number of the reference video frame. A frame type influence factor is generated based on the frame type identifier. The average value is calculated by summing the time span and the frame type influence factor to generate frame-level dependent weights; Based on all video frame nodes and their corresponding frame-level dependency weights, a frame-level dependency graph corresponding to the video stream under test is constructed.

6. The ultra-high-definition video AI restoration method based on multimodal feature fusion according to claim 5, characterized in that: In step S3, the macroblock motion vector corresponding to each video frame is extracted. The macroblock motion vector includes horizontal displacement components and vertical displacement components. Perform amplitude calculations on the macroblock motion vectors in the video frame to generate motion vector amplitudes; A macroblock is a basic processing unit obtained by dividing a video frame during the video encoding process; Statistical processing is performed on the motion vector magnitudes of all macroblocks in the video frame and the average value is calculated to generate the average motion vector magnitude corresponding to the video frame; The initial frame weights corresponding to the current video frame are calculated by summing all frame-level dependency weights and combining them with the average motion vector magnitude.

7. The ultra-high-definition video AI restoration method based on multimodal feature fusion according to claim 6, characterized in that: In step S3, the brightness structure difference between the current video frame and the corresponding reference video frame is detected, and the edge gradient change and energy attenuation are extracted. Extract the edge gradient values ​​of the current video frame and the reference video frame respectively, calculate the corresponding difference, and generate the edge gradient change. Energy decay is used to characterize the degree of loss of texture details in the current video frame relative to the reference video frame. The edge gradient change and energy decay are jointly calculated, their cumulative value is calculated and normalized to generate the distortion intensity corresponding to the current video frame. The initial frame weights are adjusted using the distortion intensity to generate adjusted frame weights; Time series statistical analysis was performed on the control frame weights corresponding to consecutive video frames, and the growth rate of the control frame weights of consecutive video frames was calculated as the distortion propagation trend.

8. The ultra-high-definition video AI restoration method based on multimodal feature fusion according to claim 1, characterized in that: In step S4, the audio signal in the video stream under test is subjected to frame segmentation processing. The audio signal is segmented according to a preset audio time window, and the corresponding audio feature data is extracted. The audio feature data includes short-time energy value and audio rhythm change value. Short-time energy value characterizes the degree of energy change of an audio signal within a unit time window; The changes in short-term energy values ​​between consecutive audio time windows are statistically analyzed to generate audio rhythm variation values; Motion analysis is performed on the target region in the video frame, and visual feature data of the corresponding region is extracted. The visual feature data includes the change in region displacement and the change in region edge. The target region refers to the local image region in a video frame used to perform visual feature analysis and distortion detection.

9. The ultra-high-definition video AI restoration method based on multimodal feature fusion according to claim 8, characterized in that: In step S4, the change in regional displacement is obtained by calculating the offset distance of the center coordinates of the corresponding target area between consecutive video frames; Edge detection processing is performed on the target area, and the edge changes between consecutive video frames are counted to generate the amount of edge change in the region; The changes in regional displacement and the changes in regional edge are accumulated to generate visual change values. The deviation between the changes in audio rhythm and the visual change values ​​is then calculated. Multiply the change deviation value by the distortion propagation trend to generate the regional anomaly assessment value; When the anomaly assessment value of a region is greater than or equal to the preset anomaly threshold, the corresponding target region is marked as a candidate distorted region; otherwise, the original image state of the corresponding target region is maintained.

10. The ultra-high-definition video AI restoration method based on multimodal feature fusion according to claim 9, characterized in that: In step S4, after the candidate distorted regions are labeled, the candidate distorted regions are input into the AI ​​image restoration model to perform reconstruction and restoration processing: The AI ​​image restoration model is used to reconstruct the edge structure, texture details and local geometry in the candidate distorted area, and the restoration result is constrained and corrected by combining the spatial texture continuity of the normal area around the candidate distorted area. AI image restoration models refer to image reconstruction models built based on deep learning networks.