Video compression and intelligent reconstruction method and system oriented to extremely low bandwidth
By constructing session time windows and semantic recoverability level maps, video content is organized hierarchically and adaptive bitrate allocation and intelligent reconstruction are performed, which solves the problem of insufficient video reconstruction quality under extremely low bandwidth and achieves stable reconstruction of key areas and efficient transmission of video content.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- UNICOM AIRLINE NETWORK CO LTD
- Filing Date
- 2026-01-26
- Publication Date
- 2026-05-05
AI Technical Summary
Under extremely low bandwidth conditions, existing video compression technologies struggle to effectively characterize the semantic importance, structural stability, and differences in reconstructability of video content, leading to the loss of information in key areas or a decline in reconstruction quality, which affects video recognition and analysis performance.
By constructing a session time window for semantic analysis, a semantic recoverability level map is generated, and the video content is organized into different recoverability levels. Adaptive bitrate allocation and compression encoding are performed under extremely low bandwidth. Combined with semantic guidance information, intelligent reconstruction processing is carried out to generate a target quality video sequence.
Under extremely low bandwidth conditions, this method ensures stable reconstruction of key semantic regions, improves the understandability and task availability of video content, reduces the occupation of invalid information, and enhances the continuity and stability of reconstruction results.
Smart Images

Figure CN121985136A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of video processing technology, and in particular to a method and system for video compression and intelligent reconstruction with extremely low bandwidth. Background Technology
[0002] With the development of video perception and intelligent analysis technologies, video data is increasingly used in remote monitoring, intelligent sensing, and low-bandwidth communication scenarios. Existing video compression technologies mainly revolve around mechanisms such as pixel redundancy elimination, motion prediction, and transform coding. By uniformly encoding and decoding video frames, the transmission and reconstruction of video content are completed. Video is usually treated as a continuous image sequence for processing, focusing on the compression efficiency and reconstruction quality at the signal level, which can meet the basic video transmission and display requirements under normal bandwidth conditions.
[0003] Under extremely low bandwidth conditions, different regions in video content exhibit significant differences in semantic importance, structural stability, and reconstruction difficulty. Traditional unified coding methods struggle to effectively characterize and utilize these differences, easily leading to the loss of key region information or a decline in reconstruction quality, which in turn affects subsequent recognition, analysis, and task execution. In scenarios where both video understandability and task usability need to be considered, existing technologies lack a processing mechanism that can comprehensively characterize the video content's generative characteristics, structural constraints, and semantic importance, making it difficult to achieve stable and controllable video reconstruction results under extremely low bandwidth conditions. Summary of the Invention
[0004] In view of the aforementioned existing problems, the present invention is proposed.
[0005] Therefore, this invention provides a video compression and intelligent reconstruction method for extremely low bandwidth to solve the problem of the difficulty in stably preserving and effectively reconstructing key semantic information of videos under extremely low bandwidth conditions.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: In a first aspect, the present invention provides a video compression and intelligent reconstruction method for extremely low bandwidth, comprising: acquiring an original video frame sequence and constructing a session time window; performing semantic analysis processing on the video content within the session time window to generate a semantic recoverability level map; dividing the video content into video regions with different recoverability levels based on the semantic recoverability level map, and organizing the video content hierarchically to generate minimum recoverability content data and enhanced description data; under the constraint of the target extremely low bandwidth, performing content-adaptive bitrate allocation on the minimum recoverability content data and enhanced description data and compressing and encoding them respectively to form a minimum recoverability bitstream and an enhanced bitstream; receiving the minimum recoverability bitstream and the enhanced bitstream at the decoding end and decoding them, while combining the semantic guidance information carried in the enhanced bitstream and performing intelligent reconstruction processing according to the corresponding recoverability level constraint to generate a target quality video sequence; performing task completion evaluation based on the target quality video sequence, and adaptively updating the semantic recoverability level division rules and bitrate allocation strategy in the session time window.
[0007] As a preferred embodiment of the video compression and intelligent reconstruction method for extremely low bandwidth described in this invention, the original video frame sequence includes multiple frames of video image data and time stamp information.
[0008] As a preferred embodiment of the video compression and intelligent reconstruction method for extremely low bandwidth described in this invention, the specific steps for constructing the session time window are as follows: The original video frame sequence is sequentially traversed according to the time stamp information, and the brightness, chroma and spatial structure changes between adjacent video frames are calculated frame by frame to form a content change metric. Based on the statistical distribution characteristics of content change measurement, a change judgment threshold is obtained. The position where the content change measurement exceeds the change judgment threshold is judged as the session boundary and divided into the initial session time window. For the initial session time window, the temporal distribution density and content change continuity of the video frames within the window are statistically analyzed, and the start and end positions of each initial session time window are corrected to obtain the session time window.
[0009] As a preferred embodiment of the video compression and intelligent reconstruction method for extremely low bandwidth described in this invention, the specific steps for generating the semantically recoverable hierarchy map are as follows: Based on the session time window, the video frames within the window are divided into several spatiotemporal regions according to a fixed spatial division and a fixed time span, and a spatiotemporal region identifier is generated for each spatiotemporal region. Based on the spatiotemporal region identifier, the changes in prediction residuals, motion changes, texture randomness, edge structure continuity, contour stability, and brightness consistency of adjacent frames are statistically analyzed for each spatiotemporal region, forming a generative statistical record and a constrainable statistical record. Based on the constrained statistical records, the existence of target categories, the occurrence of text, and the coverage ratio of task-related areas are collected in each spatiotemporal region to form a record of semantic key intensity. Based on the generative statistical records, constrainable statistical records, and semantic key strength records, each spatiotemporal region is jointly determined and mapped to different semantic recoverability levels, and a semantic recoverability level map is generated according to spatial and temporal locations.
[0010] As a preferred embodiment of the video compression and intelligent reconstruction method for extremely low bandwidth described in this invention, the specific steps for generating baseline content data and enhanced description data are as follows: Based on the semantic recoverability level map, the spatiotemporal regions are marked with different levels, namely, spatiotemporal regions with low recoverability level and spatiotemporal regions with high recoverability level. For spatiotemporal regions with low recoverability, basic pixel information, contour information, and motion trend information are extracted from the original video frame sequence and aggregated to form baseline content data. For spatiotemporal regions with high recoverability, structural guidance information, local residual information, and semantic guidance information are extracted from the original video frame sequence and aggregated to form enhanced description data.
[0011] As a preferred embodiment of the video compression and intelligent reconstruction method for extremely low bandwidth described in this invention, the specific process of forming the minimum bitstream and the enhanced bitstream is as follows: The data on guaranteed content and enhanced description are sorted and organized into fields, and the information is arranged in a structured manner according to time order and spatial location to form standardized input data. Based on the normalized input data, under the constraint of extremely low target bandwidth, priority-protected compression encoding is performed on the guaranteed content data to form a guaranteed bitstream; Based on the remaining available bandwidth after deducting the minimum bitstream usage, adaptive compression encoding is performed on the enhanced description data to form the enhanced bitstream.
[0012] As a preferred embodiment of the video compression and intelligent reconstruction method for extremely low bandwidth described in this invention, the semantic guidance information refers to auxiliary description information extracted from the original video frame sequence and transmitted with the enhanced description data.
[0013] As a preferred embodiment of the video compression and intelligent reconstruction method for extremely low bandwidth described in this invention, the specific process for generating the target quality video sequence is as follows: The decoding end receives the minimum bitstream and the enhanced bitstream, and prioritizes the decoding of the minimum bitstream to obtain the basic video sequence; Based on the basic video sequence, and according to the structural guidance information, local residual information and semantic guidance information carried in the enhanced bitstream, constrained intelligent reconstruction processing is performed on the corresponding spatiotemporal regions in the basic video sequence. The spatiotemporal regions after intelligent reconstruction are fused and corrected for consistency according to their spatial location and time sequence to form a target quality video sequence.
[0014] As a preferred embodiment of the video compression and intelligent reconstruction method for extremely low bandwidth described in this invention, the adaptive update of the semantic recoverability level classification rules and bitrate allocation strategy within the session time window is as follows: Within the session time window, the recognition consistency, temporal stability, and reconstruction continuity of the corresponding regions in the target quality video sequence are statistically analyzed to form input information; Based on the input information, perform task completion evaluation processing on the target quality video sequence in the current session time window to obtain task completion evaluation information; Based on task completion assessment information, modify the generability statistics record, constraint statistics record, and semantic criticality strength record to set semantic recoverability level classification rules; Based on the task completion assessment information, the semantic recoverability level classification rules are adaptively adjusted, and the semantic recoverability level is re-determined for the video content to obtain the corresponding semantic recoverability level distribution. The rate allocation strategy for the minimum and enhanced bitstreams is adaptively updated based on the semantically recoverable level distribution.
[0015] Secondly, this invention provides a video compression and intelligent reconstruction system for extremely low bandwidth, including the aforementioned video compression and intelligent reconstruction method for extremely low bandwidth, characterized by comprising: a session construction module, a hierarchical organization module, a bitrate allocation module, an intelligent reconstruction module, and a feedback update module; the session construction module is used to acquire the original video frame sequence and construct a session time window, perform semantic analysis of the video content within the session time window, and generate a semantic recoverability level map; the hierarchical organization module is used to divide the video content into video regions with different recoverability levels according to the semantic recoverability level map, and hierarchically organize the video content to generate baseline content data and The system includes: enhanced description data; a bitrate allocation module, used to perform content-adaptive bitrate allocation and compression encoding on the baseline content data and enhanced description data under extremely low bandwidth constraints, forming a baseline bitstream and an enhanced bitstream respectively; an intelligent reconstruction module, used to receive the baseline bitstream and enhanced bitstream at the decoding end and decode the baseline bitstream, while combining the semantic guidance information carried in the enhanced bitstream to perform intelligent reconstruction processing according to the corresponding recoverable level constraints, generating a target quality video sequence; and a feedback update module, used to perform task completion evaluation based on the target quality video sequence, and adaptively update the semantic recoverable level classification rules and bitrate allocation strategy within the session time window.
[0016] The beneficial effects of this invention are as follows: By constructing a semantically recoverable hierarchy and performing hierarchical processing on video content, limited bandwidth resources can prioritize the stable reconstruction of key semantic regions; by jointly characterizing the ease of generation, structural constraints, and semantic importance of video spatiotemporal regions, the encoding and reconstruction process is endowed with clear semantic directionality, thereby maintaining the understandability and task availability of video content even under extremely low bandwidth conditions; the video reconstruction process can implement higher-precision constraints and recovery of key regions, reduce the occupation of transmission resources by invalid information, and improve the overall reconstruction results in terms of continuity and stability. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a flowchart of a video compression and intelligent reconstruction method for extremely low bandwidth.
[0019] Figure 2 This is a schematic diagram of a video compression and intelligent reconstruction system designed for extremely low bandwidth.
[0020] Figure 3 Flowchart for generating semantic recoverability levels.
[0021] Figure 4 This is a flowchart for encoding and reconstructing extremely low bandwidth video. Detailed Implementation
[0022] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0023] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0024] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.
[0025] Reference Figures 1-4 This is one embodiment of the present invention, which provides a method for video compression and intelligent reconstruction for extremely low bandwidth, including the following steps: S1: Acquire the original video frame sequence and construct a session time window. Perform semantic analysis of the video content within the session time window using video encoding to generate a semantic recoverability level map.
[0026] S1.1: The original video frame sequence includes multiple frames of video image data and time stamp information.
[0027] By acquiring video input stream frame by frame, multiple frames of video image data are obtained. The corresponding time stamp information is recorded synchronously when acquiring each frame of video image data. By associating the time stamp information with the multiple frames of video image data one by one according to the acquisition order, the original video frame sequence is formed.
[0028] The time stamp information assigns a timestamp to each frame of video image data using a unified time standard, accurately matching each frame of video image data with the time stamp information.
[0029] S1.2: The original video frame sequence is sequentially traversed according to the time stamp information, and the brightness change, color change and spatial structure change between adjacent video frames are calculated frame by frame to form a content change metric.
[0030] Specifically, based on time stamp information, the multi-frame video image data is sorted according to time sequence and adjacent video frames are read sequentially. Each video image is converted into a grayscale image. Using a weighted average method, the difference in grayscale value is calculated pixel by pixel for two adjacent video images, that is, the brightness difference of each pixel in the two frames is calculated. The brightness difference values of all pixels are summarized and statistically analyzed to obtain the brightness change.
[0031] Each video frame is converted from the RGB color space to the HSV color space. For two adjacent video frames, the difference in hue components is calculated pixel by pixel to obtain the chromaticity change. Edge information is extracted by applying an edge detection algorithm to adjacent video frames, and the edge difference between adjacent frames is calculated to measure the spatial structure change. The brightness change, chromaticity change and spatial structure change are collected in a unified measurement format and associated with the corresponding time stamp information to form a content change measurement.
[0032] Among them, the RGB color space represents red, green, and blue; the HSV color space represents hue, saturation, and brightness.
[0033] S1.3: Obtain the change judgment threshold based on the statistical distribution characteristics of the content change measurement, determine the position where the content change measurement exceeds the change judgment threshold as the session boundary, and divide it into the initial session time window.
[0034] Specifically, for the content change measurement within the session time window, the content change values of each video frame are statistically summarized to obtain a change measurement set. The distribution range is obtained by calculating the difference between the maximum and minimum values in the content change measurement set, and the degree of dispersion is obtained by calculating the standard deviation of the content change measurement set.
[0035] By combining the distribution range and dispersion, a change judgment threshold is selected from the content change measurement set to distinguish between regular changes and abrupt changes. The content change measurement sets corresponding to adjacent video frames are compared with the change judgment threshold one by one, and the positions of adjacent video frames whose content change measurement sets exceed the change judgment threshold are marked as session boundaries. The original video frame sequence is segmented and organized according to time identification information using the session boundaries as the dividing points, and divided into initial session time windows.
[0036] It should be noted that statistical distribution characteristics refer to the range and dispersion of content change metrics within the session time window.
[0037] The threshold for determining change is calculated using the following formula: ; In the formula, This indicates the threshold for determining change. This represents the mean. This represents the adjustable coefficient, with a value ranging from 0.5 to 2.0. It represents the standard deviation.
[0038] In the formula for determining the threshold of change, the mean and standard deviation have the same dimension, so they are consistent in dimension.
[0039] S1.4: For the initial session time window, calculate the temporal distribution density and content change continuity of the video frames within the window, and correct the start and end positions of each initial session time window to obtain the session time window.
[0040] Specifically, the time stamp information corresponding to the video frames within the initial session time window is read, and the interval distribution between adjacent time stamps is calculated to statistically analyze the time distribution density of the video frames. The content change metric within the coverage area of the initial session time window is read, and the continuity level of the content change metric is calculated in chronological order. Locations with sparse time distribution density or abrupt changes in content change metric are identified as potential session boundary locations. By comparing the content change metric and time stamp information within the initial session time window, the session boundary location is repositioned, and the starting position of the initial session time window is adjusted forward or backward to the new boundary location, while the ending position is adjusted synchronously to obtain the session time window.
[0041] S1.5: Based on the session time window, the video frames within the window are divided into several spatiotemporal regions according to a fixed spatial division and a fixed time span, and a spatiotemporal region identifier is generated for each spatiotemporal region.
[0042] Specifically, the video frames within the session time window are read and sorted by time identifier information. Each video image data frame is divided into multiple spatial blocks according to a fixed spatial division method. At the same time, continuous video frames are grouped on the time axis according to a fixed time span to form multiple time segment combinations. The spatial blocks and time segment combinations are cross-combined to obtain several spatiotemporal regions. For each spatiotemporal region, the corresponding spatial location index, time segment start and end index, and session time window identifier information are recorded. The spatiotemporal region identifier is generated by concatenating the spatial location index, time segment start and end index, and session time window identifier information in a fixed field order.
[0043] It should be noted that a spatiotemporal region refers to a video area representing a specific time and spatial range within a session time window, obtained by combining fixed spatial divisions and time spans; a spatiotemporal region identifier is a unique identifier assigned to each spatiotemporal region.
[0044] S1.6: Based on the spatiotemporal region identifier, statistically analyze the changes in prediction residuals of adjacent frames, the degree of motion change, the degree of texture randomness, the continuity of edge structure, the stability of contour, and the consistency of brightness for each spatiotemporal region, forming a generative statistical record and a constrainable statistical record.
[0045] Specifically, the video frame set corresponding to the spatial block and time period combination within the session time window is located according to the spatiotemporal region identifier. The prediction residual difference between adjacent video frames is calculated frame by frame in chronological order within the video frame set to obtain the prediction residual change between adjacent frames. The displacement vector of the pixel or pixel block is obtained by performing motion analysis method of optical flow estimation between consecutive frames. The motion trend and motion amplitude are statistically analyzed based on the direction consistency and displacement magnitude of the displacement vector to obtain the degree of motion change. The dispersion of grayscale or color distribution and the similarity of distribution between adjacent frames are calculated by performing histogram statistics on the grayscale value or color component of the pixel in the block to obtain the degree of texture randomness.
[0046] Edge responses are extracted at the same spatial block locations, and edge responses of adjacent video frames are matched and compared to obtain edge structure continuity. Contour stability is obtained by comparing the positional offset and shape similarity of the contour lines of adjacent video frames. Brightness consistency is obtained by comparing the mean brightness and brightness fluctuation amplitude of adjacent video frames. The changes in prediction residuals, motion changes, and texture randomness of adjacent frames are aggregated in a fixed field order to form a generative statistical record. Edge structure continuity, contour stability, and brightness consistency are aggregated in a fixed field order to form a constrainable statistical record.
[0047] The prediction residual is the difference between the predicted frame and the actual frame obtained based on inter-frame motion compensation prediction.
[0048] S1.7: Based on the constrainable statistical records, collect the existence of target categories, the occurrence of text, and the coverage ratio of task-related areas in each spatiotemporal region to form a semantic keyness strength record.
[0049] Specifically, based on the spatiotemporal region identifiers, the video frame sets corresponding to each spatiotemporal region are located, and constrained statistical records are read to lock the structurally stable spatial location range. Within the video frame set, the target contour is identified and extracted, and the directional gradient histogram is compared with a predefined target category model to count the frequency of occurrence of the target category and obtain the existence of the target category. For text regions, edge detection is used to identify the stroke shape of characters, and combined with OCR technology, the occurrence of text is counted. For task-related regions, each spatiotemporal region is covered and labeled according to a preset spatial range, and the coverage ratio of task-related regions is calculated. The existence of target categories, the occurrence of text, and the coverage ratio of task-related regions are collected in a fixed field order to form a semantic key strength record.
[0050] It should be noted that the target category model is obtained through offline training, the prior region rules are provided by the task template configuration file and can be updated with task switching; the target category model pre-selects representative target samples according to the task scenario, and forms a standard feature template by statistically summarizing their contour morphology, directional gradient distribution and structural features; the value range is limited to a uniform numerical range, such as normalized to 0 to 1.
[0051] The preset spatial range is set by the prior region rules corresponding to the task type, and the range of values is taken from the spatial region with high correlation to the task target within the normalized coordinate range of the video frame.
[0052] S1.8: Based on the generative statistical records, constrainable statistical records, and semantic key strength records, the spatiotemporal regions are jointly determined and mapped to different semantic recoverability levels, and a semantic recoverability level map is generated according to spatial and temporal locations.
[0053] Specifically, based on the spatiotemporal region identifier, the generability statistical records, constraint statistical records, and semantic key intensity records corresponding to each spatiotemporal region are read one by one. The statistical items in the generability statistical records that represent the degree of change in prediction residuals, motion change, and texture randomness are summarized and compared to form generability judgment information. The statistical items in the constraint statistical records that represent the continuity of edge structure, contour stability, and brightness consistency are summarized and compared to form constraint judgment information.
[0054] Using semantic keyness strength records as the priority criterion, the generability and constraint information are jointly judged, and corresponding semantic recoverability levels are assigned to each spatiotemporal region. The semantic recoverability levels are then filled into the corresponding positions according to the spatial location index and time period start and end index in the spatiotemporal region identifier, generating a semantic recoverability level map.
[0055] It should be noted that the semantic recoverability level represents the importance and reconstructability of each spatiotemporal region in the video content recovery process, and is evaluated based on the generability judgment information, the constraint judgment information, and the semantic criticality strength record. The semantic recoverability level is determined based on the weighted scores of the generability statistical record, the constraint statistical record, and the semantic criticality strength record falling into different intervals.
[0056] A better approach is to encode and compress video content based on brightness changes, motion intensity, or saliency features. By jointly utilizing generative statistical records, constrained statistical records, and semantic key strength records, the semantic recoverability level of spatiotemporal regions is determined, and the encoding and reconstruction processes are coordinated and adjusted based on semantic recoverability. This improves the reconstructability and semantic preservation of video content under extremely low bandwidth conditions.
[0057] S2: Based on the semantic recoverability level map, the video content is divided into video regions with different recoverability levels, and the video content is organized hierarchically to generate minimum content data and enhanced description data. S2.1: Based on the semantic recoverability level map, classify spatiotemporal regions into low recoverability levels and high recoverability levels.
[0058] Specifically, according to the spatiotemporal region identifier, the corresponding spatial location index and time period start and end index in the semantic recoverable level map are located one by one. The semantic recoverable level of the corresponding location record is read and compared with the pre-set level grouping rules. When the spatiotemporal region simultaneously exhibits large prediction residuals, drastic motion changes, high texture randomness and weak structural constraints, and low semantic criticality, it is judged as a low recoverable level. The semantic recoverable level record that meets the low recoverable level condition is marked as a spatiotemporal region of low recoverable level. When the prediction residuals are small, the motion is stable, the texture regularity is strong, the structural constraints are clear and the semantic criticality is high, it is judged as a high recoverable level. The semantic recoverable level record that meets the high recoverable level condition is marked as a spatiotemporal region of high recoverable level.
[0059] It should be noted that the pre-set level grouping rules are set based on the comprehensive value range of the generative statistical records, constrainable statistical records, and semantic criticality intensity records corresponding to each spatiotemporal region in the semantic recoverable level map; the level grouping rules can be implemented through weighted scores G, where G comprehensively considers prediction residuals, motion changes, texture regularity, structural constraints, and semantic criticality, and scores falling into the intervals [0, τ1], [τ1, τ2], and [τ2, 1] correspond to low, medium, and high recoverable levels, respectively.
[0060] S2.2: For spatiotemporal regions with low recoverability, extract basic pixel information, contour information, and motion trend information from the original video frame sequence and aggregate them to form baseline content data.
[0061] Specifically, for spatiotemporal regions with low recoverability, the corresponding frame range and spatial position are located from the original video frame sequence based on the corresponding spatiotemporal region identifier. Within the frame range, pixel brightness and color information are read frame by frame to form basic pixel information. At the same time, edge contours in adjacent video frames are detected and matched to extract contour information. Motion trend information is obtained by calculating the direction and amplitude of pixel position changes in consecutive frames. The basic pixel information, contour information, and motion trend information are integrated according to time order and spatial position to form the baseline content data.
[0062] S2.3: For spatiotemporal regions with high recoverability, extract structural guidance information, local residual information, and semantic guidance information from the original video frame sequence, and aggregate them to form enhanced description data.
[0063] Specifically, for spatiotemporal regions with high recoverability, the corresponding frame range and spatial location are located from the original video frame sequence based on the corresponding spatiotemporal region identifier. Within the frame range, the structural features of the local region are extracted frame by frame to form structural guidance information. At the same time, the residual difference between adjacent video frames is calculated to extract local residual information. Semantic guidance information is obtained by combining the target category, text information or task-related content in the video frame. The structural guidance information, local residual information and semantic guidance information are integrated in chronological order and spatial location to form enhanced descriptive data.
[0064] S3: Under the constraint of extremely low target bandwidth, perform content-adaptive bitrate allocation on the guaranteed content data and the enhanced description data, and compress and encode them separately to form the guaranteed bitstream and the enhanced bitstream.
[0065] S3.1: Perform data organization and field arrangement for the guaranteed content data and enhanced description data respectively, and arrange various types of information in a structured manner according to time order and spatial location to form standardized input data.
[0066] Specifically, based on the spatiotemporal region identifiers, the guaranteed content data and enhanced description data are classified according to time order and spatial location. The basic pixel information, contour information, and motion trend information in the guaranteed content data are grouped into the guaranteed data field group, and the structural guidance information, local residual information, and semantic guidance information in the enhanced description data are grouped into the enhanced data field group. The relevant information in each field group is then standardized according to a unified format and arranged in field order to form standardized input data.
[0067] S3.2: Based on the normalized input data, under the constraint of extremely low target bandwidth, perform priority-protected compression encoding on the guaranteed content data to form a guaranteed bitstream.
[0068] Specifically, based on standardized input data, the original video frame sequence is processed by reducing resolution or frame rate to generate a baseline video sequence; according to the basic pixel information, contour information and motion trend information of each frame, priorities are assigned according to content importance, and priority-protected compression encoding is performed; the AV1 encoding algorithm is selected to encode the baseline video sequence to reduce the amount of data while maintaining the basic clarity and integrity of the video content, generating a baseline bitstream.
[0069] For example, after downsampling the original 1080p, 30-frame video to a minimum 480p, 10-frame video sequence, it is compressed using AV1 encoding at a bandwidth of 64kbps, so that the main target outlines and motion trends can still be identified, generating a minimum bitstream.
[0070] S3.3: Based on the remaining available bandwidth after deducting the minimum bitstream usage, perform adaptive compression encoding on the enhanced description data to form an enhanced bitstream.
[0071] Specifically, within the remaining bandwidth after deducting the minimum bitstream usage, the compressibility and bandwidth requirements of the enhanced information are evaluated based on the structural guidance information, local residual information, and semantic guidance information in the enhanced information. Adaptive compression coding is performed according to the remaining bandwidth to encode the enhanced description data. Entropy coding or encapsulation into an encodeable bitstream followed by AV1 encoding can be used to reduce the amount of data while maintaining the accuracy of intelligent reconstruction processing, thereby generating the enhanced bitstream.
[0072] S4: Receive and decode the baseline bitstream and the enhanced bitstream at the decoding end. Simultaneously, combine the semantic guidance information carried in the enhanced bitstream and perform intelligent reconstruction processing according to the corresponding recoverable level constraints to generate a target quality video sequence.
[0073] S4.1: Semantic guidance information refers to auxiliary descriptive information extracted from the original video frame sequence and transmitted with the enhanced description data.
[0074] Specifically, semantic guidance information refers to extracting structural features, edge contours, motion patterns, and scene change information of target objects from the original video frame sequence by analyzing the spatial features and temporal continuity of video frames, and using structural features, edge contours, motion patterns, and scene change information as auxiliary descriptive information.
[0075] The semantic guidance information includes the category label of the target object and its confidence level, the summary information of the text content, the spatial mask identifier of the task-related area, and the identifier of key frames or key time periods.
[0076] S4.2: The decoding end receives the minimum bitstream and the enhanced bitstream, and prioritizes the decoding of the minimum bitstream to obtain the basic video sequence.
[0077] Specifically, after receiving the minimum bitstream and the enhanced bitstream, the decoding end performs decoding processing on the minimum bitstream. By using the same AV1 encoding algorithm as during encoding, the minimum bitstream is decoded frame by frame to extract the basic pixel information, contour information and motion trend information contained therein, and a basic video sequence is generated.
[0078] S4.3: Based on the basic video sequence, and according to the structural guidance information, local residual information and semantic guidance information carried in the enhanced bitstream, perform constrained intelligent reconstruction processing on the corresponding spatiotemporal regions in the basic video sequence.
[0079] Specifically, based on the base video sequence, the decoding end performs intelligent reconstruction processing on the corresponding spatiotemporal regions in the base video sequence according to the structural guidance information, local residual information and semantic guidance information carried in the enhanced bitstream; by analyzing the structural guidance information to obtain the structural features of the spatiotemporal region, using the local residual information to repair the missing details in the video frame, and constraining the spatiotemporal region in the reconstruction process according to the semantic guidance information.
[0080] The constraints include structural constraints, detail recovery constraints, and semantic consistency constraints.
[0081] S4.4: The spatiotemporal regions after intelligent reconstruction are fused and their consistency corrected according to their spatial location and temporal order to form a target quality video sequence.
[0082] Specifically, based on the spatial location index and time period start and end index of each spatiotemporal region, the spatiotemporal regions are arranged in chronological order; consistency correction is performed on the boundaries and adjacent regions of each spatiotemporal region; by comparing the reconstructed content of adjacent spatiotemporal regions, detailed adjustments and transition processing are performed; and all processed spatiotemporal regions are merged to form a target quality video sequence.
[0083] S5: Based on the target quality video sequence, perform task completion evaluation and adaptively update the semantic recoverability level classification rules and bitrate allocation strategy in the session time window.
[0084] S5.1: Within the session time window, statistical analysis is performed on the recognition consistency, temporal stability, and reconstruction continuity of the corresponding regions in the target quality video sequence to form input information.
[0085] Specifically, within the session time window, based on the target quality video sequence, the consistency of recognition, temporal stability, and reconstruction continuity of the corresponding regions are statistically calculated frame-by-frame to assess the consistency between the recognition result of each spatiotemporal region and the expected target, thus evaluating recognition consistency; the temporal order of each spatiotemporal region and the changes between preceding and following frames are analyzed to calculate temporal stability; the spatiotemporal regions after intelligent reconstruction are compared with the original video content to evaluate reconstruction continuity; and the statistical results of recognition consistency, temporal stability, and reconstruction continuity are summarized to form input information.
[0086] The expected target is obtained from the prior semantic model or task template corresponding to the task type.
[0087] The formula for identifying consistency is: ; In the formula, Indicates the first Consistency index for identifying spatiotemporal regions This indicates the identification of a consistency identifier. Indicates the total number of video frames. Indicates in video frame In the middle, the first Actual identification information for each spatiotemporal region Indicates in video frame In the middle, the first Identification information of the expected target or object in a spatiotemporal region. Indicates the expected goal or target. It represents the identification information; the identification consistency index is calculated from the difference between the actual identification information and the expected target.
[0088] The identification information refers to the category labels and corresponding confidence scores obtained from the target objects or text content in the video.
[0089] The formula for timing stability is: ; In the formula, Indicates the first Temporal stability index for each spatiotemporal region The temporal stability identifier indicates the sequence of events in a video frame. In the middle, the first Feature values of a spatiotemporal region Represents the eigenvalue.
[0090] Reconstructing continuity, the formula is: ; In the formula, Indicates the first Reconstruction continuity indicators for each spatiotemporal region Indicates the reconstruction continuity identifier, Indicates in video frame In the middle, the first Reconstructed values for each spatiotemporal region This represents the reconstructed value; the reconstruction continuity index is calculated from the difference between adjacent reconstructed values.
[0091] Among them, the consistency formula, the time-series stability formula, and the time-series stability formula are all based on the same type of feature values for difference calculation, and all have the same physical unit or dimensionless feature.
[0092] It should be noted that the recognition result refers to the target object or scene information obtained through video frame analysis.
[0093] S5.2: Based on the input information, perform task completion evaluation processing on the target quality video sequence of the current session time window to obtain task completion evaluation information.
[0094] Specifically, based on the input information, when evaluating the task completion of the target quality video sequence within the current session time window, the system calculates and compares indicators such as recognition consistency, temporal stability, and reconstruction continuity of the corresponding regions in the video sequence. Based on accuracy, stability, and continuity, the system assesses the video sequence's ability to support the task objective, thus obtaining evaluation indicators. Based on these evaluation indicators, the system determines the effectiveness of the video sequence in achieving the task objective and generates task completion evaluation information.
[0095] S5.3: Based on the task completion assessment information, modify the rules for defining semantic recoverability levels by revising the generability statistics, constraint statistics, and semantic criticality strength records.
[0096] Specifically, when setting semantic recoverability level classification rules based on task completion assessment information to modify the generability statistics, constraint statistics, and semantic key strength records, the generability statistics for each spatiotemporal region are read one by one. The changes in prediction residuals, the degree of motion changes, and the degree of texture randomness are compared and analyzed to characterize the generation difficulty features of the spatiotemporal region. The corresponding constraint statistics are read, and the continuity of edge structure, contour stability, and brightness consistency are comprehensively analyzed to describe the constraint characteristics of the spatiotemporal region in the reconstruction process. The semantic key strength records are introduced to comprehensively evaluate the existence of target categories, the occurrence of text, and the coverage ratio of task-related areas to obtain semantic importance. The semantic importance is used as a moderating factor to jointly modify the generation difficulty features and constraint characteristics to form semantic recoverability level classification rules.
[0097] A better approach is to evaluate and optimize video quality compared to using single-dimensional features. By combining generative, constrained, and semantically critical features, the semantic recoverability classification rules are dynamically adjusted based on task requirements, thereby optimizing the video reconstruction effect under low bandwidth conditions.
[0098] S5.4: Based on the task completion assessment information, the semantic recoverability level classification rules are adaptively adjusted, and the semantic recoverability level determination is re-executed on the video content to obtain the corresponding semantic recoverability level distribution.
[0099] Specifically, based on the recognition consistency, temporal stability, and reconstruction continuity provided in the task completion assessment information, the semantic recoverability level classification rules are adjusted; according to the adjusted semantic recoverability level classification rules, the semantic recoverability level of the video content is re-determined, each spatiotemporal region is re-evaluated and mapped to a new semantic recoverability level, and an updated semantic recoverability level distribution is generated.
[0100] S5.5: Adaptively update the bitrate allocation strategy for the minimum bitstream and the enhanced bitstream based on the semantically recoverable level distribution.
[0101] Specifically, based on the semantic recoverability level distribution, when adaptively updating the bitrate allocation strategy for the minimum bitrate stream and the enhanced bitrate stream, the bitrate allocation is adjusted according to the semantic recoverability level of different spatiotemporal regions. For regions with low recoverability levels, the bandwidth of the minimum bitrate stream is allocated first to ensure the transmission of basic information. For regions with high recoverability levels, the remaining bandwidth is used to allocate the enhanced bitrate stream to optimize the improvement of video quality and complete the adaptive update.
[0102] This embodiment also provides a video compression and intelligent reconstruction system for extremely low bandwidth, including: a session construction module, a hierarchical organization module, a bitrate allocation module, an intelligent reconstruction module, and a feedback update module; the session construction module is used to acquire the original video frame sequence and construct a session time window, perform semantic analysis of the video content within the session time window, and generate a semantic recoverability level map; the hierarchical organization module is used to divide the video content into video regions with different recoverability levels according to the semantic recoverability level map, and perform hierarchical organization of the video content to generate minimum content data and enhanced description data; the bitrate allocation module is used for... Under the constraint of extremely low target bandwidth, content-adaptive bitrate allocation is performed on the guaranteed content data and enhanced description data, and they are compressed and encoded separately to form the guaranteed bitstream and the enhanced bitstream. The intelligent reconstruction module is used to receive the guaranteed bitstream and the enhanced bitstream at the decoding end and decode the guaranteed bitstream. At the same time, it combines the semantic guidance information carried in the enhanced bitstream and performs intelligent reconstruction processing according to the corresponding recoverability level constraint to generate the target quality video sequence. The feedback update module is used to perform task completion evaluation based on the target quality video sequence and adaptively update the semantic recoverability level division rules and bitrate allocation strategy in the session time window.
[0103] In summary, this invention achieves this by: constructing a semantically recoverable hierarchy and performing tiered processing of video content, ensuring the stable reconstruction of key semantic regions with limited bandwidth resources; jointly characterizing the ease of generation, structural constraints, and semantic importance of video spatiotemporal regions, giving the encoding and reconstruction processes clear semantic directionality, thereby maintaining the understandability and task availability of video content even under extremely low bandwidth conditions; and enabling higher-precision constraints and recovery of key regions during the video reconstruction process, reducing the consumption of transmission resources by invalid information, and improving the overall reconstruction results in terms of continuity and stability.
[0104] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A video compression and intelligent reconstruction method for extremely low bandwidth, characterized by: include, The original video frame sequence is acquired and a session time window is constructed. Semantic analysis of the video content within the session time window is performed to generate a semantic recoverability level map. Based on the semantic recoverability level map, the video content is divided into video regions with different recoverability levels, and the video content is organized hierarchically to generate minimum content data and enhanced description data; Under the constraint of extremely low target bandwidth, content-adaptive bitrate allocation is performed on the minimum content data and the enhanced description data, and they are compressed and encoded separately to form the minimum bitstream and the enhanced bitstream; The decoder receives and decodes the baseline bitstream and the enhanced bitstream, and combines the semantic guidance information carried in the enhanced bitstream to perform intelligent reconstruction processing according to the corresponding recoverable level constraints to generate a target quality video sequence. Based on the task completion evaluation of target quality video sequences, the semantic recoverability level classification rules and bitrate allocation strategies in the session time window are adaptively updated.
2. The video compression and intelligent reconstruction method for extremely low bandwidth as described in claim 1, characterized in that: The original video frame sequence includes multiple frames of video image data and time stamp information.
3. The video compression and intelligent reconstruction method for extremely low bandwidth as described in claim 2, characterized in that: The specific steps for constructing the session time window are as follows: The original video frame sequence is sequentially traversed according to the time stamp information, and the brightness, chroma and spatial structure changes between adjacent video frames are calculated frame by frame to form a content change metric. Based on the statistical distribution characteristics of content change measurement, a change judgment threshold is obtained. The position where the content change measurement exceeds the change judgment threshold is judged as the session boundary and divided into the initial session time window. For the initial session time window, the temporal distribution density and content change continuity of the video frames within the window are statistically analyzed, and the start and end positions of each initial session time window are corrected to obtain the session time window.
4. The video compression and intelligent reconstruction method for extremely low bandwidth as described in claim 3, characterized in that: The specific steps for generating the semantically recoverable level map are as follows. Based on the session time window, the video frames within the window are divided into several spatiotemporal regions according to a fixed spatial division and a fixed time span, and a spatiotemporal region identifier is generated for each spatiotemporal region. Based on the spatiotemporal region identifier, the changes in prediction residuals, motion changes, texture randomness, edge structure continuity, contour stability, and brightness consistency of adjacent frames are statistically analyzed for each spatiotemporal region, forming a generative statistical record and a constrainable statistical record. Based on the constrained statistical records, the existence of target categories, the occurrence of text, and the coverage ratio of task-related areas are collected in each spatiotemporal region to form a record of semantic key intensity. Based on the generative statistical records, constrainable statistical records, and semantic key strength records, each spatiotemporal region is jointly determined and mapped to different semantic recoverability levels, and a semantic recoverability level map is generated according to spatial and temporal locations.
5. The video compression and intelligent reconstruction method for extremely low bandwidth as described in claim 4, characterized in that: The specific steps for generating the guaranteed content data and enhanced description data are as follows: Based on the semantic recoverability level map, the spatiotemporal regions are marked with different levels, namely, spatiotemporal regions with low recoverability level and spatiotemporal regions with high recoverability level. For spatiotemporal regions with low recoverability, basic pixel information, contour information, and motion trend information are extracted from the original video frame sequence and aggregated to form baseline content data. For spatiotemporal regions with high recoverability, structural guidance information, local residual information, and semantic guidance information are extracted from the original video frame sequence and aggregated to form enhanced description data.
6. The video compression and intelligent reconstruction method for extremely low bandwidth as described in claim 5, characterized in that: The specific process for forming the minimum bitstream and the enhanced bitstream is as follows. The data on guaranteed content and enhanced description are sorted and organized into fields, and the information is arranged in a structured manner according to time order and spatial location to form standardized input data. Based on the normalized input data, under the constraint of extremely low target bandwidth, priority-protected compression encoding is performed on the guaranteed content data to form a guaranteed bitstream; Based on the remaining available bandwidth after deducting the minimum bitstream usage, adaptive compression encoding is performed on the enhanced description data to form the enhanced bitstream.
7. The video compression and intelligent reconstruction method for extremely low bandwidth as described in claim 6, characterized in that: The semantic guidance information refers to the auxiliary descriptive information extracted from the original video frame sequence and transmitted with the enhanced descriptive data.
8. The video compression and intelligent reconstruction method for extremely low bandwidth as described in claim 7, characterized in that: The specific process for generating the target quality video sequence is as follows. The decoding end receives the minimum bitstream and the enhanced bitstream, and prioritizes the decoding of the minimum bitstream to obtain the basic video sequence; Based on the basic video sequence, and according to the structural guidance information, local residual information and semantic guidance information carried in the enhanced bitstream, constrained intelligent reconstruction processing is performed on the corresponding spatiotemporal regions in the basic video sequence. The spatiotemporal regions after intelligent reconstruction are fused and corrected for consistency according to their spatial location and time sequence to form a target quality video sequence.
9. The video compression and intelligent reconstruction method for extremely low bandwidth as described in claim 8, characterized in that: The adaptive update of the semantic recoverability level classification rules and bitrate allocation strategy within the session time window is as follows: Within the session time window, the recognition consistency, temporal stability, and reconstruction continuity of the corresponding regions in the target quality video sequence are statistically analyzed to form input information; Based on the input information, perform task completion evaluation processing on the target quality video sequence in the current session time window to obtain task completion evaluation information; Based on task completion assessment information, modify the generability statistics record, constraint statistics record, and semantic criticality strength record to set semantic recoverability level classification rules; Based on the task completion assessment information, the semantic recoverability level classification rules are adaptively adjusted, and the semantic recoverability level is re-determined for the video content to obtain the corresponding semantic recoverability level distribution. The rate allocation strategy for the minimum and enhanced bitstreams is adaptively updated based on the semantically recoverable level distribution.
10. A video compression and intelligent reconstruction system for extremely low bandwidth, based on the video compression and intelligent reconstruction method for extremely low bandwidth as described in any one of claims 1 to 9, characterized in that: It includes a session building module, a hierarchical organization module, a bitrate allocation module, an intelligent reconstruction module, and a feedback update module; The session construction module is used to acquire the original video frame sequence and construct a session time window, perform semantic analysis of the video encoding on the video content within the session time window, and generate a semantic recoverability level map. The hierarchical organization module is used to divide the video content into video regions with different recoverability levels based on the semantic recoverability level map, and to hierarchically organize the video content to generate minimum content data and enhanced description data. The bitrate allocation module is used to perform content-adaptive bitrate allocation on the minimum content data and the enhanced description data under the constraint of extremely low target bandwidth, and to compress and encode them respectively to form the minimum bitrate stream and the enhanced bitrate stream. The intelligent reconstruction module is used to receive the minimum bitstream and the enhanced bitstream at the decoding end and decode the minimum bitstream. At the same time, it combines the semantic guidance information carried in the enhanced bitstream and performs intelligent reconstruction processing according to the corresponding recoverable level constraints to generate a target quality video sequence. The feedback update module is used to perform task completion evaluation based on the target quality video sequence and to adaptively update the semantic recoverability level classification rules and bitrate allocation strategy in the session time window.