Compression storage method and system for dynamic video
By using multimodal contextual data streams and parameterized point of interest configuration, combined with value assessment models and historical feedback calibration, the problem of single value assessment and static management strategies in existing video storage systems is solved, achieving efficient management and accuracy of dynamic video storage.
Patent Information
- Application Number
- CN202511729546.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-24
- Publication Date
- 2026-02-27
AI Technical Summary
Existing video storage systems suffer from low storage efficiency and unreliable retention of critical data due to their singular value assessment dimensions, static management strategies, and inability to adaptively optimize.
It adopts multimodal contextual data streams and parameterized point of interest configuration information, uses a value assessment model to calculate real-time basic value weights, adaptively selects video encoding parameters for compression, and combines multi-source structured metadata with associated storage. It also periodically updates the current retention value of stored video data and performs offline feedback calibration through historical business operation logs.
It achieves accuracy and comprehensiveness in multi-dimensional value assessment, ensures the identification of events with inconspicuous visual characteristics but high business value or risk, optimizes storage space utilization efficiency, and reduces the complexity of manual configuration and maintenance.
Smart Images

Figure CN121585829A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of video monitoring, in particular to a compression storage method and system of dynamic video. BACKGROUND
[0002] With the wide deployment of video monitoring systems in the fields of city management, transportation, security and the like, a large amount of video data is generated, which poses a great challenge to the storage and management of data. The traditional video storage system usually adopts fixed code rate coding and first-in-first-out (FIFO) coverage strategy. This way fails to distinguish the actual value of the video content, resulting in a large number of meaningless still or redundant pictures occupying the same storage resources as key event pictures, which not only causes serious waste of storage space, but also makes the truly valuable video evidence possibly be overwritten prematurely due to the expiration of the storage period.
[0003] To improve this situation, the prior art introduces a dynamic coding strategy based on content analysis, such as using motion detection or artificial intelligence target recognition to determine whether the picture is active, so as to use different compression rates for different pictures. However, the value judgment dimension of this kind of method is single, mainly relying on whether there is a target or movement in the visual level, which is difficult to accurately evaluate the real value of complex scenes. For example, for traffic monitoring scenes, the system may not be able to distinguish between normal vehicle traffic and abnormal stillness caused by accidents; for environmental monitoring, the system cannot include non-visual risk signals such as sudden visibility drop into the value evaluation system.
[0004] In addition, the storage management strategy of the prior art is usually static. Even if dynamic coding is used, once the video is stored, its life cycle management is still mechanical, lacking consideration of the value decay over time according to the content type. More importantly, the rules and parameters preset by these systems are fixed and cannot be adaptively adjusted according to the actual business needs of specific deployment scenarios, resulting in a deviation between the value judgment standard of the system and the actual focus of the user after a period of operation, which requires a lot of manpower for reconfiguration and maintenance.
[0005] Therefore, the present application proposes a compression storage method and system of dynamic video to solve the deficiencies of the prior art. SUMMARY
[0006] In view of the deficiencies of the prior art, the present application provides a compression storage method and system of dynamic video, which solves the problems of low storage efficiency and unreliable key data retention caused by the single value evaluation dimension, static management strategy and inability to adaptively optimize of the existing video storage system.
[0007] To achieve the above purpose, the present application realizes the following technical scheme: a compression storage method of dynamic video, comprising the following steps: Based on the multi-modal context data stream and the parameterized point of interest configuration information, a real-time basic value weight of the current video content is calculated by using a value evaluation model; According to the calculated real-time basic value weight, a matched video encoding parameter is adaptively selected to compress the video data, and the compressed video data is stored in association with the multi-source structured metadata; According to the storage duration of the stored video data and the corresponding dynamic value decay coefficient, the current retention value of the stored video data is regularly updated, and when the calculated current retention value is lower than a preset system cleaning threshold, a data pruning operation is performed; The historical business operation log of the system is extracted as a supervision signal to perform offline feedback calibration on the weight parameters in the value evaluation model.
[0008] Preferably, the step of calculating the real-time basic value weight of the current video content includes: weighting and fusing calculation of an AI perception score representing the importance of visual content, a POI activation score representing the trigger state of business rules, and an abnormal data anomaly score representing the degree of deviation of the environment state from the normal benchmark, to obtain the real-time basic value weight.
[0009] Preferably, the step of calculating the AI perception score includes: target detection on the video frame to obtain the category label and detection confidence of the target; based on the obtained category label, an importance coefficient is obtained in a preset category priority mapping table; the obtained detection confidence and the obtained importance coefficient are multiplied, and the product results of all targets in the frame are accumulated to obtain the AI perception score.
[0010] Preferably, the step of calculating the POI activation score includes: obtaining a user-defined point of interest configuration, the point of interest configuration including a spatial range, a basic business priority weight, and a parameterized rule set; the parameterized rule set includes a target speed threshold, a target stay time threshold in the area, a target moving direction limit, or a specific object category filtering condition; based on the multi-modal context data stream, it is judged whether the logic conditions in the parameterized rule set are met, and the basic business priority weight is combined to calculate the POI activation score.
[0011] Preferably, the step of adaptively selecting a matched video encoding parameter includes: comparing the real-time basic value weight with a preset high value judgment threshold and a low value judgment threshold; according to the comparison result, one corresponding video encoding parameter is selected from a multi-level encoding configuration scheme composed of a high-definition mode code rate, a standard-definition mode code rate, and a low-flow mode code rate.
[0012] Preferably, the step of periodically updating the current retention value of the stored video data comprises: extracting a subject event category label from the structured context metadata of the stored video data; looking up and determining a dynamic value decay coefficient in a preset value decay strategy table according to the extracted subject event category label; and calculating the current retention value by using an exponential decay model in combination with an initial real-time base value weight of the stored video data, a storage duration, and the determined dynamic value decay coefficient.
[0013] Preferably, the step of offline feedback calibration of the weight parameters in the value evaluation model comprises: extracting retrieval behavior, playback behavior, and recovery behavior from historical business operation logs to quantitatively calculate a real business value of a video segment; constructing a loss function based on a deviation between the calculated real business value and a retention value predicted by the value evaluation model; and iteratively updating the fusion weight coefficient and the dynamic value decay coefficient in the value evaluation model based on the constructed loss function by using a gradient descent algorithm.
[0014] Preferably, before calculating the real-time base value weight, the method further comprises: simultaneously collecting real-time video streams and heterogeneous sensor data of a monitoring area; performing real-time analysis on the real-time video streams to obtain a target density in a field of view; and when the obtained target density in the field of view exceeds a preset target density triggering threshold, performing reverse adjustment on collection parameters of the heterogeneous sensor data, and switching a sampling frequency of the heterogeneous sensor from a base sampling frequency to a high-frequency sampling upper limit.
[0015] Preferably, the data pruning operation comprises a hierarchical cleaning mode, in which, when the current retention value is lower than a system cleaning threshold but higher than a metadata retention threshold, only video track data is deleted, and structured context metadata and key frame indexes are retained.
[0016] The application also provides a dynamic video compression storage system, comprising: a content value dynamic weight distribution and prediction module configured to calculate a real-time base value weight of current video content by using the value evaluation model based on multi-modal context data streams and parameterized interest point configuration information, and to perform offline feedback calibration of weight parameters in the value evaluation model by using historical business operation logs of the system as a supervision signal; a compression storage module configured to adaptively select matching video coding parameters to compress video data according to the calculated real-time base value weight, and to store the compressed video data in association with multi-source structured metadata; A period management module is configured to periodically update a current retention value of the stored video data according to a storage duration of the stored video data and a corresponding dynamic value decay coefficient, and perform a data pruning operation when the calculated current retention value is lower than a preset system cleaning threshold.
[0017] The application provides a dynamic video compression storage method and system. 1、The application fuses AI perception score, POI activation score and heterogeneous data anomaly score to perform multi-dimensional comprehensive evaluation on the value of video content, which overcomes the limitations of single visual analysis, accurately identifies events with high business value or risk level but unobvious visual features (for example, environmental mutations perceived by heterogeneous data or slow congestion meeting specific business rules), and thus ensures the accuracy and comprehensiveness of value evaluation, providing a reliable basis for subsequent differentiated compression storage.
[0018] 2、The application proposes a dynamic value decay model based on content initial value and event type to perform full life cycle management on the stored video data, which can match a very low decay coefficient for video data of key events such as traffic accidents to achieve long-term preservation, and match a higher decay coefficient for video of redundant content such as normal traffic to accelerate cleaning, so that the fine management strategy replaces the traditional first-in-first-out mechanism, effectively prolongs the retention time of key evidence, and maximizes the utilization efficiency of storage space.
[0019] 3、The application establishes a closed-loop feedback calibration mechanism based on historical business operation logs, which can quantitatively evaluate the accuracy of previous value judgment by analyzing actual operations such as user search, playback and recovery of video, and automatically update the fusion weight coefficient and decay coefficient in the value evaluation model using gradient descent algorithm. This mechanism enables the system to have the ability of adaptive evolution, continuously meets the real business needs of specific scenarios during use, and significantly reduces the complexity and cost of manual configuration and later maintenance. BRIEF DESCRIPTION OF DRAWINGS
[0020] Figure 1 FIG. 1 is a system structure diagram of the application; Figure 2 FIG. 2 is a method flow diagram of the application; Figure 3 FIG. 3 is a code rate and value grading mapping curve diagram of the application; Figure 4 FIG. 4 is a value decay and life cycle management curve diagram of the application.
[0021] Wherein, 101, video acquisition module; 102, multi-modal heterogeneous data association and fusion module; 103, AI analysis module; 104, POI coding module; 105, content value dynamic weight distribution and prediction module; 106, compressed storage module; 107, periodic management module; 108, playback display module. DETAILED DESCRIPTION
[0022] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the specification of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0023] Referring to the drawings Figure 1 , Figure 1 It is a schematic diagram of a compressed storage system structure of dynamic video according to the present application; the present application provides a compressed storage system of dynamic video, which comprises: The video acquisition module 101 is configured to obtain real-time video stream data of a monitoring area, for transmitting the collected original video data to the downstream for processing.
[0024] The multi-modal heterogeneous data association and fusion module 102 is used for receiving video stream data and environmental data, traffic flow data or event interface data from external heterogeneous sensors; the multi-modal heterogeneous data association and fusion module 102 performs timestamp alignment and spatial coordinate mapping operations, integrates video data and heterogeneous data into multi-modal context data stream; the multi-modal heterogeneous data association and fusion module 102 has a reverse adjustment signal receiving port, which is used for adjusting the sampling frequency or data transmission priority of the heterogeneous sensor according to the received feedback instruction.
[0025] The AI analysis module 103 receives the fused multi-modal context data stream; the AI analysis module 103 is configured to run a target detection algorithm to identify the target object class, confidence and position coordinates in the video frame, and generate a structured analysis result; the AI analysis module 103 is also configured to monitor the target density, and send a reverse adjustment instruction to the multi-modal heterogeneous data association and fusion module 102 when the target density exceeds a preset threshold.
[0026] The POI coding module 104 is configured to provide a human-computer interaction interface for receiving user input of point of interest definition information; the POI coding module 104 supports user setting of the spatial range, basic label and parameterized trigger rule of POI; the parameterized POI rule data generated by the POI coding module 104 is transmitted to the content value dynamic weight distribution and prediction module 105.
[0027] The content value dynamic weight distribution and prediction module 105 is configured to calculate the basic value weight of the current video content based on the multi-modal context data stream, the AI analysis result and the parameterized POI rule.
[0028] The content value dynamic weight distribution and prediction module 105 is further configured to predict the remaining value of the video content in combination with a time decay function, and calibrate the parameters of the internal calculation model using historical business feedback data; the content value dynamic weight distribution and prediction module 105 outputs real-time value weight instructions to the compression storage module 106 and the period management module 107.
[0029] The compression storage module 106 is configured to select corresponding encoding parameters from a preset code rate set to compress the video data according to the received value weight instructions; the compression storage module 106 further encapsulates the compressed video data and the structured metadata in time alignment and stores them in the storage medium.
[0030] The period management module 107 is configured to monitor the storage duration of the video data in the storage medium; the period management module 107 updates the current value of the stored video according to the value weight and the decay strategy provided by the content value dynamic weight distribution and prediction module 105, and performs a clipping or deletion operation when the current value is lower than a preset threshold; the period management module 107 further sends the clipping operation record as feedback data to the content value dynamic weight distribution and prediction module 105.
[0031] The playback display module 108 is configured to read the video data and the associated structured metadata from the storage medium in response to a retrieval request, and superimpose the POI information and the heterogeneous environment data on the video screen on the display interface.
[0032] Referring to the accompanying drawings Figure 2 , Figure 2 is a flow chart of a dynamic video compression storage method according to the present application. The present application provides a dynamic video compression storage method, comprising the following steps: S100, real-time video stream and heterogeneous sensor data of a monitored area are collected in parallel, and the collection parameters of the heterogeneous sensor data are adjusted in reverse based on the target detection feedback signal generated in the subsequent steps; S200, deep learning algorithm is used to analyze the video frame in real time to obtain structured detection result, and the real-time video stream, the structured detection result and the heterogeneous sensor data are time stamped and aligned and spatially coordinate mapped to generate multi-modal context data stream; S300, user-defined parameterized interest point (POI) configuration information is obtained, and the business rule containing trigger condition parameters and basic weight value is parsed; S400, based on the multi-modal context data stream and the parameterized interest point configuration information, calculating a real-time basic value weight of the current video content by using a value evaluation model, and combining a time decay function to predict a retention value changing over time; S500, according to the calculated real-time basic value weight, adaptively selecting a matched video encoding parameter to compress the video data, and storing the compressed video data in association with the multi-source structured metadata; S600, according to the storage duration of the stored video data and the corresponding dynamic value decay coefficient, regularly updating the current value weight of the video data, and performing a data pruning operation when the current value weight is lower than a preset threshold; S700, extracting the historical business operation log of the system as a supervision signal to perform offline feedback calibration on the weight parameters in the value evaluation model.
[0033] To further illustrate the implementation of each technical link of the present application, the following will be described in detail the implementation of each function module involved above, and its internal processing flow.
[0034] Referring to the drawings Figure 2 , step S100 mainly involves parallel collection of monitoring area data and dynamic regulation of collection strategy; this step specifically includes acquisition of video stream data, acquisition of heterogeneous sensor data, and reverse adjustment of heterogeneous data collection parameters based on feedback signals; the following will be described in detail in combination with steps S101 to S103.
[0035] In S101, real-time video stream data collection is performed; the video collection module 101 is connected to the front-end camera device through wired or wireless network protocol to acquire continuous video image frame sequences ; the collection of this video stream data follows preset basic frame rate and resolution standards, and the specific communication protocol and codec format of video collection belong to the existing technology well known to those skilled in the art, which will not be described here.
[0036] In S102, parallel acquisition of heterogeneous sensor data is performed; the multi-modal heterogeneous data association and fusion module 102 establishes a communication connection with the heterogeneous sensor set arranged in the monitoring area and external data sources through a data interface; the heterogeneous sensor data includes but is not limited to the following data types: environmental physical quantity data such as illumination intensity, environmental temperature, and visibility value; traffic running state data such as vehicle flow detected by a geomagnetic coil and radar speed data; and external event data obtained through an application program interface (API), such as manually reported event records or warning information issued by a meteorological department. In the initial state, the system polls or subscribes to read the above-mentioned heterogeneous sensor data at a preset basic sampling frequency.
[0037] In S103, heterogeneous data reverse adjustment based on target density is performed. This step is a key control link in the multimodal data acquisition process, which aims to solve the problem that a fixed sampling frequency cannot adapt to dynamic scene changes. The multimodal heterogeneous data association and fusion module 102 is equipped with a feedback signal receiving interface to receive the subsequently generated real-time target density feedback signal. When the monitoring scene is in a low-activity state, the system maintains a low basic sampling frequency to reduce the computational load and communication bandwidth usage. When a high-density target or a specific triggering event occurs in the monitoring scene, the system needs higher-precision environmental context information to assist in value assessment, which triggers the adjustment of the acquisition parameters.
[0038] The specific reverse adjustment logic is implemented through a sampling frequency control model; let... The real-time sampling frequency of the heterogeneous sensor is This frequency is controlled by the target density index fed back by the subsequent AI analysis module 103. The sampling frequency is adjusted according to the following formula: ; In the formula, This indicates the basic sampling frequency of the heterogeneous sensor, corresponding to the low-frequency acquisition mode under normal monitoring conditions; This indicates the upper limit of high-frequency sampling for heterogeneous sensors, corresponding to high-precision acquisition modes in high-value or high-risk scenarios; express The target density value within the field of view is fed back by the AI analysis module 103 at all times. This value reflects the degree of aggregation of targets of interest in the current monitoring screen. This indicates the preset target density trigger threshold, used to distinguish between normal and high-density scenes; Indicates a characteristic function, when the logical condition within the parentheses... If true, the function value is 1; otherwise, the function value is 0.
[0039] According to the above model, when the AI analysis module 103 detects the target density within the current field of view... Exceeding the preset threshold When the indicator function outputs 1, the control logic immediately changes the sampling frequency of the heterogeneous sensor from... Switch to This adjustment mechanism ensures that during high-value events (such as traffic congestion or crowd gatherings), the system can acquire environmental parameters with higher temporal resolution (such as denser vehicle speed change data or real-time visibility fluctuation data), thereby providing accurate contextual basis for subsequent content value calculation.
[0040] In addition to the adjustment of the sampling frequency, the reverse adjustment mechanism also includes the adjustment of the data transmission priority; when receiving a high-density target feedback signal, the multi-modal heterogeneous data association and fusion module 102 generates a priority control instruction, and the data packet transmission priority of the heterogeneous sensor data is raised to the highest level, ensuring that such data can be transmitted and processed preferentially in the case of limited network bandwidth, avoiding situational awareness deviation caused by data delay.
[0041] Referring to the drawings Figure 2 , step S200 mainly involves deep intelligent processing of multi-source data, aiming to build a unified data structure that can fully represent the current monitoring scene state; this step specifically includes deep learning-based visual content analysis, standardized processing of heterogeneous data, and spatio-temporal fusion of multi-modal data; steps S201 to S203 will be described in detail below.
[0042] In S201, AI situational analysis based on deep learning is performed; the AI analysis module 103 receives the real-time image frames output by the video acquisition module 101 and loads a pre-trained target detection model; the target detection model uses a convolutional neural network architecture (such as the YOLO series or Faster-R-CNN) and is configured to perform feature extraction and region proposal regression on image frames; the AI analysis module 103 performs inference calculation on the image frames and outputs a structured feature vector containing all targets of interest within the field of view ; the structured feature vector contains detailed attribute information of the target objects identified in the current frame, which can be mathematically expressed as follows: ; wherein, denotes the total number of targets detected in the current frame; denotes the class label of the target (such as vehicle, pedestrian, non-motor vehicle); denotes the confidence probability value of the target belonging to the class , with a value range of [0, 1]; denotes the bounding box position coordinates of the target in the image coordinate system, usually consisting of center point coordinates and width and height parameters ; In addition, the AI analysis module 103 calculates the target density of the current field of view based on and transmits this parameter as a feedback signal to the multi-modal heterogeneous data association and fusion module 102 in the previous step, which is used to trigger the reverse adjustment of heterogeneous data.
[0043] In S202, spatiotemporal alignment and vectorization processing of heterogeneous data are performed; the multimodal heterogeneous data association and fusion module 102 acquires data at the same time. The heterogeneous sensor raw data is used; since the sampling frequency of different sensors is inconsistent with the video frame rate, the multimodal heterogeneous data association and fusion module 102 first performs a time axis alignment operation. This operation uses a zero-order hold or linear interpolation algorithm to map the discrete sensor data to the video frame. The timestamps ensure the synchronization of multi-source data across time. Subsequently, the multimodal heterogeneous data association and fusion module 102 normalizes physical quantity data of different dimensions, eliminates dimensional differences, and generates standardized heterogeneous feature vectors. Heterogeneous feature vectors It includes environmental characteristic components (such as normalized visibility and illumination values) and traffic characteristic components (such as normalized flow rate and average speed values).
[0044] In S203, multimodal contextual data fusion mapping is performed; this step aims to combine visual perception results with non-visual environment perception results to generate a unified descriptor for subsequent value calculation; the multimodal heterogeneous data association and fusion module 102 utilizes a preset fusion mapping function. The original image data AI structured feature vectors and heterogeneous feature vectors Mapped to multimodal context vector The mathematical expression for the multimodal fusion mapping function is: ; In the formula, express The comprehensive context descriptor at any given moment, this vector represents the complete state of the current monitoring scene in the feature space, including both visual target distribution information and environmental physical state information; Represents the raw video image frame data at the current moment; This represents the set of structured features output by the AI analysis module 103, which includes target category, confidence level, and location information. This represents the feature vector of the heterogeneous sensor after time alignment and normalization. This represents a multimodal feature fusion operator. In specific implementations, this operator is configured as a feature concatenation operation, which... Statistical characteristics (such as number of targets, average confidence level) and The dimensions are concatenated to form a high-dimensional feature vector, which serves as the input parameter for the subsequent content value dynamic weight allocation and prediction module 105.
[0045] Through the above fusion processing, the system establishes a strong association between video content and external environment data, so that the subsequent value evaluation is no longer limited to the visual picture itself, but is based on global information containing environmental context to make judgments; for example, the same vehicle queue picture (similar visual features) combined with heavy rain weather data and the same vehicle queue picture combined with sunny weather data will present different feature distributions, thereby supporting the system to make differentiated value judgments.
[0046] Referring to the accompanying Figure 2 , step S300 mainly involves the analysis and digital mapping of user-defined business logic; this step aims to convert business focus points understood by humans into parameterized rules executable by computers, thereby providing a logical benchmark for subsequent automated value evaluation; this step specifically includes the definition of spatial regions, the configuration of parameterized business rules, and the calculation of rule trigger scores; steps S301 to S303 will be described in detail below.
[0047] In S301, the spatial definition and basic attribute configuration of parameterized points of interest (POI) are performed; the POI coding module 104 receives user input instructions through a graphical human-computer interaction interface; unlike traditional technologies that only perform simple name labeling or static position marking on points of interest, the POI definition in this embodiment includes geometric description in spatial dimension and priority assignment in business dimension; the user defines the boundary of the interest region by drawing a rectangular frame, polygon, or free curve on the video preview screen; the system maps the geometric figure in the screen coordinate system to a set of spatial coordinates in the video frame coordinate system; at the same time, the system receives the basic business priority weight assigned by the user for the interest region , which reflects the inherent importance of the region in the monitoring business (for example, the set value of a certain key intersection is higher than that of an ordinary road segment).
[0048] In S302, the parameterized rule configuration of complex business logic is performed; the system allows the user to associate a set of logical trigger rule sets for each defined interest region ; the trigger rule set is no longer limited to a single spatial trigger (i.e., only judging whether the target enters the region), but supports composite logical judgment based on physical attributes and time dimension; the user sets specific rule parameters through the configuration interface including but not limited to target speed threshold, target dwell time threshold in the region, target moving direction restriction, and filter conditions for specific object categories; for example, a user can define a rule as "detecting a vehicle with speed lower than 10km / h and duration longer than 300 seconds", which represents the business concept of "congestion"; the POI coding module 104 converts these natural language descriptions of business logic into mathematical expressions containing logical operators (AND / OR) and comparison operators, and stores them as metadata files in JSON or other structured formats.
[0049] In S303, rule trigger score calculation based on multi-modal context data is performed; this step is the bridge connecting user-defined rules and system automatic evaluation; the content value dynamic weight distribution and prediction module 105 reads the multi-modal context vector generated in step S200 and the rule set generated in step S302, and calculates the instantaneous activation score of each POI region in real time.
[0050] In order to accurately quantify the degree of satisfaction of the current scene to the user-defined rules, the system adopts a logical product model for calculation; for the th POI region, its POI activation score at time follows the calculation formula: In the formula, represents the activation score of the th POI region at the current time, which directly reflects the business urgency of the events in the region; represents the user's preset basic business priority weight of the POI region, which is used as the reference gain coefficient for calculation; represents the total number of parameterized rules associated with the POI region; represents the user configuration parameter (such as the specific threshold value) corresponding to the th rule; represents the multi-modal context vector containing visual target features and heterogeneous environment features; represents the logic judgment function of the th rule, whose output value range is [0, 1].
[0051] In specific implementation, the logic judgment function can be a binary Boolean function (output 1 when the condition is met, otherwise output 0), or a continuous activation function (such as Sigmoid function) for describing the confidence of condition satisfaction; the multiplication symbol in the formula Acting as a logical AND, this means that the product of this term is only high when all the user-defined parameterized conditions (e.g., satisfying location, speed, and time requirements) are simultaneously verified by multimodal contextual data, thus enabling... Achieving high scores; this calculation mechanism ensures that the system's value assessment strictly follows the complex business logic set by the user, avoids false alarms caused by single-condition triggers, and achieves accurate quantitative perception of specific business scenarios (such as "abnormal parking in a specific area").
[0052] See attached document Figure 2 Step S400 mainly involves real-time value quantification calculation based on multi-source information fusion. As the core of the system's decision-making, this step aims to transform visual features, business rule triggering status, and environmental anomaly degree into a unified scalar value, namely the real-time basic value weight. This weight directly determines the encoding quality and storage strategy of subsequent video data. This step specifically includes AI perception score calculation, heterogeneous data anomaly quantification, and multi-factor weighted fusion calculation. The following will elaborate on steps S401 to S403.
[0053] In S401, the quantitative calculation of AI perception score is performed; the content value dynamic weight allocation and prediction module 105 receives the structured feature vector from the AI analysis module 103 and evaluates the visual content value in the current video frame; this evaluation is no longer based solely on the number of targets, but combines the preset importance and recognition confidence of the target category; the system maintains a target category importance mapping table, which defines the basic score coefficients corresponding to different categories of objects (such as "pedestrian", "motor vehicle", "flame").
[0054] AI perception score The calculation follows a weighted summation logic, and its mathematical expression is as follows: ; In the formula, express The AI perception score at any given moment represents the overall importance of visual targets in the image; This indicates the total number of valid targets detected in the current frame; Indicates the first The detection confidence of each target, with a value range of [0,1], is derived from the Softmax output of the target detection network; Indicates the first Category labels for each target; Indicates category The corresponding preset importance coefficient; for example, in a security scenario, the importance coefficient of "intruder" is... The value is set higher than that of a stray cat; the coefficient supports dynamic updating through a configuration file.
[0055] Regarding the preset importance coefficient The embodiment adopts a mapping mechanism based on a lookup table to implement specific determination and acquisition logic; the content value dynamic weight distribution and prediction module 105 maintains a category priority mapping table in a non-volatile memory; the mapping table establishes an index relationship between the category label (Label) output by the AI analysis module 103 and the numerical weight, and the value range of the numerical weight is preset to [0, 1].
[0056] In the system initialization phase, the system loads the default mapping table configuration according to the current deployment scene mode; for example, when the system deployment scene is set to “highway monitoring”, in the default mapping table corresponding to this scene, the value of the category “pedestrian” is set to 0.95 (representing extremely high danger), the value of the category “stationary vehicle” is set to 0.8, and the value of the category “moving vehicle” is set to 0.5; when the system deployment scene is set to “wildlife observation”, the value of the category “animal” in the default mapping table is set to 0.9, and the value of the category “vehicle” is reduced to 0.2.
[0057] In addition, the determination logic is user-configurable; the system provides a parameter configuration interface to allow users to modify the values in the category priority mapping table in real time; after the AI analysis module 103 outputs the category label of the first target, the content value dynamic weight distribution and prediction module 105 directly queries the corresponding value (Value) in the current category priority mapping table with as the index key (Key), that is, the value of the target is obtained; if the detected target category is not defined in the mapping table, the system assigns it a preset default low weight value (such as 0.1) to prevent unknown targets from interfering with the value evaluation.
[0058] In S402, statistical calculation of heterogeneous data anomaly score is performed; in order to include environmental data in non-visual dimensions into the value evaluation system, the multi-modal heterogeneous data association and fusion module 102 performs anomaly degree analysis on heterogeneous sensor data; the system maintains the historical statistical distribution (such as the mean vector and the covariance matrix ) of each sensor data based on a sliding time window.
[0059] Heterogeneous data anomaly score Statistical distance (such as Mahalanobis distance) is used to measure the degree to which the current environmental state deviates from the normal baseline. The calculation formula is as follows: ; In the formula, express The heterogeneous data anomaly score at any given time indicates that the higher the score, the more abnormal the current environmental state (such as visibility and traffic flow) is, and the higher the potential monitoring value. express Time-normalized feature vectors of heterogeneous sensors; A vector representing the mean of sensor data within a historical time window; This represents the inverse of the historical data covariance matrix.
[0060] By calculating statistical distance, the system can uniformly quantify the degree of anomaly of different physical quantities (such as a sudden rise in temperature or a sudden drop in flow rate) without having to hard-code thresholds for a single sensor.
[0061] In step S403, a multi-factor fusion calculation of the dynamic weight of content value is performed. This step weights and fuses the AI perception score, heterogeneous anomaly score, and POI activation score calculated in step S300 to generate the final real-time basic value weight. This weight is a normalized scalar used to drive subsequent compression and storage operations.
[0062] ; The weight normalization constraint must be satisfied: ; In the formula, express Real-time basic value weight of video frames at any given moment; This represents a normalization function (such as Min-Max normalization) used to map AI scores and heterogeneous scores to the [0,1] interval to eliminate the influence of dimensions; This means taking the maximum activation score among all currently defined POI regions. Using the maximum value instead of the average value is to ensure that the triggering of a single high-priority POI region (such as an accident at a key intersection) can dominate the overall value judgment and avoid being diluted by other low-value regions. These represent the fusion weight coefficients for the AI perception dimension, POI business rules dimension, and heterogeneous environment dimension, respectively.
[0063] Fusion weighting coefficient Instead of being fixed constants, they are stored in the configuration space as tunable parameters of the system; in the initial running phase of the system, they take preset values; in the running process of the system, they will be iteratively updated according to the offline feedback calibration mechanism in step S700; by adjusting the proportion of the three coefficients, the system can adapt to different business preferences: for example, increasing the value of makes the system more inclined to record content rich in visual targets; increasing the value of makes the system strictly follow the user-defined business rules; and increasing the value of makes the system more sensitive to environmental mutations; this parameterized weighted fusion mechanism realizes the mathematical mapping from multi-dimensional data input to a single decision indicator.
[0064] Referring to the accompanying drawings Figure 2 and the accompanying drawings Figure 3 , step S500 mainly involves converting the calculated abstract value weight into a specific physical storage strategy and establishing a strong coupling relationship between video data and multi-dimensional metadata; this step specifically includes value weight-based encoding parameter adaptive decision, dynamic compression processing of video streams, and associated encapsulated storage of structured data; the following will be described in detail in combination with steps S501 to S503.
[0065] In S501, adaptive decision mapping of video encoding parameters is performed; the compression storage module 106 receives the real-time basic value weight calculated by the previous step and inputs it as a control variable into the code rate decision logic; in order to balance between limited storage space and video quality, the system does not use a fixed code rate (CBR) for encoding, but uses a dynamic variable code rate (VBR) strategy based on value classification; the compression storage module 106 is preconfigured with multiple levels of encoding configuration schemes, and different configuration schemes correspond to different target code rates, frame rates, and quantization parameters (QP).
[0066] The system determines the target encoding code rate at the current time through a piecewise mapping function; the mapping logic is shown in the following formula: ; In the formula, represents the target bit rate of the video encoder at time ; , and respectively represent the preset high-definition mode code rate (e.g. 8 Mbps), the standard-definition mode code rate (e.g. 2 Mbps), and the low-flow mode code rate (e.g. 512 Kbps); and represent the high-value judgment threshold and the low-value judgment threshold for dividing the value level, and satisfy ;when When the value exceeds a high-value threshold, the system determines that the current content has critical preservation value (such as the moment an accident occurs), and therefore allocates the highest bitrate to retain image details; when When in the middle range, allocate the standard bitrate; when When the content is below the low-value threshold, the system determines that the current content is meaningless redundancy (such as a long period of static image) and allocates the lowest bitrate to save storage space. In addition to the bitrate, the system can also adjust the keyframe interval synchronously and reduce the GOP length in the high-value segment to improve random access performance.
[0067] In step S502, dynamic compression encoding of the video data is performed; the video encoding engine in the compression storage module 106 (following standards such as H.264 / AVC, H.265 / HEVC, or AV1) outputs the data according to step S501. The instruction adjusts the encoder's rate-distortion optimization (RDO) parameters in real time; in practice, the encoder dynamically adjusts the macroblock-level quantization step size to ensure that the actual bitrate of the output video stream closely approximates the target. For low-value segments, the encoder also performs frame rate downsampling (e.g., reducing from 30fps to 5fps) to further compress the data volume; the output of this process is a compressed video base stream.
[0068] In S503, the spatiotemporal alignment and encapsulation of video stream and metadata are performed; in order to realize subsequent intelligent retrieval and contextual playback, the system must ensure that the video image and its corresponding context information are closely associated in the storage medium; the compressed storage module 106 constructs a composite storage container structure or associated file system.
[0069] This module performs time-axis aligned synchronous writing of the following three types of data: Compressed video data: Video stream from step S502; Structured contextual metadata: containing AI analysis results at the same time. (Target bounding box, category), feature vectors of heterogeneous sensors And the triggered POI rule identifier; Value assessment metadata: includes calculated real-time base value weights. Numerical value.
[0070] The specific implementation of the association storage adopts a multi-track packaging technology; taking MP4 or MKV container format as an example, the system writes compressed video data into a video track, and at the same time, serializes structured situational metadata and value evaluation metadata into binary or JSON format into a private data track or a subtitle track in the container; each metadata packet is marked with the same presentation timestamp as the corresponding video frame; this packaging mechanism ensures that when the system reads a certain segment of video, the environmental state, AI recognition result and value score at that time can be read millisecond-level synchronously without complex database queries across files; this “data-metadata” integrated storage structure also provides direct data support for subsequent value decay-based life cycle management, so that the system only needs to read the metadata track to determine whether to perform a deletion operation when scanning the storage file, without decoding video images, significantly reducing IO overhead.
[0071] Referring to the drawings Figure 2 With the drawings Figure 4 Step S600 mainly involves the full life cycle value management of the stored video data; this step discards the traditional first-in-first-out (FIFO) mechanical coverage strategy, and instead introduces a two-dimensional value decay mechanism based on time and content attributes, to realize intelligent release of storage space and long-term retention of key evidence; this step specifically includes event type recognition and decay coefficient matching, dynamic calculation of current retention value, and storage pruning operation based on threshold; the following will be described in detail in combination with steps S601 to S603.
[0072] In S601, event type analysis and decay coefficient matching of the storage object are performed; the periodic management module 107 periodically (for example, every morning or when the storage space occupancy rate reaches the warning line) traverses the video file index in the storage medium; for each video segment to be evaluated, the periodic management module 107 reads the associated structured situational metadata (packaged in step S500) thereof, and extracts the subject event category label output by the AI analysis module 103 at the recording time of the segment; the system internally presets a value decay strategy table, which establishes a mapping relationship between event categories and time decay coefficients ; the periodic management module 107 looks up and determines the decay coefficient corresponding to the video segment according to the extracted event category in the strategy table ; the decay coefficient The value is set to a very small value or zero, meaning that the value of this type of video decays very slowly over time and tends to be preserved permanently; for low-importance event categories (such as "normal traffic flow" or "static background"), Setting the value to a large value means that the value of this type of video will decrease rapidly over time.
[0073] In S602, the dynamic calculation of the current retention value of the video content is performed; the period management module 107 calculates the actual retention value of the video segment at the current moment using the exponential decay model based on the initial recording time of the video segment, the current system time, and the matched decay coefficient; this calculation process can quantify the value flow of video data in the time dimension.
[0074] Video clip at the current moment Retention value Follow the following dynamic decay model formula: ; In the formula, Indicates the current moment of the video clip. The retention value of a document is the sole criterion for determining whether or not to keep it. This indicates the time the video clip was recorded. The calculated real-time basic value weight (i.e., in step S400) The stored value represents the initial importance during video generation; Indicates the initial recording timestamp of the video clip; This indicates the duration the video clip has been stored, calculated using the current system time. Subtract recording time ; Indicates the first The dynamic decay coefficient corresponding to the type of event is determined by sub-step S601; is the base of the natural logarithm.
[0075] According to the model above, the current value of a video depends not only on the importance of its initial moment (as determined by...) The decision also depends on whether its content type can stand the test of time (by...). (Decision); This mechanism ensures that even if a video receives a high initial score due to a false alarm during recording, its value will rapidly decay to a low level if it belongs to a low-value category (such as a high score caused by a false alarm); conversely, true key evidence can remain above the retention threshold for a long time even if the initial score is not full, thanks to its extremely low decay rate.
[0076] In S603, storage space pruning and cleanup decisions are executed based on retention value; the periodic management module 107 calculates the current retention value. a preset system cleaning threshold are compared, and a corresponding lifecycle management action is performed according to a comparison result.
[0077] lifecycle management logic The execution rule is shown in the following formula: ; In the formula, indicates that a deletion operation is performed, and the period management module 107 sends an instruction to the file system to release the physical storage blocks occupied by the video segment and the associated video track data; indicates that a retention operation is performed, and the current storage state is maintained unchanged; indicates a system cleaning threshold, which can be a fixed value or a floating value dynamically adjusted according to the remaining space of the current storage medium (for example, the threshold is automatically raised to accelerate the elimination of low-value data when the remaining space is less).
[0078] As a supplement to the above-mentioned deletion operation, the deletion operation further includes a hierarchical cleaning mode: when is lower than the cleaning threshold but higher than the metadata retention threshold, the system only deletes the video track data occupying a large space, while retaining the structured context metadata and key frame index occupying a very small space; this hierarchical processing mode makes it possible for users to know the events and their context information that have occurred at that time through metadata retrieval even if the original video is cleaned, ensuring the traceability of historical data; the period management module 107 generates a log of each execution of the deletion or hierarchical cleaning operation (including the ID of the deleted video, the initial value, the survival time, etc.) as input data for the offline feedback calibration mechanism.
[0079] Referring to the accompanying drawings Figure 2 , step S700 mainly involves establishing a closed-loop learning mechanism based on real business feedback, and using the actual operation log generated during system operation to correct the internal parameters of the value evaluation model and the lifecycle management strategy in reverse; this step aims to solve the deviation problem that may exist between the preset parameters and the actual business needs, and realize the adaptive evolution of the system; this step specifically includes the quantitative definition of the real value of the business, the construction of the loss function, and the gradient update of the model parameters; steps S701 to S703 will be described in detail below.
[0080] In S701, the true value of the execution business is quantitatively extracted; the content value dynamic weight allocation and prediction module 105 periodically (e.g., weekly or monthly) reads historical records from the system's operation log database; the operation log records the user's interaction behavior with the stored video data, specifically including retrieval behavior, playback behavior and recovery behavior; based on these objective human operation data, the system reverse-engineers the true importance of video content in actual business.
[0081] To transform discrete user behaviors into scalars that can be used for computation, the system employs a weighted behavior analysis model to compute video segments. The true business value The calculation follows the formula below: ; In the formula, Indicates the first The real business value of a video clip is used as a label in supervised learning. This indicates a search indicator variable. The value is 1 when the video clip is successfully searched and clicked by a user using keywords or timeline; otherwise, it is 0. This indicates the actual duration for which the video clip was played; This indicates the total duration of the video clip. It represents the completion rate, reflecting the depth of user attention to the content; This variable indicates the recovery indicator. When the video clip was previously marked as low value by the system (about to be deleted or degraded) but is manually marked as "permanently retained" by the user or a "quality restoration" operation is performed, the value is 1; otherwise, it is 0. , , These represent the normalized weight coefficients for retrieval, viewing, and recovery behaviors, respectively. These coefficients are set according to the importance of different operations in the business scenario (for example, in an evidence collection scenario, the recovery operation usually means extremely high value, therefore...). (Settings are relatively large).
[0082] In S702, the loss function for assessing the bias is constructed; the system uses the predicted value (i.e., the real-time base value weight) calculated in step S400. (and the real business value obtained in step S701) By comparison, a loss function is constructed to measure the accuracy of the model; this loss function measures not only the accuracy of value assessment, but also the rationality of the decay strategy in lifecycle management.
[0083] Overall loss function of system construction As shown below: ; wherein the prediction model is expanded as: ; wherein, denotes the mean square error loss function of the current batch of samples; denotes the parameter set to be optimized, which includes the fusion weight coefficient in step S400 and the decay coefficient corresponding to different event categories in step S600; ; denotes the total number of historical video segment samples for calibration; denotes the predicted retention value of the i-th sample at the user operation moment calculated based on the current parameter set; denotes the historical AI perception score, POI activation score and heterogeneous data anomaly score (these data are read from the stored structured metadata) corresponding to the i-th sample, respectively; denotes the time interval from recording to user operation of the i-th sample; denotes the regularization term for constraining the parameter range (such as ensuring ), preventing overfitting. In S703, gradient update and self-optimization of model parameters are performed; the content value dynamic weight distribution and prediction module 105 adopts the stochastic gradient descent (SGD) or Adam optimization algorithm to iteratively update the parameter set according to the calculated loss function .
[0084] For the update logic of the fusion weight coefficient (taking as an example), the following applies: ; For the update logic of the decay coefficient (taking
[0085] as an example), the following applies: ; wherein, denotes the AI perception dimension fusion weight coefficient updated via the gradient descent algorithm, which will be applied to the real-time value assessment calculation in the next period; denotes the AI perception dimension fusion weight coefficient being used in the current period and to be updated; denotes the dynamic decay coefficient corresponding to the i-th event category after update; denotes the dynamic decay coefficient corresponding to the i-th event category being used in the current period; a first-order partial derivative of the total loss function with respect to the parameter a first-order partial derivative of the total loss function with respect to the parameter , which mathematically characterizes the sensitivity and rate direction of the loss function value change with respect to , and is used to indicate the adjustment gradient of the weight coefficient; a first-order partial derivative of the total loss function with respect to the parameter a first-order partial derivative of the total loss function with respect to the parameter , which is used to quantify the error gradient between the current decay strategy and the real business retention demand; a learning rate; through the above back propagation process, the system can automatically identify the root cause of the value evaluation deviation; for example, if a large number of users frequently retrieve and play video clips with low AI scores but high scores of heterogeneous sensors (such as sound or vibration), the loss function will guide the model to increase (heterogeneous weight) and decrease (AI weight), so as to more accurately capture such features in future real-time evaluation; similarly, if the videos of a certain type of event are often manually rescued (restored) by users before being automatically cleaned up, the system will automatically reduce the decay coefficient of this category , and extend its default retention period.
[0086] After completing the parameter update, the system deploys the new parameter set to the content value dynamic weight allocation and prediction module 105 and the period management module 107, completing a closed-loop self-evolution process; this mechanism ensures that the system can become more and more consistent with the actual business preferences of a specific place as time goes on, without the need for frequent manual adjustment of configuration parameters.
[0087] Although embodiments of the present application have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made therein without departing from the principles and spirit of the application, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for compressing and storing dynamic video, characterized in that, Includes the following steps: Based on multimodal contextual data streams and parameterized point of interest configuration information, the real-time basic value weight of the current video content is calculated using a value assessment model. Based on the calculated real-time basic value weight, the video data is compressed by adaptively selecting matching video encoding parameters, and the compressed video data is associated with and stored with multi-source structured metadata. Based on the storage duration of the stored video data and the corresponding dynamic value decay coefficient, the current retention value of the stored video data is updated periodically, and data pruning is performed when the calculated current retention value is lower than the preset system cleanup threshold. Historical business operation logs of the system are extracted as monitoring signals to perform offline feedback calibration of the weight parameters in the value assessment model.
2. The method for compressing and storing dynamic video according to claim 1, characterized in that, The steps for calculating the real-time basic value weight of the current video content include: The AI perception score, which represents the importance of visual content, the POI activation score, which represents the triggering state of business rules, and the heterogeneous data anomaly score, which represents the degree to which the environmental state deviates from the normal benchmark, are weighted and fused to obtain the real-time basic value weight.
3. The method for compressing and storing dynamic video according to claim 2, characterized in that, The steps for calculating the AI perception score include: Perform object detection on video frames and obtain the object's category label and detection confidence score; Based on the obtained category labels, the importance coefficient is obtained by querying the preset category priority mapping table; The obtained detection confidence score is multiplied by the obtained importance coefficient, and the product results of all targets in the frame are accumulated to obtain the AI perception score.
4. The method for compressing and storing dynamic video according to claim 2, characterized in that, The steps for calculating the POI activation score include: Retrieve user-defined point of interest (POI) configurations, which include spatial range, basic business priority weights, and parameterized rule sets. The parameterized rule set includes a target velocity threshold, a target duration threshold within the area, a target movement direction restriction, or a filtering condition for a specific object category. Based on the multimodal context data stream, it is determined whether the logical conditions in the parameterized rule set are met, and the POI activation score is calculated by combining the basic business priority weight.
5. The method for compressing and storing dynamic video according to claim 1, characterized in that, The step of adaptively selecting matching video encoding parameters includes: The real-time basic value weight is compared with the preset high-value judgment threshold and low-value judgment threshold; Based on the comparison results, a corresponding video encoding parameter is selected from a multi-level encoding configuration scheme consisting of high-definition mode bitrate, standard-definition mode bitrate, and low-stream mode bitrate.
6. The method for compressing and storing dynamic video according to claim 1, characterized in that, The step of periodically updating the current retention value of the stored video data includes: Extract main event category labels from the structured contextual metadata of the stored video data; Based on the extracted main event category labels, the dynamic value decay coefficient is searched and determined in the preset value decay strategy table; The current retention value is calculated by using an exponential decay model, combining the initial real-time base value weight of the stored video data, the storage duration, and the determined dynamic value decay coefficient.
7. The method for compressing and storing dynamic video according to claim 1, characterized in that, The steps for offline feedback calibration of the weight parameters in the value assessment model include: Extract retrieval, playback, and recovery behaviors from historical business operation logs, and quantify the true business value of video clips. A loss function is constructed based on the deviation between the calculated actual business value and the retention value predicted by the value assessment model. Based on the constructed loss function, the gradient descent algorithm is used to iteratively update the fusion weight coefficient and dynamic value decay coefficient in the value assessment model.
8. The method for compressing and storing dynamic video according to claim 1, characterized in that, Before calculating the real-time base value weights, the method further includes: Parallel acquisition of real-time video streams and heterogeneous sensor data from the monitored area; The real-time video stream is analyzed in real time to obtain the target density within the field of view; When the acquired target density within the field of view exceeds the preset target density trigger threshold, the acquisition parameters of the heterogeneous sensor data are adjusted in reverse, switching the sampling frequency of the heterogeneous sensor from the basic sampling frequency to the upper limit of high-frequency sampling.
9. The method for compressing and storing dynamic video according to claim 1, characterized in that, The data cropping operation includes: A tiered cleanup mode is proposed, in which, when the current retention value is lower than the system cleanup threshold but higher than the metadata retention threshold, only video track data is deleted, while structured contextual metadata and keyframe indexes are retained.
10. A dynamic video compression and storage system, applied to the method described in any one of claims 1-9, characterized in that, include: The content value dynamic weight allocation and prediction module is configured to calculate the real-time basic value weight of the current video content based on the multimodal context data stream and parameterized interest point configuration information, and extract the system's historical business operation logs as supervision signals to perform offline feedback calibration of the weight parameters in the value assessment model. The compression storage module is configured to adaptively select matching video encoding parameters to compress video data based on the calculated real-time basic value weight, and to associate and store the compressed video data with multi-source structured metadata. The periodic management module is configured to periodically update the current retention value of stored video data based on the storage duration of the stored video data and the corresponding dynamic value decay coefficient, and to perform data pruning operation when the calculated current retention value is lower than the preset system cleanup threshold.