Video monitoring data storage management system based on big data

By constructing cross-layer feature fusion calculation of multi-source heterogeneous data acquisition and three-dimensional tensor structure, real-time joint perception and dynamic strategy optimization of encoding, transmission and storage in video surveillance system is realized, the response lag problem caused by data silos is solved, and the system's adaptability and decision-making integrity are improved.

CN120263944AActive Publication Date: 2025-07-04SHANDONG HENENG TECH CO LTD

Patent Information

Application Number
CN202510461159.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-04
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

In the existing video surveillance system, the encoding, transmission and storage modules are optimized separately, resulting in data silos, unable to achieve global resource coordination, resulting in response lag and suboptimal decision-making, and it is difficult to meet the requirements of high compression rate, low transmission jitter and high storage reliability.

Method used

Through multi-source heterogeneous data acquisition, cross-layer feature fusion calculation, comprehensive index fusion and adaptive regulation and execution, a three-dimensional tensor structure is built to realize real-time joint perception and dynamic strategy optimization of encoding, transmission and storage, and form a closed-loop control mechanism.

Benefits of technology

Real-time joint perception of encoding characteristics, network state and storage load is realized, dynamically matches the best management strategy, improves the system's adaptability to the dynamic environment, solves the problems of local optimization limitations and response lag in traditional methods, and ensures the system's decision integrity and optimization effect in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120263944A_ABST
    Figure CN120263944A_ABST
Patent Text Reader

Abstract

The invention discloses a video monitoring data storage management system based on big data, and particularly relates to the field of data analysis, comprising multi-source heterogeneous data acquisition, cross-layer feature fusion calculation, comprehensive index fusion, collaborative decision generation and adaptive regulation and control execution. Through multi-level data synchronous acquisition and cross-domain feature fusion, a three-dimensional tensor structure is constructed to eliminate dimensional difference, global state perception of coding, transmission and storage is realized, local optimization limitation caused by a traditional data island is solved, a dynamic weight adjustment model automatically matches an optimal strategy based on a three-order collaborative management coefficient calculated in real time, and a dynamic weight adjustment model is improved. The joint elastic regulation and control of coding parameters and network protocols and the dynamic reconstruction of storage resources are realized, and the defect of response lag of a static threshold mechanism under burst traffic is overcome.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data analysis technology, and more specifically, to a video surveillance data storage management system based on big data. Background Art

[0002] With the rapid development of ultra-high-definition video, low-latency live broadcast, smart security and other businesses, the system needs to meet the stringent requirements of high compression rate, low transmission jitter and high storage reliability. Current industry solutions generally adopt an independent optimization mode, and each module interacts with data through a simple interface, lacking a resource coordination mechanism from a global perspective.

[0003] Mainstream technical architectures usually configure the encoder to a fixed bit rate control mode. The network layer relies on traditional congestion control algorithms for bandwidth adaptation. The storage layer performs data segmentation and hot and cold stratification based on preset rules. This serial processing mode causes the encoding efficiency to be out of touch with the network status. The transmission strategy cannot perceive the topological distribution of storage nodes, and the operation and maintenance system can only trigger single-point policy adjustments through threshold alarms.

[0004] The existing scheme has three significant defects: the fixed encoding parameters cause quality fluctuations under burst traffic, the mismatch between the transport layer retransmission mechanism and the storage node load status causes data scheduling conflicts, the offline decision model cannot respond to dynamic changes in the system in real time, and the isolated optimization of each module forms a local optimal trap, which makes it difficult to achieve the best allocation of cross-domain resources, seriously restricting the comprehensive energy efficiency improvement of large-scale video systems. Summary of the invention

[0005] In order to overcome the above-mentioned defects of the prior art, the present invention provides a video surveillance data storage and management system based on big data, which solves the problems of delayed response caused by data islands in the encoding, transmission and storage modules proposed in the above-mentioned background technology, the inability of static strategies to adapt to dynamic environmental changes, and suboptimal decision-making caused by the lack of a closed-loop feedback mechanism through the following scheme.

[0006] To achieve the above object, the present invention provides the following technical solution: a video surveillance data storage and management system based on big data, comprising:

[0007] Multi-source heterogeneous data acquisition module: By deploying acquisition tools of the video coding layer, network transmission layer and distributed storage layer in parallel, the original data can be captured synchronously. The video coding layer uses video stream analysis tools to extract frame structure and motion parameters. The network transmission layer captures protocol features and channel status through programmable network devices. The distributed storage layer obtains physical media status and access logs based on the storage system interface.

[0008] Cross-layer Feature Fusion Calculation Module: Perform streaming feature extraction on the raw data provided by the acquisition layer, and calculate three types of metrics for the video coding efficiency monitoring group, the video transmission feature analysis group, and the storage feature analysis group respectively;

[0009] Comprehensive Index Fusion Module: Perform non-linear aggregation on the feature matrix output by the feature calculation layer, and use three groups of formulas with calibrated parameters to generate the video coding efficiency comprehensive index, the transmission efficiency aggregation metric, and the storage system entropy change index respectively;

[0010] Collaborative Decision Generation Module: Based on the three-dimensional state vector of the index analysis layer, calculate the global management coefficient through a third-order collaborative management decision model, and generate hierarchical management instructions in combination with the preset threshold conditions and state space division rules;

[0011] Adaptive Regulation Execution Module: Convert the policy instructions generated by the decision-making layer into three core operations: dynamic optimization of video coding parameters, real-time switching of network transmission protocols, and reconstruction of storage resource distribution strategies. Drive the update of codec parameters, the hot loading of network protocol stack configuration, and the execution of storage node data migration tasks through a distributed control system.

[0012] Preferably, the acquisition layer uses the storage system monitoring interface to collect physical storage status information. All data is standardized and stored in the time series database to provide basic data input for subsequent feature calculations, and its output is directly associated with the data processing module of the feature calculation layer.

[0013] Preferably, the multi-source heterogeneous data acquisition module uses the probe function of FFmpeg to parse the video stream metadata, extracts the type and size of each frame through the ffprobe - show_frames command, cooperates with the --analysis - save parameter of the x265 encoder to output macroblock-level entropy values and motion vector residual data, uses the VMAF toolbox to batch encode and record the PSNR value under the preset bitrate ladder, obtains the key frame density by statistically analyzing the I-frame interval time in the GOP structure, and directly parses the resolution information from the video container header metadata. Finally, aggregate the I-frame byte sequence, P-frame size sequence, macroblock entropy value matrix, motion residual time series data, multi-bitrate PSNR dataset, and resolution parameters through a Python script and store them in the time series database.

[0014] Preferably, the multi-source heterogeneous data acquisition module deploys P4 programmable switch mirror traffic, uses TShark to capture packets and sets -d udp.port==1234, uses an rtp filter to parse the RTP / RTCP protocol, counts the packet loss rate and retransmission times of spatio-temporal shards through the rtp_analysis module, samples and calculates the network jitter coefficient using the / proc / net / tcpprobe interface of the Linux tcpprobe module, obtains the protocol header overhead by disassembling the TCP / IPv6 / RTP header length field, records the channel bandwidth fluctuation using the JSON output mode of iperf3, extracts the video bitrate change characteristics from the SEI information output by the encoder, obtains the FEC redundancy by parsing the ULP FEC parameter of the RTP extension header, extracts the hierarchical coding depth from the DONC unit of the SVC encoded stream, measures the RTT data through the ACK reception interval of the QUIC protocol, and injects all data into the InfluxDB time series library in real time through the Kafka pipeline.

[0015] Preferably, the multi-source heterogeneous data acquisition module enables the monitoring of bluestore in the Ceph cluster, exports the object access heat map and physical location mapping table through the ceph-dencoder tool, converts and obtains the SSD erase count from the Percentage_Used field of the SMART log, uses the IPMI sensor-util command to collect the storage temperature in real time, analyzes the compressed block size distribution from the --trace log of the zstd compressor, reads the storage page size through the blockdev --getbsz command, calculates the cross-node data similarity by combining the detection results of Ceph's rados list-inconsistent-pg with the MinHash algorithm, and statistically calculates the data block temperature entropy value through the access log sliding window. Finally, all original metrics are aggregated by the Prometheus exporter and stored in the OpenTSDB cluster.

[0016] Preferably, the calculation results of the feature calculation layer form a structured feature matrix and are transmitted to the index analysis layer, constituting the data basis for subsequent comprehensive evaluation, and its processing logic directly depends on the data completeness of the acquisition layer.

[0017] Preferably, the video coding efficiency monitoring group includes spatio-temporal compression ratio, entropy coding fluctuation coefficient, motion prediction residual gradient, and bitrate-quality inflection point slope; the video transmission feature analysis group includes spatio-temporal scheduling correlation, protocol stack efficiency index, channel feature matching degree, and fault tolerance diffusion coefficient; the storage feature analysis group includes hot and cold data vorticity, coding storage coupling rate, medium loss balance factor, and topology-aware compression rate.

[0018] Preferably, the spatio-temporal compression ratio calculates the average value of P frames using a sliding window and analyzes the distribution of key frames in combination with the GoP structure, which is specifically expressed as: STR represents the spatio-temporal compression ratio, I^(v) represents the byte size of a single I frame, \bar{P}^(v) represents the average byte size of P frames in the last 10 seconds, and p_t^(v) represents the key frame density per unit time; the entropy coding fluctuation coefficient is calculated macro-block by macro-block based on the Shannon entropy formula and outliers are filtered using the quantile method, which is specifically expressed as: EVF represents the entropy coding fluctuation coefficient, H_max^(mb) represents the maximum macro-block entropy value within a single frame, H_min^(mb) represents the minimum macro-block entropy value within a single frame, and \bar{H}^(mb) represents the average value of the macro-block entropy values within a frame; the motion prediction residual gradient calculates the difference through the motion compensation residual matrix and accelerates the summation using SIMD instructions, which is specifically expressed as: MRG represents the motion prediction residual gradient, ΔR_i^(t)=R_i^(t)-R_i^(t-1) represents the change in the residual of the i-th macro-block, A represents the frame area, and N represents the total number of macro-blocks; the slope of the rate-quality inflection point constructs a cubic spline interpolation function for derivation and the resolution term is standardized, which is specifically expressed as: RQS represents the slope of the rate-quality inflection point, represents the derivative of the PSNR-rate curve, and W and h represent the width and height of the video resolution.

[0019] Preferably, the spatio-temporal scheduling correlation evaluates the network jitter component based on the Kalman filter, which is specifically expressed as: STC represents the spatio-temporal scheduling correlation, λ_s represents the spatial slice packet loss rate, τ_t represents the number of time slice retransmissions, J represents the network jitter coefficient, and ε represents the anti-zero constant; the protocol stack efficiency index extracts the header fields using a protocol parser and calculates the MTU matching degree, which is specifically expressed as: PSE represents the protocol stack efficiency index, η_p represents the payload ratio, H_T represents the TCP header overhead, H_R represents the RTP header overhead, and α_f represents the fragmentation aggregation degree factor; the channel feature matching degree calculates the Shannon entropy using a sliding window, which is specifically expressed as: CCM=1-|E B -E V |, CCM represents the channel feature matching degree, E_B represents the channel bandwidth distribution entropy, and E_V represents the video bitrate change entropy; the fault tolerance diffusion coefficient analyzes the redundant structure through the FEC matrix and traverses the hierarchical coding tree, which is specifically expressed as: FDC represents the fault tolerance diffusion coefficient, β_F represents the forward error correction redundancy, d_L represents the hierarchical coding depth, and RTT represents the network round-trip delay.

[0020] Preferably, the vorticity of hot and cold data constructs an access frequency field to calculate the gradient, and spherical coordinates are used for positioning, specifically expressed as: HDV represents the vorticity of hot and cold data, f represents the access frequency gradient, σ_L represents the storage location dispersion, and H_T represents the data block temperature entropy; the encoding storage coupling rate uses a memory-mapped file to detect the boundary alignment situation, specifically expressed as: SCR represents the encoding storage coupling rate, N_align represents the number of block boundary alignments, S_c represents the compressed block size, S_p represents the storage page size, and N_total represents the total number of data blocks; the medium loss equalization factor establishes a wear leveling state matrix and combines the temperature sensor data, specifically expressed as: WEF represents the medium loss equalization factor, W_max / min represents the maximum / minimum number of erase / write cycles, \bar{W} represents the average number of erase / write cycles, and δ_T represents the temperature difference coefficient; the topology-aware compression rate uses locality-sensitive hashing to detect cross-node similarity, specifically expressed as: TCR represents the topology-aware compression rate, C_std represents the standard compression rate, D_kj represents the data similarity between nodes k and j, and n represents the total number of nodes.

[0021] Preferably, the cross-layer feature fusion calculation module constructs a cross-layer joint feature matrix by defining a three-dimensional tensor structure: taking the time axis as the first dimension, dividing fixed-length sliding windows; in the feature dimension, mapping STR, EVF, MRG, and RQS to tensor channels 1-4, STC, PSE, CCM, and FDC to channels 5-8, and HDV, SCR, WEF, and TCR to channels 9-12; using a double-buffer mechanism to achieve timing alignment, triggering the splicing operation when and only when the window timestamps of the three data streams overlap by more than 95%, and optimizing the memory layout using the SIMD instruction set, finally generating a tensor with the shape of [T×12×N], where T is the number of time windows and N is the batch size, and normalizing in the channel dimension through LayerNorm to ensure the dimensional consistency of cross-layer features.

[0022] Preferably, the three comprehensive indicators generated by the index analysis layer constitute a three-dimensional characterization vector of the system state, and their calculation results are transmitted to the decision-making layer through a data pipeline.

[0023] Preferably, the video coding efficiency comprehensive index is specifically expressed as: VECI represents the video coding efficiency comprehensive index, α represents the encoder architecture adaptation constant, and β represents the time-domain continuity compensation coefficient.

[0024] Preferably, the transmission efficiency aggregation metric is specifically expressed as: TEM represents the transmission efficiency aggregation metric, γ represents the network topology complexity factor, and δ represents the fault tolerance efficiency critical threshold.

[0025] Preferably, the storage system entropy change index is specifically expressed as: SSEI represents the storage system entropy change index, η represents the storage medium aging rate, θ represents the optimal coupling threshold, and σ represents the process fluctuation standard deviation.

[0026] Preferably, the three - order collaborative management decision - making model is specifically expressed as:

[0027] TCDM represents the global management coefficient, K represents the coding quality weight, μ represents the transmission stability coefficient, v represents the storage pressure damping factor, ω represents the mutation tolerance threshold, VECIΔ represents the coding efficiency change rate in the recent 3 minutes, and TEMΔ represents the transmission efficiency gradient in the recent 5 RTT windows.

[0028] Preferably, the collaborative decision - making generation module divides the system state into three response levels through a hierarchical management trigger mechanism based on the TCDM value by setting two critical thresholds of 80 and 50: when the TCDM value is higher than 80, the system enables the joint optimization of the elastic coding mechanism and the BBRv3 congestion control algorithm, automatically switches to the constant quality factor coding mode on the premise of maintaining video quality, and simultaneously optimizes network queue management to reduce transmission delay; if the TCDM value is between 50 and 80, the time - domain hierarchical coding technology is activated to achieve bit - rate adaptation, and high - popularity storage nodes are pre - heated in advance to balance the access load; when the TCDM value drops below the 50 threshold, a forced I - frame request is immediately executed to quickly reconstruct the integrity of the video stream, and cross - availability - zone data migration is started to ensure storage availability.

[0029] Preferably, the system state changes generated by the operation execution of the regulation and execution layer are real - time fed back to the acquisition layer through the monitoring loop to form a closed - loop control.

[0030] Preferably, the regulation and execution layer implements three - dimensional control based on the hierarchical management instruction set generated by the decision - making layer, including the coding side, the transmission side, and the storage side. The coding side realizes real - time optimization by dynamically reconstructing the GOP structure of the picture group, uses the exponential function model to automatically calculate the GOP length according to VECI, and automatically shortens the GOP when the VECI value decreases to enhance fault tolerance. Its mathematical expression is The GOP_size represents the Group of Pictures size, ensuring a dynamic balance between coding efficiency and error recovery ability; on the transmission side, an error correction level FEC hierarchical control mechanism is established, and the error correction level is determined through a composite function of TEM and RTT. When the value of TEM / √RTT breaks through 4.7, it automatically upgrades to level 5 FEC to enhance the packet loss resistance ability, and when it is lower than 2.9, it downgrades to the basic error correction mode to save bandwidth; on the storage side, a garbage collection GC strategy with spatio-temporal awareness is implemented, and the GC trigger threshold is dynamically adjusted using a sine periodic function combined with the system running time t. The expression is GC_threshold = 0.65 + 0.12·sin(2πt / 86400), where GC_threshold represents the garbage collection threshold, enabling the storage reorganization operation to enhance activity during the low-access period and automatically suppressing the GC frequency during the peak period to ensure I0 performance. All control operations take effect synchronously through the distributed control bus, and the execution status data is returned to the acquisition layer in real-time through the buried point system, forming a complete closed-loop from decision-making to execution.

[0031] Technical effects and advantages of the present invention:

[0032] 1. Through the construction of a multi-level data fusion mechanism, the present solution realizes real-time joint perception of coding features, network status, and storage load. The three-dimensional tensor structure is used to eliminate the cross-domain feature dimension difference, and the non-linear aggregation algorithm is used to accurately represent the global state of the system, solving the local optimization limitation caused by data islands in traditional methods and significantly improving the decision-making integrity in complex scenarios;

[0033] 2. The present invention introduces a dynamic weight adjustment model, which automatically matches the best management strategy according to the real-time calculated third-order collaborative management coefficient. The coding parameters and network protocols are jointly and flexibly regulated, and the storage strategy can be dynamically reconstructed according to the transmission load, overcoming the response lag defect of the traditional static threshold mechanism under burst traffic and effectively enhancing the system's adaptability to the dynamic environment;

[0034] 3. A full closed-loop control loop from data acquisition to strategy execution is established. The decision-making instructions are converted into operable control signals through the distributed execution module, and the system state changes are fed back to the feature calculation layer in real-time to trigger parameter calibration, forming a continuous optimization mechanism, breaking through the sub-optimal strategy solidification problem existing in the traditional offline decision-making mode, and realizing the continuous convergence and optimization of the system operation state. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 It is a schematic diagram of the overall structure of the present invention.

[0036] Figure 2 It is a schematic diagram of the complete embodiment structure of the present invention.

[0037] Figure 3 It is a schematic diagram of the swimlane diagram structure of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0038] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0039] Refer to Figures 1 - 3 A video surveillance data storage and management system based on big data shown in the figure, including:

[0040] Multi-source heterogeneous data acquisition module: By deploying acquisition tools for the video encoding layer, network transmission layer, and distributed storage layer in parallel, the synchronous capture of raw data is realized. The video encoding layer uses a video stream analysis tool to extract the frame structure and motion parameters. The network transmission layer captures protocol features and channel status through programmable network devices. The distributed storage layer obtains the physical medium status and access logs based on the storage system interface.

[0041] The acquisition layer uses the storage system monitoring interface to collect physical storage status information. All data is standardized and stored in the time series database, providing basic data input for subsequent feature calculations. Its output is directly associated with the data processing module of the feature calculation layer.

[0042] The multi-source heterogeneous data acquisition module uses the probe function of FFmpeg to parse video stream metadata, extracts the type and size of each frame through the ffprobe - show_frames command, cooperates with the --analysis - save parameter of the x265 encoder to output macroblock - level entropy values and motion vector residual data, uses the VMAF tool kit to batch - encode and record PSNR values under a preset bitrate ladder, obtains the key frame density by statistically analyzing the I - frame interval time in the GOP structure, and directly parses the resolution information from the video container header metadata. Finally, the I - frame byte sequence, P - frame size sequence, macroblock entropy value matrix, motion residual time series data, multi - bitrate PSNR data set, and resolution parameters are aggregated through a Python script and stored in the time series database.

[0043] The multi-source heterogeneous data acquisition module deploys the mirror traffic of the P4 programmable switch, uses TShark to capture packets and sets -d udp.port==1234, parses the RTP / RTCP protocol with an rtp filter, calculates the packet loss rate and retransmission times of spatio-temporal shards through the rtp_analysis module, samples and calculates the network jitter coefficient using the / proc / net / tcpprobe interface of the Linux tcpprobe module, obtains the protocol header overhead by disassembling the TCP / IPv6 / RTP header length fields, records the channel bandwidth fluctuation using the JSON output mode of iperf3, extracts the video bitrate change characteristics from the SEI information output by the encoder, obtains the FEC redundancy by parsing the ULP FEC parameters of the RTP extension header, extracts the hierarchical coding depth from the DONC unit of the SVC coding stream, measures the RTT data through the ACK reception interval of the QUIC protocol, and all data is injected into the InfluxDB time series library in real time through the Kafka pipeline.

[0044] The multi-source heterogeneous data acquisition module enables the monitoring of bluestore in the Ceph cluster, exports the object access heat map and physical location mapping table through the ceph-dencoder tool, converts and obtains the SSD write count from the Percentage_Used field of the SMART log, uses the IPMI sensor-util command to collect the storage temperature in real time, analyzes the compression block size distribution from the --trace log of the zstd compressor, reads the storage page size through the blockdev --getbsz command, calculates the cross-node data similarity by combining the detection results of Ceph's rados list-inconsistent-pg with the MinHash algorithm, and statistically calculates the data block temperature entropy value through the access log sliding window. Finally, all the original metrics are aggregated by the Prometheus exporter and stored in the OpenTSDB cluster.

[0045] Cross-layer feature fusion calculation module: performs streaming feature extraction on the raw data provided by the acquisition layer, and calculates three types of metrics for the video coding efficiency monitoring group, the video transmission feature analysis group, and the storage feature analysis group respectively.

[0046] The calculation results of the feature calculation layer form a structured feature matrix and are transmitted to the metric analysis layer, constituting the data basis for subsequent comprehensive evaluation. Its processing logic directly depends on the data completeness of the acquisition layer.

[0047] The video encoding efficiency monitoring group includes the spatio-temporal compression ratio, the entropy coding fluctuation coefficient, the motion prediction residual gradient, and the bitrate-quality inflection point slope; the video transmission feature analysis group includes the spatio-temporal scheduling correlation, the protocol stack efficiency index, the channel feature matching degree, and the fault tolerance diffusion coefficient; the storage feature analysis group includes the hot and cold data vorticity, the encoding storage coupling rate, the medium loss balance factor, and the topology-aware compression rate.

[0048] The spatio-temporal compression ratio calculates the average value of P frames using a sliding window and analyzes the key frame distribution in combination with the GoP structure. Specifically, it is expressed as: STR represents the spatio-temporal compression ratio, I^(v) represents the byte size of a single I frame, \bar{P}^(v) represents the average byte size of P frames in the last 10 seconds, and ρ_t^(v) represents the key frame density per unit time; the entropy coding fluctuation coefficient is calculated macro-block by macro-block based on the Shannon entropy formula, and outliers are filtered using the quantile method. Specifically, it is expressed as: EVF represents the entropy coding fluctuation coefficient, H_max^(mb) represents the maximum macro-block entropy value within a single frame, H_min^(mb) represents the minimum macro-block entropy value within a single frame, and \bar{H}^(mb) represents the average value of the macro-block entropy values within the frame; the motion prediction residual gradient calculates the difference through the motion compensation residual matrix and uses SIMD instructions to accelerate the summation. Specifically, it is expressed as: MRG represents the motion prediction residual gradient, ΔR_i^(t) = R_i^(t) - R_i^(t - 1) represents the change in the residual of the i-th macro-block, A represents the frame area, and N represents the total number of macro-blocks; the bitrate-quality inflection point slope constructs a cubic spline interpolation function for derivation, and the resolution term is standardized. Specifically, it is expressed as: RQS represents the bitrate-quality inflection point slope, represents the derivative of the PSNR-bitrate curve, and W and h represent the width and height of the video resolution.

[0049] The spatio-temporal scheduling correlation evaluates the network jitter component based on the Kalman filter. Specifically, it is expressed as: STC represents the spatio-temporal scheduling correlation, λ_s represents the spatial slice packet loss rate, τ_t represents the number of time slice retransmissions, J represents the network jitter coefficient, and ε represents the anti-zero constant; the protocol stack efficiency index uses a protocol parser to extract the header fields and calculates the MTU matching degree. Specifically, it is expressed as: PSE represents the protocol stack efficiency index, n_p represents the payload ratio, H_T represents the TCP header overhead, H_R represents the RTP header overhead, and α_f represents the fragmentation aggregation degree factor; the channel feature matching degree calculates the Shannon entropy using a sliding window. Specifically, it is expressed as: CCM = 1 - |E B -E V|, CCM represents the channel feature matching degree, E_B represents the channel bandwidth distribution entropy, and E_V represents the video bitrate change entropy; the fault-tolerant diffusion coefficient analyzes the redundant structure through the FEC matrix and traverses the hierarchical coding tree, and is specifically expressed as: FDC represents the fault-tolerant diffusion coefficient, β_F represents the forward error correction redundancy, d_L represents the hierarchical coding depth, and RTT represents the network round-trip delay.

[0050] The hot and cold data vorticity constructs an access frequency field to calculate the gradient and uses spherical coordinates for positioning, and is specifically expressed as: HDV represents the hot and cold data vorticity, represents the access frequency gradient, σ_L represents the storage location dispersion, and H_T represents the data block temperature entropy; the coding storage coupling rate uses a memory-mapped file to detect the boundary alignment situation, and is specifically expressed as: SCR represents the coding storage coupling rate, N_align represents the number of block boundary alignments, Sc represents the compressed block size, S_p represents the storage page size, and N_total represents the total number of data blocks; the medium loss equalization factor establishes a wear leveling state matrix and combines temperature sensor data, and is specifically expressed as: WEF represents the medium loss equalization factor, W_max / min represents the maximum / minimum number of erase / write cycles, \bar{W} represents the average number of erase / write cycles, and δ_T represents the temperature difference coefficient; the topology-aware compression rate uses locality-sensitive hashing to detect cross-node similarity, and is specifically expressed as: TCR represents the topology-aware compression rate, C_std represents the standard compression rate, D_kj represents the data similarity between nodes k and j, and n represents the total number of nodes.

[0051] The cross-layer feature fusion calculation module constructs a cross-layer joint feature matrix by defining a three-dimensional tensor structure: taking the time axis as the first dimension, dividing it into fixed-length sliding windows; in the feature dimension, mapping STR, EVF, MRG, and RQS to tensor channels 1-4, STC, PSE, CCM, and FDC to channels 5-8, and HDV, SCR, WEF, and TCR to channels 9-12; using a double-buffer mechanism to achieve timing alignment, and triggering the splicing operation when and only when the window timestamps of the three data streams overlap by more than 95%, optimizing the memory layout using the SIMD instruction set, and finally generating a tensor with a shape of [T×12×N], where T is the number of time windows and N is the batch size, and normalizing in the channel dimension through LayerNorm to ensure the dimensional consistency of cross-layer features.

[0052] Comprehensive index fusion module: Nonlinearly aggregates the feature matrix output by the feature calculation layer, and uses three groups of formulas with calibrated parameters to generate the comprehensive index of video coding efficiency, the aggregated metric of transmission efficiency, and the entropy change index of the storage system respectively.

[0053] The three comprehensive indicators generated by the indicator analysis layer constitute a three-dimensional representation vector of the system state, and their calculation results are transmitted to the decision-making layer through the data pipeline.

[0054] The comprehensive video coding efficiency index is specifically expressed as: VECI represents the comprehensive video coding efficiency index, α represents the encoder architecture adaptation constant, and β represents the temporal continuity compensation coefficient.

[0055] The comprehensive video coding efficiency index constructs a basic efficiency factor with the product of STR and RQS, reflecting the synergistic effect of compression ability and quality sensitivity. It applies a 1.5th power amplification to RQS to affect non-linear quality degradation. The geometric mean of EVF and MRG is introduced into the denominator to suppress the evaluation deviation caused by high-frequency texture and motion mutations. The logarithmic term is used to compensate for the time continuity loss in low key-frame density scenarios. The β coefficient is determined through regression analysis of the H.264 / HEVC encoder on the CMCD standard test sequence set. The α value is derived from the principal component analysis of the performance of different encoders on UGC content, and the normalized value of the first principal component loading is taken.

[0056] The transmission efficiency aggregation metric is specifically expressed as: TEM represents the transmission efficiency aggregation metric, γ represents the network topology complexity factor, and δ represents the fault tolerance efficiency critical threshold.

[0057] The transmission efficiency aggregation metric is based on the orthogonal characteristics of protocol efficiency and channel matching, constructs a coupling term of PSE and CCM. The 0.7th power of CCM weakens the marginal benefit of perfect matching. STC is used as the denominator to reflect the scheduling conflict loss. The exponential term introduces the influence factor (1 - FDC / δ) of the fault tolerance mechanism to achieve adaptive weight adjustment. The δ value is derived from the FDC boundary condition corresponding to 99.999% reliability in the 5G NR URLLC standard. The arctan function maps the ratio of the number of retransmissions to the packet loss rate to the interval [0, π / 2) to suppress the metric distortion under extreme network conditions. The γ coefficient is obtained by fitting the simulation data of different cell densities according to the 3GPP 38.901 channel model and is corrected using the Weibul1 distribution shape parameter.

[0058] The storage system entropy change index is specifically expressed as: SSEI represents the storage system entropy change index, η represents the storage medium aging rate, θ represents the optimal coupling threshold, and σ represents the process fluctuation standard deviation.

[0059] The entropy change index of the storage system highlights the impact of access pattern mutations with the 2.2 power of HDV, forms a wear-leveling constraint with the WEF, the logarithmic term reflects the diminishing marginal benefit of compression ratio improvement, the error function erf() converts the degree of SCR deviation from the ideal threshold θ into a penalty coefficient, the value of θ is taken from the best write alignment experimental data of 3D NAND flash at the 24nm / 16nm process, σ is determined by the distribution of the Page Error rate in the SSD durability test, the η coefficient integrates multiple failure models such as the NAND cell erase / write cycle, data retention time, and temperature accelerated aging, and is calibrated by the Arrhenius equation for the 85°C accelerated aging experiment.

[0060] Collaborative decision-making generation module: Based on the three-dimensional state vector of the index analysis layer, calculate the global management coefficient through the third-order collaborative management decision-making model, and generate hierarchical management instructions in combination with the preset threshold conditions and state space partitioning rules.

[0061] The decision-making layer converts the system state quantity into a specific operation strategy, and its output instruction set directly drives the operation module of the regulation execution layer.

[0062] The third-order collaborative management decision-making model is specifically expressed as: TCDM represents the global management coefficient, K represents the coding quality weight, μ represents the transmission stability coefficient, v represents the storage pressure damping factor, ω represents the mutation tolerance threshold, VECIΔ represents the change rate of coding efficiency in the past 3 minutes, and TEMΔ represents the gradient of transmission efficiency in the past 5 RTT windows.

[0063] The construction of the third-order collaborative management decision-making model starts with the quantitative modeling of the non-linear interaction relationship among VECI, TEM, and SSEI. First, establish the basic decision-making factors to characterize the balance relationship between the system's positive efficiency and storage constraints. Among them, the exponential terms κ = 0.6, μ = 1.2, and v = 0.9 are determined by calculating the marginal contribution of each index to QoE on the video session dataset through the Shapley value decomposition method to ensure that the coding quality weight is lower than the transmission stability requirement; introduce ∈ = 10^-5 to prevent the denominator from approaching zero abnormally in the low storage pressure scenario, and the mutation suppression term is constructed based on the Lyapunov stability theory, where ω = 15 is determined by analyzing the 90th percentile of index mutations in the large-scale video cloud operation and maintenance data. When the index change rate exceeds 15 units / time window, the output of the tanh function approaches 1, causing the overall decision value to decay and suppressing the policy oscillation caused by sudden loads.

[0064] The collaborative decision-making generation module divides the system state into three response levels based on the hierarchical management trigger mechanism of the TCDM value by setting two critical thresholds of 80 and 50: when the TCDM value is higher than 80, the system enables the joint optimization of the elastic coding mechanism and the BBRv3 congestion control algorithm, automatically switches to the constant quality factor coding mode while maintaining the video quality, and at the same time optimizes the network queue management to reduce the transmission delay; if the TCDM value is between 50 and 80, the time-domain hierarchical coding technology is activated to achieve bitrate adaptability, and the high-heat storage nodes are preheated in advance to balance the access load; when the TCDM value drops below the 50 threshold, a forced I-frame request is immediately executed to quickly reconstruct the integrity of the video stream, and cross-availability zone data migration is started to ensure storage availability.

[0065] Adaptive regulation execution module: It converts the policy instructions generated by the decision-making layer into three core operations, namely dynamic optimization of video coding parameters, real-time switching of network transmission protocols, and reconstruction of storage resource distribution strategies, and drives the update of codec parameters, hot loading of network protocol stack configuration, and execution of storage node data migration tasks through a distributed control system.

[0066] The system state changes generated by the operation execution of the regulation execution layer are fed back to the acquisition layer in real time through the monitoring loop to form a closed-loop control.

[0067] The regulation execution layer implements three-dimensional control based on the hierarchical management instruction set generated by the decision-making layer, including the coding side, the transmission side, and the storage side. On the coding side, real-time optimization is achieved by dynamically reconstructing the Group of Pictures (GOP) structure. The GOP length is automatically calculated according to the VECI using an exponential function model. When the VECI value decreases, the GOP is automatically shortened to enhance fault tolerance. Its mathematical expression is GOP_size represents the size of the Group of Pictures, ensuring the dynamic balance between coding efficiency and error recovery ability; on the transmission side, an error correction level Forward Error Correction (FEC) hierarchical control mechanism is established, and the error correction level is determined by the composite function of TEM and RTT. When the value of TEM / √RTT breaks through 4.7, it automatically upgrades to level 5 FEC to enhance the packet loss resistance ability, and when it is lower than 2.9, it degrades to the basic error correction mode to save bandwidth; on the storage side, a garbage collection (GC) strategy with space-time awareness is implemented, and the GC trigger threshold is dynamically adjusted using a sine periodic function combined with the system running time t. The expression is GC_threshold = 0.65 + 0.12·sin(2πt / 86400), where GC_threshold represents the garbage collection threshold, enabling the storage reorganization operation to enhance activity during the access low period and automatically suppressing the GC frequency during the peak period to ensure the I / O performance. All regulation operations take effect synchronously through the distributed control bus, and the execution status data flows back to the acquisition layer in real time through the buried point system, forming a complete closed-loop from decision-making to execution.

[0068] What needs to be specifically explained in this embodiment is that the constants in all the formulas in this embodiment are determined by the two-stage "theoretical modeling + experimental calibration": first, a differential equation or a probability model is established to describe the system dynamics (such as the coding rate distortion equation, the network fluid model, the storage wear equation), and the theoretical constraint range of the parameters is derived; then an orthogonal experimental matrix is ​​designed, and 2000+ sample points are generated using the actual system (FFmpeg encoder, P4 network simulation platform, Ceph cluster), and the Bayesian optimization algorithm is used to find the parameter value that maximizes the Pearson correlation coefficient between the formula output and the actual QoE score.

[0069] The present invention firstly realizes the synchronous capture of multi-source heterogeneous data by deploying acquisition tools of the video coding layer, the network transmission layer and the distributed storage layer in parallel, uses FFmpeg to parse the video stream metadata, adopts the P4 programmable switch mirror traffic to obtain the network transmission characteristics, and enables the storage monitoring function in the Ceph cluster to collect the physical medium status. All the original data are stored in the time series database after standardization, and then enters the cross-layer feature fusion stage. Through streaming calculation, 12 indicators of the video coding efficiency monitoring group, the transmission feature analysis group and the storage feature analysis group are extracted, and a three-dimensional tensor structure is constructed to realize the time window alignment and feature dimension normalization. In the comprehensive indicator fusion link, a nonlinear aggregation algorithm is used to generate a comprehensive video coding efficiency. The index, transmission efficiency aggregation measurement and storage system entropy change index ensure the reliability of the index through parameter calibration. The collaborative decision-making model based on the three-dimensional state vector combines the threshold conditions to generate hierarchical management instructions. When the global management coefficient is higher than 80, the elastic coding and BBRv3 congestion control joint optimization are enabled. The time domain layered coding and storage preheating strategy are activated in the range of 50 to 80. If it is lower than 50, I frame requests and cross-region data migration are enforced. Finally, the control execution layer converts the decision instructions into dynamic optimization of coding parameters, FEC hierarchical control and time-space aware garbage collection strategy. The codec parameter update, protocol stack hot loading and storage resource reconstruction are realized through the distributed control system, forming a full-link management process from data collection to closed-loop control.

[0070] Secondly: In the drawings of the embodiments disclosed in the present invention, only the structures related to the embodiments disclosed in the present invention are involved, and other structures can refer to the general design. In the absence of conflict, the same embodiment and different embodiments of the present invention can be combined with each other;

[0071] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A video surveillance data storage and management system based on big data, characterized in that, Including: Multi-source heterogeneous data acquisition module: By deploying acquisition tools for the video encoding layer, network transmission layer, and distributed storage layer in parallel, it realizes the synchronous capture of raw data. The video encoding layer uses video stream analysis tools to extract frame structures and motion parameters. The network transmission layer captures protocol features and channel states through programmable network devices. The distributed storage layer obtains physical medium states and access logs based on the storage system interface; Cross-layer feature fusion calculation module: It performs streaming feature extraction on the raw data provided by the acquisition layer and calculates three types of metrics for the video encoding efficiency monitoring group, video transmission feature analysis group, and storage feature analysis group respectively; Comprehensive index fusion module: It performs non-linear aggregation on the feature matrix output by the feature calculation layer and uses three groups of formulas with calibrated parameters to generate the video encoding efficiency comprehensive index, transmission efficiency aggregation metric, and storage system entropy change index respectively; Collaborative decision-making generation module: Based on the three-dimensional state vector of the index analysis layer, it calculates the global management coefficient through a third-order collaborative management decision model and generates hierarchical management instructions in combination with preset threshold conditions and state space partitioning rules; Adaptive regulation execution module: It converts the policy instructions generated by the decision-making layer into three core operations: dynamic optimization of video encoding parameters, real-time switching of network transmission protocols, and reconstruction of storage resource distribution strategies. It drives the update of codec parameters, hot loading of network protocol stack configurations, and execution of storage node data migration tasks through a distributed control system.

2. The video surveillance data storage and management system based on big data according to claim 1, wherein: The video encoding efficiency monitoring group includes the spatio-temporal compression ratio, entropy coding fluctuation coefficient, motion prediction residual gradient, and bitrate-quality inflection point slope; the video transmission feature analysis group includes spatio-temporal scheduling correlation, protocol stack efficiency index, channel feature matching degree, and fault tolerance diffusion coefficient; the storage feature analysis group includes hot and cold data vorticity, encoding storage coupling rate, medium loss balance factor, and topology-aware compression rate.

3. The video surveillance data storage and management system based on big data according to claim 2, characterized in that: The spatio-temporal compression ratio calculates the average value of P frames using a sliding window and analyzes the distribution of key frames in combination with the GoP structure, which is specifically expressed as: STR represents the spatio-temporal compression ratio, I^(v) represents the byte size of a single I frame, \bar{P}^(v) represents the average byte size of P frames in the most recent 10 seconds, and ρ_t^(v) represents the key frame density per unit time; the entropy coding fluctuation coefficient is calculated macro-block by macro-block based on the Shannon entropy formula and outliers are filtered using the quantile method, which is specifically expressed as: EVF represents the entropy coding fluctuation coefficient, H_max^(mb) represents the maximum macro-block entropy value within a single frame, H_min^(mb) represents the minimum macro-block entropy value within a single frame, and \bar{H}^(mb) represents the average value of macro-block entropy values within a frame; the motion prediction residual gradient calculates the difference through the motion compensation residual matrix and uses SIMD instructions to accelerate the summation, which is specifically expressed as: MRG represents the motion prediction residual gradient, ΔR_i^(t)=R_i^(t)-R_i^(t-1) represents the change in the residual of the i-th macro-block, A represents the frame area, and N represents the total number of macro-blocks; the slope of the rate-quality inflection point constructs a cubic spline interpolation function for derivation and the resolution term is normalized, which is specifically expressed as: RQS represents the slope of the rate-quality inflection point, represents the derivative of the PSNR-rate curve, and W and h represent the width and height of the video resolution.

4. A video surveillance data storage and management system based on big data according to claim 2, characterized in that: The spatio-temporal scheduling correlation evaluates the network jitter component based on a Kalman filter and is specifically expressed as: STC represents the spatio-temporal scheduling correlation, λ_s represents the spatial shard packet loss rate, T_t represents the number of time shard retransmissions, J represents the network jitter coefficient, and ε represents the anti-zero constant; the protocol stack efficiency index extracts the header fields using a protocol parser and calculates the MTU matching degree, which is specifically expressed as: PSE represents the protocol stack efficiency index, ηp represents the payload ratio, H_T represents the TCP header overhead, H_R represents the RTP header overhead, and α_f represents the shard aggregation degree factor; the channel feature matching degree calculates the Shannon entropy using a sliding window and is specifically expressed as: CCM = 1 - |E B -E V |, CCM represents the channel feature matching degree, E_B represents the channel bandwidth distribution entropy, and E_V represents the video bitrate change entropy; the fault tolerance diffusion coefficient analyzes the redundant structure through an FEC matrix and traverses the hierarchical coding tree, which is specifically expressed as: FDC represents the fault tolerance diffusion coefficient, β_F represents the forward error correction redundancy, d_L represents the hierarchical coding depth, and RTT represents the network round-trip delay.

5. A video surveillance data storage and management system based on big data according to claim 2, characterized in that: The vorticity of hot and cold data constructs an access frequency field to calculate the gradient, which is located using spherical coordinates and is specifically expressed as: HDV represents the vorticity of hot and cold data, represents the access frequency gradient, σ_L represents the storage location dispersion, and H_T represents the data block temperature entropy; the encoding storage coupling rate uses a memory-mapped file to detect boundary alignment, and is specifically expressed as: SCR represents the encoding storage coupling rate, N_align represents the number of block boundary alignments, Sc represents the compressed block size, S_p represents the storage page size, and N_total represents the total number of data blocks; the media loss equalization factor establishes a wear leveling state matrix in combination with temperature sensor data, and is specifically expressed as: WEF represents the media loss equalization factor, W_max / min represents the maximum / minimum number of erase / write cycles, \bar{W} represents the average number of erase / write cycles, and δ_T represents the temperature difference coefficient; the topology-aware compression rate uses locality-sensitive hashing to detect cross-node similarity, and is specifically expressed as: TCR represents the topology-aware compression rate, C_std represents the standard compression rate, D_kj represents the data similarity between nodes k and j, and n represents the total number of nodes.

6. A video surveillance data storage and management system based on big data according to claim 1, characterized in that: The cross-layer feature fusion calculation module constructs a cross-layer joint feature matrix by defining a three-dimensional tensor structure: taking the time axis as the first dimension, it divides fixed-length sliding windows; in the feature dimension, it maps STR, EVF, MRG, and RQS to tensor channels 1-4, STC, PSE, CCM, and FDC to channels 5-8, and HDV, SCR, WEF, and TCR to channels 9-12; it uses a double-buffer mechanism to achieve time series alignment. When and only when the window timestamps of the three data streams overlap by more than 95%, it triggers the splicing operation, optimizes the memory layout using the SIMD instruction set, and finally generates a tensor with the shape of [T×12×N], where T is the number of time windows and N is the batch size, and normalizes it in the channel dimension through LayerNorm to ensure the dimensional consistency of cross-layer features.

7. A video surveillance data storage management system based on big data according to claim 1, characterized in that: The comprehensive video coding efficiency index is specifically expressed as: VECI represents the comprehensive video coding efficiency index, α represents the encoder architecture adaptation constant, and β represents the time-domain continuity compensation coefficient; The specific expression of the transmission efficiency aggregation metric is as follows: TEM represents the transmission efficiency aggregation metric, γ represents the network topology complexity factor, and δ represents the fault tolerance efficiency critical threshold; The entropy change index of the storage system is specifically expressed as: SSEI represents the entropy change index of the storage system, η represents the aging rate of the storage medium, θ represents the optimal coupling threshold, and σ represents the standard deviation of process fluctuations.

8. A video surveillance data storage and management system based on big data according to claim 1, characterized in that: The specific expression of the three - order collaborative management decision - making model is as follows: TCDM represents the global management coefficient, K represents the coding quality weight, μ represents the transmission stability coefficient, v represents the storage pressure damping factor, ω represents the mutation tolerance threshold, VECIΔ represents the coding efficiency change rate in the recent 3 minutes, and TEMΔ represents the transmission efficiency gradient in the recent 5 RTT windows.

9. A video surveillance data storage and management system based on big data according to claim 1, characterized in that: The collaborative decision-making generation module, based on the hierarchical management trigger mechanism of the TCDM value, divides the system state into three response levels by setting two critical thresholds of 80 and 50: when the TCDM value is higher than 80, the system enables the joint optimization of the elastic coding mechanism and the BBRv3 congestion control algorithm, automatically switches to the constant quality factor coding mode while maintaining the video quality, and at the same time optimizes the network queue management to reduce the transmission delay; if the TCDM value is between 50 and 80, the time-domain hierarchical coding technology is activated to achieve bitrate adaptation, and the high-heat storage nodes are preheated in advance to balance the access load; when the TCDM value drops below the 50 threshold, a forced I-frame request is immediately executed to quickly reconstruct the integrity of the video stream, and at the same time, cross-availability zone data migration is started to ensure storage availability.

Citation Information

Patent Citations

  • Industrial big data computing task scheduling management system

    CN115981804A

  • Intelligent factory data acquisition platform and method thereof

    CN119293689A

  • Unattended video remote monitoring system

    CN119653042A

  • Analytics-aware video compression control using end-to-end learning

    US20240275996A1

Cited By

  • Cross-platform database heterogeneous migration and fault-tolerant control method and system

    CN120743888A

  • Cross-platform database heterogeneous migration and fault-tolerant control method and system

    CN120743888B

  • Cross-system cooperative control method and system based on digital thread

    CN120762325A

  • Intelligent information acquisition method and system based on scene design feedback

    CN121328679A

  • Industrial data full life cycle management method and system

    CN121765405A