Intelligent coding storage optimization method and system for video monitoring
Through the multi-scale convolutional neural network and object detection algorithm adaptively adjusting the encoding parameters, combined with environmental sensor data for image enhancement and differential encoding, the video quality and storage transmission efficiency of the video surveillance system in complex environments is solved, and efficient and low-cost video surveillance optimization is achieved.
Patent Information
- Application Number
- CN202510538247.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-07-22
AI Technical Summary
The existing video surveillance system cannot adaptively adjust the encoding parameters in a complex and changeable monitoring environment, resulting in a decline in video quality and low storage and transmission efficiency, serious waste of resources, and difficult to meet the comprehensive requirements of intelligent monitoring systems for efficiency, adaptability and cost control.
Multi-scale convolutional neural network is used to extract video features, combine object detection algorithm to identify dynamic target areas, adaptively adjust encoding parameters, combine environmental sensor data for image enhancement, and optimize video data through differential encoding and keyframe storage, and dynamically select transmission paths.
It realizes dynamic optimization of compression rates in different video scenarios, ensuring image quality while reducing data redundancy, improving the system's video quality and transmission efficiency under low illumination and strong interference conditions, and reducing storage and transmission costs.
Smart Images

Figure CN120358345A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video surveillance, and specifically to an intelligent encoding and storage optimization method and system for video surveillance. Background Art
[0002] With the wide application of video surveillance systems in fields such as urban security, traffic management, and industrial production, how to achieve efficient storage and transmission of large-scale video data while ensuring video quality has become a technical bottleneck that urgently needs to be broken through. Currently, mainstream solutions generally adopt standard compression coding formats such as H.264 or H.265, and uniformly compress video content through fixed GOP structures, quantization parameters (QP), and bitrate control strategies to meet basic storage and transmission requirements.
[0003] However, the existing technical solutions have obvious limitations in dealing with complex and changing surveillance environments. Due to the fixed encoding parameters, which cannot be dynamically adjusted according to the scene content, it is easy to cause image quality degradation, especially under low illumination or strong interference conditions, and the video quality drops significantly. At the same time, such solutions lack effective identification and suppression of image redundancy, resulting in low transmission efficiency and serious resource waste, ultimately increasing the overall operation cost of the system and making it difficult to meet the comprehensive requirements of intelligent surveillance systems for efficiency, adaptability, and cost control. Summary of the Invention
[0004] In view of the deficiencies of the prior art, the present invention provides an intelligent encoding and storage optimization method and system for video surveillance, which solves the problems of video quality degradation and low storage and transmission efficiency caused by the inability of encoding parameters to be adaptively adjusted, poor environmental adaptability, and high data redundancy rate in the prior art.
[0005] To achieve the above objectives, the present invention is realized through the following technical solutions: An intelligent encoding and storage optimization method and system for video surveillance, including the following steps: a) Obtain video frames, extract video features through a multi-scale convolutional neural network on an edge computing node, and identify dynamic target regions in combination with a target detection algorithm; b) Based on the extracted features and target detection results, adaptively adjust the encoding parameters and improve the encoding quality of the dynamic target regions for content-based encoding optimization; c) Combine environmental sensor data to adaptively enhance the image attributes such as brightness and contrast of the video frames, and improve the image clarity in complex environments with low illumination; d) Use spatio-temporal redundancy between video frames for differential encoding, and at the same time select key frames for high-quality storage, and the remaining frames are compressed and stored in a differential manner to reduce the data volume; e) Dynamically select the optimal transmission path according to the network state.
[0006] Preferably, in the step b), the following sub-steps are specifically included: b1) Calculate the content complexity: According to the video frame features extracted from step a), calculate the content complexity index C of the current frame t , and the content complexity index is based on the image texture changes and details within the video frame, and is specifically calculated by the following formula: where I(i,j) is the pixel value of the i-th row and j-th column of the current frame, is the image gradient of this pixel point, and M and N are the number of rows and columns of the video frame respectively; b2) Adaptively adjust the bit rate: According to the calculated content complexity C t , adjust the bit rate R of the video encoding t , where the adjustment of the bit rate satisfies the following relationship: where α is the system adjustment coefficient, indicating the scaling degree of the bit rate, and this adjustment enables the video encoding to be adaptively optimized according to the content complexity, so as to balance the video quality and the compression ratio; b3) Optimize the encoding of the dynamic target area: For the identified dynamic target area, allocate a relatively high encoding quality and adopt a lower quantization step QP fg to maintain high image quality. Specifically, the quantization parameter QP of the target area fg is adjusted in the following way: QP fg = QP base - δ; where QP base is the reference quantization parameter, and δ is the adjustment coefficient, indicating the quantization accuracy of the target area.
[0007] This strategy reduces redundant data by reducing the encoding accuracy of non-target areas and optimizes the video storage and transmission efficiency.
[0008] Preferably, the step c) includes: Calculate the initial value of the target parameter y through the following formula: y0 = f(x1,x2,...,x n ), where x1 to x n are input parameters, and f is a known function; Iteratively update the value of y through the update formula until the target condition is met: y i+1 = g(y i ), where g is the update function and i is the number of iterations.
[0009] Preferably, step c) further includes: Calculating the target value y through the following formula: y = h(x1, x2,..., x n ), where h is the objective function, and x1 to x n are input parameters; adjusting y using an optimization algorithm to minimize the error E, where E is the difference between the target value and the actual value, and is calculated through the following formula: where y i is the predicted value, and
[0010] is the actual value.
[0010] Preferably, in step e), by monitoring network state parameters, the score value of the transmission path is calculated in real time, and the score function S satisfies the following relationship: S = α·B - β·D + γ·L; where B is the current path bandwidth, D is the path delay, L is the path packet loss rate, and α, β, and γ are weighting coefficients. The system optimizes path selection by adjusting the weighting coefficients and selects the path with the highest score for video stream transmission.
[0011] Preferably, step e) further includes: Performing real-time sampling and prediction on network state parameters, estimating the future network bandwidth Bt, delay Dt, and packet loss rate Lt through the Kalman filtering algorithm, and updating the score function S t : where T is the sampling time window, and B i , D i , L i respectively represent the bandwidth, delay, and packet loss rate of the i-th sampling. The optimal real-time transmission path is selected through this prediction method.
[0012] An intelligent coding and storage optimization system for video surveillance, comprising: A multi-scale convolutional neural network unit for performing multi-level feature extraction on video frames; An encoding control unit for dynamically adjusting video encoding parameters according to the feature extraction results; A target detection unit for identifying dynamic targets in video frames and assigning a higher encoding quality to the target areas; An environment perception unit for receiving data from environment sensors and adaptively adjusting the encoding parameters of video frames according to environmental conditions; A storage unit for storing video data, including video data optimized through differential encoding and selected key frames; an edge computing unit for processing video data at the edge nodes of the video surveillance system and reducing bandwidth consumption; A transmission control unit for selecting an optimal transmission path according to real-time bandwidth conditions for real-time transmission of video streams.
[0013] Preferably, the encoding control unit dynamically adjusts parameters such as the coding rate and quantization step size of video encoding according to the output of the multi-scale convolutional neural network.
[0014] Preferably, the storage unit adopts a distributed storage architecture, which can perform redundant storage of video data and perform fast retrieval when necessary.
[0015] Preferably, the transmission control unit selects an optimal transmission path using an intelligent routing algorithm based on real-time network conditions.
[0016] The present invention provides an intelligent encoding storage optimization method and system for video surveillance, having the following beneficial effects: 1. The present invention adopts an intelligent encoding parameter adjustment mechanism driven by deep features, realizing content adaptation of encoding control. It can dynamically optimize the compression ratio in different video scenarios, effectively reducing data redundancy while ensuring image quality. Compared with existing solutions that rely on fixed parameter settings, it breaks through the limitation of insufficient encoding efficiency in complex scenarios.
[0017] 2. The present invention introduces a joint strategy of environment perception and image enhancement, enabling the system to still output clear images under complex conditions such as low illuminance and strong interference, ensuring that surveillance information is not lost. Most current solutions have weak environmental adaptability, effectively avoiding the problem of target recognition failure caused by deteriorated image quality.
[0018] 3. The present invention significantly compresses inter-frame data with high repetition by integrating an end-to-end data redundancy elimination model. In practical applications, it can significantly reduce the network backhaul traffic. Compared with traditional processing methods based on frame difference filtering, it avoids the risk of erroneously deleting valid data and improves the link throughput efficiency at the same time.
[0019] 4. The present invention demonstrates advantages through the coordination of compression quality and resource scheduling, especially in the long-term storage scenario of video surveillance. It supports efficient storage compression while maintaining stable image quality output. Compared with existing technical paths that can only unidirectionally optimize the compression ratio or clarity, it effectively balances image quality and cost. Description of the Drawings
[0020] Figure 1 It is a flowchart of the method steps of the present invention. Detailed Embodiments
[0021] Next, in combination with the accompanying drawings of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0022] Please refer to the attached Figure 1 , the embodiment of the present invention provides an intelligent coding storage optimization method for video surveillance, including the following steps: a) Obtain video frames, extract video features through a multi-scale convolutional neural network on an edge computing node, and identify dynamic target regions in combination with a target detection algorithm; First, obtain the original video frames through a video acquisition device. The video acquisition device includes an image acquisition module, and the image acquisition module is connected to the video input end of the edge computing node.
[0023] The edge computing node includes an image preprocessing module, a feature extraction module, and a target detection module. The image preprocessing module is connected to the image acquisition module and is used for performing brightness normalization and image size adjustment processing on the input video frames.
[0024] Subsequently, the preprocessed video frames are input into the feature extraction module. The feature extraction module includes a multi-scale convolutional neural network for extracting multi-scale texture and edge features in the image. This multi-scale convolutional neural network includes a backbone network and a multi-scale branch network.
[0025] The backbone network is constructed based on the ResNet structure and is used to capture global structure information; the multi-scale branch network adopts a dilated convolution structure to perform parallel modeling on image regions at different spatial scales.
[0026] In the image feature tensor output by the feature extraction module, H, W, and C respectively represent the height, width, and number of channels of the feature map.
[0027] The image feature tensor is input into the target detection module. The target detection module adopts the YOLO structure and outputs prediction results including target category, position information, and confidence in combination with the detection head.
[0028] Specifically, for each frame of image, the target detection module outputs N bounding boxes, and each bounding box is defined as: B i =(x i ,y i ,w i ,h i ,c i ); where x i ,yi is the center coordinate of the bounding box, w i , h i are the width and height respectively, and c i is the confidence level.
[0029] The feature tensor output by the feature extraction module and the detection results are jointly used to construct a dynamic target region mask. The mask matrix M ∈ {0, 1} H×W If the pixel position is within a certain detection bounding box, then set M ij = 1, otherwise M ij = 0.
[0030] This dynamic target region mask is used to guide the region bitrate allocation strategy in the encoding module and serves as an important input for content complexity calculation and encoding parameter adjustment in subsequent step b).
[0031] Through the above structural design, the effective distinction between the target region and the non-target region in the video frame is achieved, and it can provide region-level context awareness information for subsequent encoding optimization.
[0032] In actual deployment, the edge computing node includes an embedded GPU computing module for real-time execution of the above neural network structure. The image acquisition module is connected to the computing node through the MIPI interface, and the neural network model is loaded and inferred through the TensorRT engine to ensure low-latency running performance.
[0033] In summary, by constructing a multi-scale convolutional feature extraction structure and a real-time target detection algorithm, the input video frame can be encoded with structured features, providing an accurate spatial context basis for encoding control.
[0034] b) Based on the proposed features and target detection results, adaptively adjust the encoding parameters and improve the encoding quality of the dynamic target region for content-based encoding optimization; b1) Calculate the content complexity: According to the video frame features extracted from step a), calculate the content complexity metric C t of the current frame. The content complexity metric is based on the image texture changes and details within the video frame and is specifically calculated through the following formula: where I(i, j) is the pixel value at the i-th row and j-th column of the current frame, is the image gradient of this pixel point, and M and N are the number of rows and columns of the video frame respectively; b2) Adaptively adjust the bitrate: According to the calculated content complexity C t , adjust the bitrate R of the video encoding t, where the adjustment of the bit rate satisfies the following relationship: Among them, α is the system adjustment coefficient, representing the scaling degree of the bit rate. This adjustment enables the video encoding to be adaptively optimized according to the content complexity, thereby balancing the video quality and the compression ratio; b3) Dynamic target area encoding optimization: For the identified dynamic target area, a relatively high encoding quality is assigned, and a lower quantization step QP is adopted fg to maintain high image quality. Specifically, the quantization parameter QP of the target area fg is adjusted in the following way: QP fg = QP base - δ; Among them, QP base is the reference quantization parameter, and δ is the adjustment coefficient, representing the quantization accuracy of the target area.
[0035] This strategy reduces redundant data by reducing the encoding accuracy of the non-target area and optimizes the video storage and transmission efficiency.
[0036] The image enhancement module adjusts the brightness and contrast of the image according to the environmental sensor data and outputs the enhanced image. The enhanced image data will be transmitted to the subsequent video encoding module for further compression and encoding processing.
[0037] In addition, the image enhancement module can also perform denoising processing according to the noise level of the image. In environments such as low light or high humidity, there may be noise in the image. The image enhancement module smooths the image through denoising algorithms (such as Gaussian filtering, mean filtering, etc.) to reduce the impact of noise on the encoding effect.
[0038] c) Combining the environmental sensor data, adaptively enhance the image attributes such as the brightness and contrast of the video frame to improve the image clarity in complex environments such as low illumination; In step c), first obtain the data of the environmental sensor module, which includes multiple sensors, such as temperature sensor T env , humidity sensor H env and light sensor L env . These sensors are connected to the edge computing node by wireless or wired means. The real-time collected environmental data will be used to adjust operations such as the brightness, contrast, and denoising of the image.
[0039] The environmental light intensity L received by the image enhancement module env is used to dynamically adjust the image brightness. The specific brightness adjustment is carried out through the following formula: I new = α·Iorig + β·L env ; Wherein, I orig is the pixel value of the original image, I new is the enhanced pixel value, α and β are the brightness enhancement coefficient and the weighted coefficient of the illumination intensity respectively, and L env is the ambient illumination intensity. This adjustment can adaptively adjust the brightness of the image according to the environmental changes, thereby improving the image quality under low illuminance.
[0040] Meanwhile, according to the formula of dependent claim 3, in the contrast adjustment of the image, the environmental temperature and humidity information are used to calculate the final image contrast adjustment factor. The specific calculation formula is: Wherein, C represents the contrast of the original image, and C ′ represents the adjusted contrast, T env is the environmental temperature, H env is the environmental humidity, T base is the temperature reference value, T m ax and T m in are the maximum and minimum values of the temperature, H m ax is the maximum value of the humidity, and H m ax is the maximum value of the humidity, and α and β are the adjustment coefficients of the temperature and humidity. Through this formula, the contrast of the image is adaptively adjusted according to the current environmental temperature and humidity, and the detailed performance of the image can be optimized under different environments.
[0041] According to the content of dependent claim 4, further, the image enhancement module can also perform image denoising processing according to the noise level collected by the sensor. Especially in an environment with low light or high humidity, the image may be affected by noise, and the denoising algorithm will smooth the image by means of Gaussian filtering, mean filtering, etc. to remove the noise points.
[0042] Finally, the enhanced image will be transmitted to the subsequent video coding module for compression processing. During the encoding process, the encoding bit rate R t will be dynamically adjusted according to the content complexity defined in step b). The enhanced image will improve the quality of the encoded video and ensure that it can adapt to different network conditions during the transmission process.
[0043] d) Using the spatio-temporal redundancy between video frames for differential coding, and at the same time selecting key frames for high-quality storage, and the remaining frames are compressed and stored in a differential manner to reduce the data volume; In step d), first, the video coding module performs differential coding on consecutive video frames. Specifically, by comparing the current frame F t and the previous frame Ft-1 Compare them, calculate the difference between two frames, and obtain the difference image ΔF t This difference image only retains the changed parts between the two frames, thus effectively reducing the amount of redundant data and improving the coding efficiency.
[0044] The calculation of the difference image is mainly achieved by comparing the pixel values of the current frame and the previous frame pixel by pixel. In this way, only the changed parts in the image (such as motion, scene changes, etc.) will be encoded and compressed, avoiding repeated encoding of the same parts.
[0045] Next, motion estimation technology is adopted to further compress the difference image. This process calculates the motion vector of each image block by finding the best matching region between the previous frame and the current frame. The motion vector reflects the offset of the image block in time, which is used to predict the pixel values in the current frame and reduce the amount of encoded data.
[0046] For each image block, the prediction result of the previous frame is compared with the current frame through the motion compensation method to obtain the difference part. This difference part is compressed by differential coding. The coding bit rate is dynamically adjusted according to the size and complexity of the difference image to ensure the coding efficiency and quality of the video.
[0047] Through the above steps, differential coding and motion compensation can effectively reduce the amount of redundant data, lower the bit rate required for video coding, and at the same time ensure the balance between video quality and data compression. This process can adapt to the changes in video scenes and optimize the coding efficiency on the premise of ensuring video quality.
[0048] Finally, the compressed data will be transmitted or stored for subsequent decoding and playback.
[0049] e) Dynamically select the optimal transmission path according to the network state; In step e), first, the sensor network monitors the current network state parameters, including network bandwidth B, delay D, packet loss rate L, etc. These network state parameters are fed back to the network path selection module in real time for path evaluation.
[0050] The network path selection module calculates the score value S of different paths through the analysis of the current network state. The scoring function S comprehensively considers factors such as bandwidth, delay, and packet loss rate to ensure the selection of the optimal transmission path under various network conditions. The scoring function S is calculated by the following formula: S = α·B - β·D + γ·L; Among them, B is the current path bandwidth, D is the path delay, L is the path packet loss rate, and α, β, and γ are weighting coefficients, which respectively represent the influence degrees of bandwidth, delay, and packet loss rate on path selection. The system dynamically adjusts the weight coefficients according to these parameters to adapt to different network environments.
[0051] Perform real-time sampling and prediction on network state parameters, estimate the future network bandwidth Bt, delay Dt, and packet loss rate Lt through the Kalman filtering algorithm, and update the scoring function S t : Where T is the sampling time window, B i , D i , L i respectively represent the bandwidth, delay, and packet loss rate of the i-th sampling, and select the optimal real-time transmission path through this prediction method.
[0052] In order to further improve the accuracy of path selection, the network path selection module also predicts the future network state through historical data. Perform real-time sampling and prediction on network bandwidth, delay, and packet loss rate through the Kalman filtering algorithm, and calculate the future values of network state. The prediction is carried out through the following formula: Among them, represents the estimated value of the network state, A is the state transition matrix, B is the control matrix, u k is the control input, and k is the time step. The Kalman filter uses this historical data and model to predict the future network bandwidth Bt, delay Dt, and packet loss rate Lt, thereby optimizing path selection.
[0053] According to the current network state and prediction results, the path selection module calculates the score S of each path t and selects the path with the highest score for video data transmission. Finally, the selected path will be used for the transmission of the video stream, and the stability of the path will be monitored in real time. If the network state changes significantly, the path selection module will recalculate the path score and reselect the optimal path according to the new network state.
[0054] The intelligent coding storage optimization system for video surveillance includes: A multi-scale convolutional neural network unit for performing multi-level feature extraction on video frames; The multi-scale convolutional neural network (CNN) unit plays a crucial role in the video surveillance system, mainly used for extracting deep features from each frame of video data. This unit uses multiple convolutional layers with different scales, combined with pooling layers, to process image features at different levels, enabling it to extract rich information from local details to global structures. Through this multi-scale convolutional structure, the system can identify complex dynamic scenes and subtle motion patterns, such as targets under low-light conditions, fast-moving objects, and the distinction between static backgrounds and dynamic targets.
[0055] The network captures spatio-temporal information in video frames through hierarchical convolutional operations and feature fusion, providing richer semantic inputs for subsequent video encoding and object detection. These features not only help the object detection unit identify objects in the scene but also provide a basis for the encoding control unit to dynamically adjust the encoding strategy. The design of the convolutional neural network adopts a residual structure to solve the problem of gradient disappearance in the training of deep networks, improving the network training efficiency and feature extraction accuracy.
[0056] An encoding control unit, used to dynamically adjust video encoding parameters according to the feature extraction results; The core function of the encoding control unit is to dynamically adjust video encoding parameters based on the feature results extracted by the multi-scale convolutional neural network. This unit determines the encoding strategy for different regions by analyzing the content complexity, motion frequency, and importance of the target area of the video frame in real time. For complex motion scenes or key target areas, the encoding control unit will select a higher bitrate for fine encoding to ensure that the video image can still maintain high clarity and details under limited bandwidth.
[0057] The bitrate adjustment formula in the encoding control unit is: Where: R t is the bitrate of the video frame at time t; C t is the content complexity of the video frame (such as the degree of motion, image details, etc.); α is an adjustment coefficient, adjusted based on factors such as network bandwidth and storage capacity.
[0058] An object detection unit, used to identify dynamic targets in video frames and assign higher encoding quality to the target areas; The target detection unit uses deep learning models (such as YOLO, Faster R-CNN, RetinaNet, etc.) to identify and locate targets in video frames. This unit can real-time identify target objects in the video stream, such as people, vehicles, animals, etc., and judge their importance through intelligent algorithms, allocating higher coding quality to key areas. For dynamic targets (such as fast-moving objects or important event areas), this unit will mark these target areas as "key areas" and avoid distortion by improving their coding quality, ensuring high availability of the targets in subsequent video analysis and archiving.
[0059] The confidence formula for target recognition in the target detection unit is: Where: S target is the target score output by the target detection model; P target is the detection confidence of this target.
[0060] The environmental perception unit is used to receive data from environmental sensors and adaptively adjust the encoding parameters of video frames according to environmental conditions; The environmental perception unit integrates multiple sensors (such as temperature and humidity sensors, light sensors, wind speed sensors, etc.) to real-time perceive and monitor environmental changes. The data of these sensors provides auxiliary information for video encoding, enabling the encoding control to automatically adjust according to environmental conditions. For example, the light sensor can monitor the ambient light intensity. If the light is weak, the system will automatically increase the brightness of the video frame or adjust the exposure time; the temperature and humidity sensor can reflect the changes in environmental temperature and humidity. Combining the video content and sensor data, the system can timely adjust the frame rate or resolution of the video to cope with environmental interference and maintain the quality and stability of the video picture.
[0061] By continuously collecting and analyzing environmental data, the environmental perception unit can more intelligently adjust the video encoding strategy to ensure the best performance of video monitoring under different environmental conditions, especially in outdoor monitoring, warehouse environments, or scenarios affected by special weather.
[0062] The adjustment formula for the impact of environmental light conditions on encoding is: Where: L is the current ambient light intensity (unit: Lux); L th is the preset light threshold; ΔI is the light change amount; γ is the light adjustment coefficient.
[0063] The storage unit is used to store video data, including video data optimized by differential encoding and selected key frames; The storage unit is responsible for efficiently storing video data and optimizing the storage space. This unit not only reduces storage redundancy through traditional differential coding but also optimizes the storage strategy through an intelligent key frame selection mechanism. Based on the results of the target detection unit, the system assigns high-quality coding to key target areas (such as moving targets or important scenes) and stores the corresponding key frames. In contrast, for static areas or irrelevant scenes, more efficient differential coding or low-bitrate coding is used for storage to save space.
[0064] In addition, the storage unit supports hierarchical storage management of data, allocating video data to different storage media according to access frequency and importance. Frequently accessed data is stored in high-speed storage devices (such as solid-state drives or high-speed cloud storage), while less frequently accessed data is stored in storage media with larger capacities but slower response times (such as HDDs or cold storage), ensuring reasonable allocation of resources. The edge computing unit is used to process video data at the edge nodes of the video surveillance system and reduce bandwidth consumption. The edge computing unit is the distributed computing core of the entire system, responsible for processing a large amount of video data at the edge nodes of the video surveillance system, reducing the computing pressure on the central server and improving the response speed. The edge computing unit can perform real-time processing on the video stream locally, including operations such as image enhancement, target detection, preliminary coding, and data filtering. Only the optimized data will be transmitted to the central server or the cloud for further processing. This processing method not only reduces bandwidth consumption but also significantly reduces the impact caused by network latency, enhancing the real-time response ability of the system.
[0065] The processing delay formula of the edge computing unit is: Where: D video is the amount of video data to be processed (unit: bit); B edge is the processing bandwidth of the edge computing unit (unit: bit / second); T proc is the processing time of edge computing.
[0066] The transmission control unit is used to select the optimal transmission path for real-time transmission of the video stream according to real-time bandwidth conditions; the role of the transmission control unit is to select the optimal transmission path according to real-time network conditions such as network bandwidth, latency, and packet loss rate to ensure the stable transmission of the video stream. This unit adopts an intelligent transmission algorithm to dynamically adjust parameters such as the bitrate and resolution of the video stream according to the bandwidth requirements of the video content, ensuring that video data can still be transmitted efficiently and with low latency under bandwidth constraints.
[0067] The transmission control unit can also monitor the status of the network channel in real time and automatically switch to the optimal path according to the changing network conditions to avoid data loss caused by network congestion or failures. If there is a problem with the main transmission path, the system will automatically switch to the alternative path, thus ensuring the continuous transmission of the video stream. The transmission control unit also has an adaptive function and can dynamically adjust the transmission strategy according to different video transmission requirements (such as real-time video stream, event playback, etc.) to ensure the high efficiency and stability of the system.
[0068] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An intelligent coding and storage optimization method for video surveillance, characterized in that, It includes the following steps: a) Obtain video frames, extract video features on the edge computing node through a multi-scale convolutional neural network, and identify the dynamic target area by combining with the object detection algorithm; b) Based on the extracted features and object detection results, adaptively adjust the encoding parameters and improve the encoding quality of the dynamic target area for content-based encoding optimization; c) Combine the environmental sensor data to adaptively enhance the image attributes such as brightness and contrast of the video frames, and improve the image clarity in complex environments with low illuminance; d) Use the spatio-temporal redundancy between video frames for differential encoding, and at the same time select key frames for high-quality storage, and the remaining frames are compressed and stored in a differential manner to reduce the data volume; e) Dynamically select the optimal transmission path according to the network status.
2. The intelligent encoding and storage optimization method for video surveillance according to claim 1, characterized in that, In the step b), it specifically includes the following sub-steps: b1) Calculate the content complexity: Calculate the content complexity metric C of the current frame based on the video frame features extracted in step a) t , where the content complexity metric is based on the image texture changes and details within the video frame and is specifically calculated using the following formula: where I(i,j) is the pixel value at the i-th row and j-th column of the current frame, is the image gradient of this pixel point, and M and N are the number of rows and columns of the video frame respectively; b2) Adaptively adjust the bit rate: According to the calculated content complexity C t , adjust the bit rate R of video encoding t , where the adjustment of the bit rate satisfies the following relationship: Among them, α is the system adjustment coefficient, indicating the scaling degree of the bit rate. This adjustment enables the video encoding to be adaptively optimized according to the content complexity, so as to balance the video quality and compression ratio; b3) Optimize the encoding of the dynamic target area: For the identified dynamic target region, a relatively high coding quality is assigned, and a lower quantization step size QP is adopted fg to maintain high image quality. Specifically, the quantization parameter QP of the target region fg is adjusted in the following way: QP fg = QP base - δ; Among them, QP base is the reference quantization parameter, and δ is the adjustment coefficient, representing the quantization accuracy of the target region; This strategy reduces redundant data by reducing the encoding accuracy of non-target areas and optimizes the video storage and transmission efficiency.
3. The intelligent coding and storage optimization method for video surveillance according to claim 1, characterized in that The step c) includes: Calculate the initial value of the target parameter y through the following formula: y0 = f(x1, x2,..., x n ), where x1 to x n are input parameters and f is a known function; Iteratively update the value of y through the update formula until the target condition is met: y i+1 = g(y i ), where g is the update function and i is the iteration number.
4. The intelligent encoding and storage optimization method for video surveillance according to claim 1, characterized in that The step c) also includes: Calculate the target value y through the following formula: y = h(x1, x2,..., x n ), where h is the target function, and x1 to x n are input parameters; use the optimization algorithm to adjust y to minimize the error E, where E is the difference between the target value and the actual value, and is calculated through the following formula: where y i is the estimated value, is the actual value.
5. The intelligent encoding and storage optimization method for video surveillance according to claim 1, characterized in that In the step e), by monitoring the network status parameters, the score value of the transmission path is calculated in real time, and the score function S satisfies the following relationship: S = α·B - β·D + γ·L; Among them, B is the current path bandwidth, D is the path delay, L is the path packet loss rate, and α, β, and γ are weighting coefficients. The system optimizes the path selection by adjusting the weighting coefficients and selects the path with the highest score for video stream transmission.
6. The intelligent coding storage optimization method for video surveillance according to claim 1, characterized in that The step e) further includes: Real-time sampling and prediction of network state parameters are carried out, and the future network bandwidth Bt, delay Dt, and packet loss rate Lt are estimated through the Kalman filtering algorithm, and the scoring function S is updated t : where T is the sampling time window, B i , D i , L i respectively represent the bandwidth, delay, and packet loss rate of the i-th sampling, and the optimal real-time transmission path is selected through this prediction method.
7. An intelligent encoding and storage optimization system for video surveillance, which is used for the intelligent encoding and storage optimization method for video surveillance according to any one of claims 1-6, characterized in that, It includes: A multi-scale convolutional neural network unit for performing multi-level feature extraction on video frames; An encoding control unit for dynamically adjusting video encoding parameters according to the feature extraction results; An object detection unit for identifying dynamic targets in video frames and assigning higher encoding quality to the target areas; An environmental perception unit for receiving data from environmental sensors and adaptively adjusting the encoding parameters of video frames according to environmental conditions; A storage unit for storing video data, including video data optimized by differential encoding and selected key frames; An edge computing unit for processing video data at the edge nodes of the video surveillance system and reducing bandwidth consumption; A transmission control unit for selecting the optimal transmission path according to the real-time bandwidth conditions for real-time transmission of video streams.
8. The intelligent encoding and storage optimization system for video surveillance according to claim 7, characterized in that The encoding control unit dynamically adjusts the parameters of the video coding bit rate and quantization step according to the output of the multi-scale convolutional neural network.
9. The intelligent coding and storage optimization system for video surveillance according to claim 7, wherein The storage unit adopts a distributed storage architecture, which can perform redundant storage of video data and perform fast retrieval when necessary.
10. The intelligent encoding and storage optimization system for video surveillance according to claim 7, characterized in that The transmission control unit selects the optimal transmission path based on real-time network conditions using an intelligent routing algorithm.
Citation Information
Patent Citations
HEVC intra-frame code rate control method for optimizing monitoring video perception quality
CN112291564A
Network state evaluation method, device and system, equipment and storage medium
CN118301429A
Video stream acquisition method based on deep learning
CN119299703A
Image recognition system based on machine vision
CN119399534A
Edge-type surveillance camera AI abnormal situation detection and control device and method
KR102647328B1
Cited By
Video monitoring system data storage and intelligent compression optimization method and system
CN121217912A