A static scene video transmission optimization method
Patent Information
- Application Number
- CN202611092349.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-22
- Publication Date
- 2026-09-29
AI Technical Summary
[0006]综上所述,现有技术存在以下核心缺陷:静态识别精度不足,多依赖单一像素特征与固定阈值判定,无法区分环境光影干扰与真实动态变化,微动态场景误判率较高;压缩方式粗放,部分方案通过丢帧换取带宽,牺牲帧率流畅度;优化维度单一,仅针对编码或传输单环节做优化,未联动硬件功耗与数据存储,综合收益有限;传输与存储安全性弱,多输出标准编码码流,通用播放器可直接解码播放,视频数据易被拷贝、篡改,无法满足高安全场景的防护要求
本发明通过采用多级分层静态场景识别与环境自适应阈值机制,实现了对完全静态、弱微静态、强微静态及动态切换场景的精准判定,有效降低了环境干扰下的误判率;通过基于长效基准参考帧的分级增量编码策略,在保持原始帧率和主观画质无感知下降的前提下,实现了静态场景下的码率大幅压缩;通过联动调整编码参数、传输参数和硬件功耗参数,实现了编码模块、传输模块和功耗模块对当前场景的协同适配,有效降低了设备功耗与存储成本;通过对编码后的码流采用私有协议帧结构进行封装加扰,生成不兼容标准解码器的安全传输数据,实现了视频数据的防篡改和防拷贝能力。
Smart Images

Figure CN122845762A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of video surveillance and transmission technology, and in particular to a method for optimizing video transmission in static scenes. Background Technology
[0002] In the current security monitoring field, network cameras and analog cameras are widely used in scenarios such as park warehouses, forest fire prevention, building corridors, and unattended computer rooms. According to statistics, static or low-dynamic scenes account for 60% to 85% of the daily output in these scenarios. However, the existing video transmission system still uses encoding and transmission strategies for general dynamic scenes without deep adaptation for static scenes, exposing multiple problems: low bandwidth utilization, static scenes are still transmitted at a high base bitrate, and a large number of redundant frames occupy bandwidth resources; high device power consumption, encoding modules run at full load continuously, the battery life of solar-powered or battery-powered edge cameras is shortened, and the risk of thermal failure increases in high-temperature environments; inefficient storage and retrieval, with full storage of highly repetitive static scenes, driving up storage costs.
[0003] Currently, the existing implementations most similar to this invention mainly fall into the following three categories: Option 1 is a fixed-threshold static frame skipping transmission scheme. It uses an absolute error sum algorithm to calculate the pixel difference between consecutive frames, sets a fixed pixel difference threshold, and classifies frames below the threshold as static frames, transmitting only the frame index and metadata. The receiving end reconstructs the image from the buffered frames. Frames above the threshold are then transmitted as complete frames. This scheme can reduce bandwidth usage from a high level to a low level in completely static scenarios, but its fixed threshold cannot adapt to micro-dynamic interference, fluctuations in the difference value are prone to misjudgment, and there is no reduction in power consumption during static periods.
[0004] Option 2 is a video transmission scheme based on narrowband conditions. It establishes a background model through an inter-frame differential algorithm, skips intermediate frame transmission when the scene is static, and only periodically updates the background reference frame. The receiving end reconstructs the scene by buffering the background frames, and resumes transmission of complete frames when a moving target is detected. This scheme can reduce the number of transmitted frames by 60% to 70% in completely static scenes, but it relies on only a single pixel feature for static judgment, resulting in a high false judgment rate under environmental interference; it reduces bandwidth by dropping frames, and the frame rate decrease during static periods causes a stuttering effect; it only optimizes a single transmission link, without improving encoding power consumption and storage redundancy; and it outputs a standard encoded bitstream that can be directly decoded by general-purpose players, lacking proprietary protocols and security verification mechanisms, making the video data susceptible to leakage and tampering.
[0005] Option 3 is a timed static period bitrate reduction scheme. Users preset static periods, during which the bitrate and frame rate are uniformly reduced. Outside of preset periods, the normal parameters are restored. This scheme can reduce device power consumption within the preset period and has a low deployment threshold, but it relies on manual preset and cannot cope with dynamic switching between static and dynamic periods; the parameter adjustment is not dynamically optimized according to the duration of the static period.
[0006] In summary, existing technologies suffer from the following core defects: insufficient accuracy in static recognition, relying heavily on single pixel features and fixed thresholds, failing to distinguish between ambient light and shadow interference and real dynamic changes, resulting in a high false positive rate in micro-dynamic scenes; crude compression methods, with some solutions sacrificing frame rate smoothness by dropping frames to gain bandwidth; limited optimization dimensions, focusing only on the encoding or transmission stage without considering hardware power consumption and data storage, leading to limited overall benefits; and weak transmission and storage security, often outputting standard encoded bitstreams that can be directly decoded and played by general-purpose players, making video data easily copied and tampered with, thus failing to meet the protection requirements of high-security scenarios. Summary of the Invention
[0007] The main objective of this invention is to provide a method for optimizing video transmission in static scenes.
[0008] Another objective of this invention is to provide a static scene video transmission optimization device.
[0009] The third objective of this invention is to provide a computer device.
[0010] The fourth objective of this invention is to provide a non-transitory computer-readable storage medium.
[0011] To achieve the above objectives, a first aspect of the present invention proposes a method for optimizing video transmission in static scenes, comprising: Multi-level hierarchical static scene recognition is performed on the acquired video frames. By filtering block-level pixel differences, comparing texture features and verifying motion vectors, combined with environmental adaptive threshold updates, the scene type is output. Based on a long-term reference frame, a hierarchical incremental coding strategy is adopted for coding according to the scene type. Among them, a frame-level zero-pixel incremental coding is adopted for completely static scenes, a macroblock-level incremental coding is adopted for weakly static scenes, and a frame-level differential incremental coding is adopted for strongly static scenes. The encoding parameters, transmission parameters, and hardware power consumption parameters are adjusted in conjunction with the scenario type to enable the encoding module, transmission module, and power consumption module to adapt to the current scenario in a coordinated manner. The encoded bitstream is encapsulated and scrambled using a proprietary protocol frame structure to generate secure transmission data that is incompatible with standard decoders, and then output to the receiving end or storage medium.
[0012] In one embodiment of the present invention, the step of performing multi-level hierarchical static scene recognition on the acquired video frames includes: Calculate the pixel difference between the macroblocks corresponding to the current frame and the previous frame, and count the maximum global difference, the average difference, and the proportion of blocks exceeding the threshold to quickly eliminate dynamic images and lock suspected static frames. Extract the texture features of the frame, calculate the texture similarity with historical frames, and distinguish between pixel brightness fluctuations and changes in real content to filter out environmental interference. Calculate motion vectors for preset sensitive areas, and combine edge continuity verification to determine whether the change is the movement of a physical target in order to eliminate non-physical interference; Based on historical data and scene baselines, the judgment threshold range is adaptively adjusted, and location-specific threshold self-learning is supported. Based on the above recognition results, the scene determination results are output as either completely static, weakly static, strongly static, or dynamically switching.
[0013] In one embodiment of the present invention, the step of encoding based on a long-term reference frame and employing a hierarchical incremental coding strategy according to the scene type includes: When the scene type is completely static, a baseline reference frame is generated as a global long-term reference. Subsequent continuous static frames are not pixel-level encoded, but only frame-level incremental metadata is output. The frame-level incremental metadata is transmitted in the form of a verification data packet plus timestamp metadata, and the baseline frame is periodically verified and updated.
[0014] When the scene type is weak static, the frame is divided into macroblocks with the reference frame as the prediction reference. Static macroblocks reuse the corresponding pixels of the reference frame and skip residual coding. Micro-variable macroblocks only encode the residual incremental data with the reference frame. Dynamic macroblocks perform full coding. When the scene type is strong, weak, and static, the reference frame is used as the long-term reference frame. Each frame only encodes the overall differential increment with the reference frame. The quantization parameters are increased in static areas, while the standard image quality is maintained in dynamic areas. At the same time, the reference frame refresh interval is increased and the key frame insertion frequency is reduced.
[0015] In one embodiment of the present invention, the step of adjusting the encoding parameters, transmission parameters, and hardware power consumption parameters in conjunction with the scene type includes: When the scene type is completely static, the encoding module enters a low-power standby mode to perform only integrity verification, the transmission module only transmits metadata, and the power consumption module causes the encoding module to reduce its frequency and the network module to reduce its power. When the scene type is weak static, the encoding module reserves some computing resources, and the transmission module transmits the layered incremental encoded bitstream and encapsulates and scrambles it through a private protocol. When the scene type is strong, weak, static, the encoding module maintains the normal operating level and performs a slight frequency reduction, while the transmission module transmits the complete incremental encoded bitstream and encapsulates and scrambles it with a private protocol, combined with frame header compression optimization. When the scene type is dynamic switching, the pre-cached full set of dynamic scene parameters are directly called, the incremental coding mode is exited and the standard inter-frame predictive coding is restored.
[0016] In one embodiment of the present invention, the encapsulation and scrambling of the encoded bitstream using a private protocol frame structure includes: The encoded bitstream is re-encapsulated according to the privately defined frame header, frame payload, and frame trailer structure. The frame header contains a frame type identifier, timestamp, and length field; the frame payload contains scrambled encoded data; and the frame trailer contains an integrity check value. The frame payload data is scrambled using a preset scrambling algorithm, so that the encapsulated data cannot be directly parsed by the standard decoder.
[0017] In one embodiment of the present invention, it further includes: After the device is powered on, it reads the computing power level of the main control chip and activates the corresponding recognition mode and compression strategy according to the computing power level.
[0018] In one embodiment of the present invention, the storage and retrieval of video data in static scenes are further optimized: In a completely static scene, only the metadata of the base frame is stored; In weak static scenes, only the data of the changing pixel blocks are stored, and each incremental data is bound to the corresponding reference frame through a reference identifier to form a chain storage structure; When a dynamic switching event is detected, an event tag record is automatically generated and stored independently. Generate an independent integrity check value for each baseline index and incremental data entry, and bind and store it with the data. Physical storage media are managed by logical partitions, and the oldest static incremental data is overwritten first using a cyclic overwrite strategy.
[0019] To achieve the above objectives, a second aspect of the present invention provides a static scene video transmission optimization device, comprising: The scene recognition module is used to perform multi-level hierarchical static scene recognition on the acquired video frames. It outputs the scene type by filtering block-level pixel differences, comparing texture features and verifying motion vectors, combined with environmental adaptive threshold updates. The incremental coding module is used to encode based on a long-term reference frame and according to the scene type using a hierarchical incremental coding strategy. The linkage adjustment module is used to adjust the encoding parameters, transmission parameters, and hardware power consumption parameters in a linkage manner according to the scene type. The secure encapsulation module is used to encapsulate and scramble the encoded bitstream using a proprietary protocol frame structure, generating secure transmission data that is incompatible with standard decoders.
[0020] To achieve the above objectives, a third aspect of this application provides a computer device, including a processor and a memory; wherein the processor reads executable program code stored in the memory to run a program corresponding to the executable program code, for implementing a static scene video transmission optimization method as described in the first aspect embodiment.
[0021] To achieve the above objectives, a fourth aspect of this application provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a static scene video transmission optimization method as described in the first aspect embodiment.
[0022] The embodiments of the present invention have the following beneficial effects: This invention employs a multi-level hierarchical static scene recognition and environmental adaptive threshold mechanism to accurately determine completely static, weakly static, strongly slightly static, and dynamically switching scenes, effectively reducing the false judgment rate under environmental interference. Through a hierarchical incremental coding strategy based on a long-term reference frame, it achieves significant bitrate compression in static scenes while maintaining the original frame rate and imperceptible decrease in subjective image quality. By coordinating and adjusting coding parameters, transmission parameters, and hardware power consumption parameters, it achieves collaborative adaptation of the coding module, transmission module, and power consumption module to the current scene, effectively reducing device power consumption and storage costs. By encapsulating and scrambling the encoded bitstream using a proprietary protocol frame structure, it generates secure transmission data incompatible with standard decoders, achieving anti-tampering and anti-copying capabilities for video data. Attached Figure Description
[0023] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein: Figure 1 A flowchart illustrating a static scene video transmission optimization method provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the static frame storage structure provided in an embodiment of the present invention; Figure 3 An architecture diagram of a static scene video transmission optimization system provided in an embodiment of the present invention; Figure 4 This is a structural diagram of a static scene video transmission optimization device provided in an embodiment of the present invention. Detailed Implementation
[0024] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0025] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0026] The following describes a static scene video transmission optimization method according to an embodiment of the present invention, with reference to the accompanying drawings.
[0027] Example 1 This embodiment provides a method for optimizing video transmission in static scenes, such as... Figure 1 As shown, the method includes the following steps: S1 performs multi-level hierarchical static scene recognition on the acquired video frames. Through block-level pixel difference filtering, texture feature comparison and motion vector verification, combined with environment adaptive threshold update, the scene type is output.
[0028] S2, based on the long-term reference frame, a hierarchical incremental coding strategy is adopted for coding according to the scene type. Among them, the completely static scene adopts frame-level zero-pixel incremental coding, the weakly static scene adopts macroblock-level incremental coding, and the strongly static scene adopts frame-level differential incremental coding.
[0029] S3, adjust the encoding parameters, transmission parameters and hardware power consumption parameters in conjunction with the scene type, so that the encoding module, transmission module and power consumption module can work together to adapt to the current scene.
[0030] S4 encapsulates and scrambles the encoded bitstream using a proprietary protocol frame structure to generate secure transmission data that is incompatible with standard decoders, and outputs it to the receiving end or storage medium.
[0031] Specifically, step S1 includes: First, a continuous frame sequence is extracted at a configurable period (200ms to 1000ms), denoised by Gaussian filtering, converted to YUV420 format, and divided into 16×16 pixel macroblocks.
[0032] Next, a first-level block-level coarse screening is performed. The sum of absolute errors (SAD) for each macroblock is calculated. The pixel values of the corresponding macroblocks in the current frame and the previous frame are subtracted pixel by pixel, and the absolute values are summed to obtain the SAD value for each macroblock. The maximum SAD value among all macroblocks is taken as the global maximum SAD value. The arithmetic mean of the SAD values of all macroblocks is calculated as the SAD mean. The proportion of macroblocks with SAD values exceeding a preset threshold is taken as the percentage of macroblocks exceeding the threshold. Using these three metrics, highly dynamic scenes are quickly eliminated, and suspected static frames are identified.
[0033] Then, the second level of texture refinement is performed. For each pixel within each macroblock, a 3×3 neighborhood is taken centered on that pixel. The gray values of the neighboring pixels are compared with the gray value of the center pixel to generate a binary code. The LBP encoded histogram of all pixels within the macroblock is then used as the texture feature vector of that macroblock. The cosine similarity between the texture feature vectors of each macroblock in the current frame and the texture feature vectors of the corresponding macroblocks in the previous 5 frames is calculated. If the cosine similarity is lower than a preset texture change threshold, it is determined to be a texture change region. This distinguishes pixel brightness fluctuations from changes in actual content and filters out environmental interference such as light flickering and gradual changes in illumination.
[0034] Furthermore, a third-level semantic verification is performed. Corner feature points are extracted from preset sensitive areas such as doorways and passageways, and the motion vectors of feature points between adjacent frames are calculated using the LK optical flow method. Cluster analysis is performed on the motion vectors. If the clustered motion vectors have the same direction and the amplitude exceeds a preset motion threshold, the edge contours of the motion region are further extracted, and the continuity index of the edge contours is calculated. If the continuity index is higher than a preset continuity threshold, it is determined to be the motion of a physical target; otherwise, it is determined to be non-physical interference, thereby eliminating non-physical interference such as noise and light drift.
[0035] Regarding threshold updates, it has a built-in ambient brightness fluctuation factor that adaptively adjusts the judgment threshold range based on nearly 10 frames of historical data and scene baseline; it also supports 24-hour scene self-learning to generate location-specific thresholds.
[0036] Finally, based on the above recognition results, the method outputs scene determination results for completely static, weakly static, strongly slightly static, or dynamically switching scenarios. This method can achieve a scene recognition accuracy of no less than 98%, with a false judgment rate of no more than 2% under environmental interference.
[0037] Specifically, step S2 includes: This step builds a hierarchical incremental coding mechanism based on a long-term reference frame, and matches incremental coding strategies with different granularities for different static scenes.
[0038] When the scene type is completely static, frame-level zero-pixel incremental encoding is used. A high-quality baseline reference frame is generated as a global long-term reference. Subsequent continuous static frames do not undergo pixel-level encoding; only frame-level incremental metadata is output. The metadata includes the LBP hash checksum, frame number, and timestamp, with the data size of a single frame not exceeding 32 bytes. The baseline frame is automatically checked and updated every 30 seconds. Under this strategy, the bitrate compression ratio for completely static scenes is no less than 93%.
[0039] When the scene type is weakly static, macroblock-level incremental coding is used. Using the reference frame as the prediction reference, the image is divided into 16×16 macroblocks. Static macroblocks fully reuse the corresponding pixels of the reference frame, skipping residual coding and retaining only the macroblock skip marker. Micro-variable macroblocks use QP=35 to 40, encoding only the residual incremental data from the reference frame. A small number of dynamic macroblocks maintain the standard QP=28, performing full coding to preserve details. Static area data volume is compressed by more than 90%, and the overall bitrate is reduced by 70% to 80%.
[0040] When the scene type is strongly static, frame-level differential incremental coding is used. The reference frame is used as the long-term reference frame, and each frame only encodes the overall differential increment with the reference frame; the QP is appropriately increased in static areas, while the standard image quality is maintained in dynamic areas; at the same time, the reference frame refresh interval is increased and the key frame insertion frequency is reduced, resulting in an overall bitrate reduction of 40% to 60%.
[0041] When the scene type is dynamic switching, exit incremental coding mode and resume standard GOP inter-frame predictive coding, with a switching delay of no more than 100ms.
[0042] Specifically, step S3 includes: Based on the scene recognition results in step S1, this step adjusts the operating parameters of the encoding module, transmission module, and hardware power consumption module in a coordinated manner, so that the three modules can work together to adapt to the current scene and achieve end-to-end resource optimization.
[0043] When the scene type is completely static, the encoding module enters a low-power standby mode, only performing hash integrity verification, and the encoding computing power usage does not exceed 5%; the transmission module switches to heartbeat transmission mode, only transmitting 16-byte LBP hash heartbeat packets and timestamp metadata, and all transmitted data is encapsulated and scrambled using a private protocol frame structure. The receiving end maintains the original 25fps output frame rate based on the cached reference frame; the power consumption module shuts down some GPU cores, causing the encoding module to reduce its frequency and go into standby mode, and the network module to reduce its power consumption to less than 0.3W, with the total power consumption of the whole machine not exceeding 1W.
[0044] When the scene type is weak and static, the encoding module reserves some GPU cores for macroblock-level incremental encoding. Static macroblocks skip residual encoding and only retain the skip mark. Micro-variable macroblocks encode residual incremental data, and dynamic macroblocks perform complete encoding. The transmission module switches to a hierarchical incremental transmission mode to transmit hierarchical incremental encoded bitstreams. All bitstreams are encapsulated and scrambled using a private protocol, reducing the overall bitrate by 70% to 80%. The power consumption module reduces the power of the network module to 0.5W.
[0045] When the scene type is strong-slight-static, the encoding module maintains the normal operating level and performs a slight frequency reduction, increases the reference frame refresh interval and reduces the key frame insertion frequency; the transmission module switches to the full incremental transmission mode, transmits the full incremental encoded bitstream, and is encapsulated and scrambled by a private protocol and optimized with frame header compression, reducing the overall bit rate by 40% to 60%.
[0046] When the scene type is dynamic switching, the encoding module, transmission module and power consumption module directly call the pre-cached full set of dynamic scene parameters to restore the standard GOP inter-frame predictive coding, standard transmission parameters and standard power consumption level, with a switching delay of no more than 100ms.
[0047] Specifically, step S4 includes: The encoded bitstream is re-encapsulated according to a privately defined frame header, frame payload, and frame trailer structure. The frame header contains a frame type identifier, timestamp, and length field; the frame payload contains scrambled encoded data; and the frame trailer contains an integrity check value. The frame payload data is scrambled using a preset scrambling algorithm, making the encapsulated data unparseable by standard H.265 or H.264 decoders.
[0048] Furthermore, this embodiment also includes storage and retrieval optimization steps. A four-level storage structure of "baseline index + incremental mounting + event mapping + security verification" is adopted for video data in static scenes. In completely static scenes, only four types of metadata are stored: the unique hash identifier of the base frame, timestamp, resolution, and encoding parameters. The size of a single index data entry does not exceed 64 bytes. When the static scene is continuous and the image features do not change substantially, the base frame is not repeatedly generated; only the original hash index is used. In slightly static scenes, the complete frame image is not stored; only the coordinates, pixel values, and data length of the changed pixel blocks are stored. Each incremental data entry is bound to the corresponding base frame through a base hash ID, forming a chained storage structure of "1 base index + N incremental data entries," with the storage size of a single frame compressed by more than 90% compared to a complete frame. When a dynamic switching event is detected, an event tag record is automatically generated, including the event type, trigger time, corresponding base frame hash ID, and associated incremental frame sequence number. The event tag is stored independently in the event index area. During retrieval, the target segment is directly located through the tag, and the retrieval time does not exceed 1 second. The security verification unit independently generates SHA-256 integrity verification values for each base index and each incremental data, and stores them in data-bound format; it automatically verifies the data during read and playback, and can immediately identify and alert if the data has been tampered with.
[0049] like Figure 2 As shown, the physical storage medium is divided into four logical partitions: the base index area, the incremental data area, the event index area, and the dynamic frame storage area. It is managed according to the cyclic overwrite strategy, which prioritizes overwriting the earliest static incremental data to ensure the retention priority of event data and the base index.
[0050] In addition, this embodiment also includes an adaptive switching step for the operating mode. After the device is powered on, the computing power level of the main control chip is read; if the computing power is greater than or equal to 4 TOPS, the high-performance mode is enabled, which enables complete three-level hierarchical recognition, real-time threshold update, environment self-learning and AI scene classification; if the computing power is less than or equal to 2 TOPS, the lightweight mode is enabled, which is simplified to two-level recognition (SAD and LBP), fixed periodic threshold update and basic hierarchical compression strategy.
[0051] In summary, this embodiment achieves accurate scene recognition with a success rate of no less than 98% and a false judgment rate of no more than 2% under environmental interference by employing a three-level hierarchical static scene recognition and environmental adaptive threshold mechanism. Through a hierarchical incremental coding strategy based on a long-term benchmark reference frame, it achieves an extreme compression effect of no less than 93% bitrate compression for fully static scenes and a 70% to 85% reduction in bitrate for micro-static scenes, while maintaining the original 25fps frame rate and imperceptible decrease in subjective image quality. By adjusting encoding, transmission, and power consumption parameters in a coordinated manner, it achieves end-to-end resource optimization, reducing device power consumption by more than 50% and storage costs by more than 60% for static scenes. Finally, by encapsulating and scrambling the encoded bitstream using a private protocol and performing SHA-256 integrity verification, it achieves anti-tampering and anti-copy capabilities for transmitted and stored data.
[0052] Example 2 This embodiment demonstrates the application of 4K high-definition monitoring in an unmanned warehouse scenario. The equipment consists of a 4K high-definition network camera with a mains computing power of 5 TOPS, powered by mains electricity, and deployed during non-operational warehouse hours. The scenario is characterized by predominantly weak static data during 8 hours of nighttime, with minor dynamic interference such as flickering lights.
[0053] After adopting the method of this embodiment, light interference is effectively filtered through three-level recognition, and the false judgment rate is controlled within 1.5%; after the scene is determined to be weak static scene, macroblock-level incremental coding is enabled, and the overall bit rate is reduced from 4Mbps to 0.8Mbps, a reduction of 80%, with no significant decrease in subjective image quality; after private protocol encapsulation, general players cannot decode it, and the risk of data leakage is greatly reduced; device power consumption is reduced from 3.5W to 1.3W; 8 hours of storage is reduced by 65% compared with the original solution; the dynamic switching delay for personnel intrusion is 82ms, and there is no blurry uphill period in the picture.
[0054] Example 3 This embodiment demonstrates the application of a solar-powered camera in a forest fire prevention scenario. The equipment consists of a 1080P low-power network camera with a main control computing power of 1.5 TOPS, powered by solar cells, and deployed for nighttime monitoring in a mountainous forest. The scenario is characterized by a completely static environment for 12 hours at night, with no significant dynamic targets and only occasional minor movements of branches and leaves.
[0055] After adopting the method of this embodiment, the scene is identified as completely static. Frame-level zero-pixel incremental coding and heartbeat transmission mode are enabled, reducing the single-channel bit rate from 2Mbps to 100kbps, achieving a compression ratio of 95%, and maintaining a full frame rate output of 25fps. Data transmission is protected against eavesdropping and tampering after scrambling with a private protocol. Device power consumption is reduced from 2.8W to 0.9W, and battery life is extended from 18 days to 32 days. The switching latency for wild animal intrusion is 90ms, meeting the requirements for real-time monitoring.
[0056] Example 4 This embodiment provides a static scene video transmission optimization system, such as Figure 3 As shown, this system is integrated into the camera's main control chip and implemented based on an embedded system. It has a three-layer architecture: an input layer, a core processing layer, and an output layer. The core processing layer includes a scene recognition module, a full-link control module, a storage and retrieval module, and a mode switching module. These modules interact via an internal data bus and transmit commands via a control bus.
[0057] The input layer is used to capture raw video frames through the camera image sensor and lens, and output raw YUV format images to the core processing layer.
[0058] The scene recognition module includes a feature extraction unit, a dynamic threshold calculation unit, and an environmental interference adaptation unit. It is responsible for frame sampling preprocessing, multi-level feature extraction, environmental adaptive threshold update, and scene classification determination, and outputs scene type labels that are completely static, weakly static, strongly static, or dynamically switched.
[0059] The end-to-end control module comprises an encoding control unit, a transmission control unit, a power consumption control unit, a pre-buffered parameter unit, and a proprietary protocol encapsulation unit. Upon receiving a scene type tag, this module adjusts the encoding compression strategy, transmission data format, and hardware computing power level accordingly. During static periods, it pre-stores multiple levels of dynamic parameters for rapid retrieval during dynamic switching. The proprietary protocol encapsulation unit is responsible for encapsulating the encoded bitstream with a proprietary frame structure and performing scrambling to generate secure transmission data incompatible with standard decoders.
[0060] The storage and retrieval module comprises a feature indexing unit, an incremental storage unit, an event retrieval unit, and a security verification unit. This module stores a baseline index for static frames and incremental blocks for micro-static frames, while simultaneously generating event tags to support fast retrieval. The security verification unit generates integrity verification values for each frame's index and incremental data, ensuring tamper-proof storage.
[0061] The mode switching module includes a computing power detection unit and a strategy matching unit, which are used to automatically read the device's computing power after the device is powered on and match the corresponding lightweight or high-performance operating mode.
[0062] The output layer is used to output the encoded video stream to a remote monitoring platform and write the index and video data to the local storage medium.
[0063] Example 5 This invention also provides a static scene video transmission optimization device, such as... Figure 4 As shown, the device 10 includes: The scene recognition module 100 is used to perform multi-level hierarchical static scene recognition on the acquired video frames. Through block-level pixel difference filtering, texture feature comparison and motion vector verification, combined with environment adaptive threshold update, the scene type is output.
[0064] The incremental coding module 200 is used to encode based on a long-term reference frame and according to the scene type using a hierarchical incremental coding strategy.
[0065] The linkage adjustment module 300 is used to adjust the encoding parameters, transmission parameters and hardware power consumption parameters in linkage according to the scene type.
[0066] The secure encapsulation module 400 is used to encapsulate and scramble the encoded bitstream using a private protocol frame structure to generate secure transmission data that is incompatible with standard decoders.
[0067] Example 6 To implement the methods of the above embodiments, the present invention also provides a computer device, which includes a memory and a processor; wherein the processor runs a program corresponding to the executable program code by reading executable program code stored in the memory, so as to implement the various steps of the methods described above.
[0068] Example 7 To implement the above embodiments, this application also proposes a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the method described in the foregoing embodiments.
[0069] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
[0070] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0071] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.
Claims
1. A method for optimizing video transmission in static scenes, characterized in that, Includes the following steps: Multi-level hierarchical static scene recognition is performed on the acquired video frames. By filtering block-level pixel differences, comparing texture features and verifying motion vectors, combined with environmental adaptive threshold updates, the scene type is output. Based on a long-term reference frame, a hierarchical incremental coding strategy is adopted for coding according to the scene type. Among them, a frame-level zero-pixel incremental coding is adopted for completely static scenes, a macroblock-level incremental coding is adopted for weakly static scenes, and a frame-level differential incremental coding is adopted for strongly static scenes. The encoding parameters, transmission parameters, and hardware power consumption parameters are adjusted in conjunction with the scenario type to enable the encoding module, transmission module, and power consumption module to adapt to the current scenario in a coordinated manner. The encoded bitstream is encapsulated and scrambled using a proprietary protocol frame structure to generate secure transmission data that is incompatible with standard decoders, and then output to the receiving end or storage medium.
2. The method according to claim 1, characterized in that, The process of performing multi-level hierarchical static scene recognition on the acquired video frames includes: Calculate the pixel difference between the macroblocks corresponding to the current frame and the previous frame, and count the maximum global difference, the average difference, and the proportion of blocks exceeding the threshold to quickly eliminate dynamic images and lock suspected static frames. Extract the texture features of the frame, calculate the texture similarity with historical frames, and distinguish between pixel brightness fluctuations and changes in real content to filter out environmental interference. Calculate motion vectors for preset sensitive areas, and combine edge continuity verification to determine whether the change is the movement of a physical target in order to eliminate non-physical interference; Based on historical data and scene baselines, the judgment threshold range is adaptively adjusted, and location-specific threshold self-learning is supported. Based on the above recognition results, the scene determination results are output as either completely static, weakly static, strongly static, or dynamically switching.
3. The method according to claim 1, characterized in that, The encoding based on the long-term reference frame, using a hierarchical incremental coding strategy according to the scene type, includes: When the scene type is completely static, a baseline reference frame is generated as a global long-term reference. Subsequent continuous static frames are not pixel-level encoded, but only frame-level incremental metadata is output. The frame-level incremental metadata is transmitted in the form of a verification data packet plus timestamp metadata, and the baseline frame is periodically verified and updated. When the scene type is weak static, the frame is divided into macroblocks with the reference frame as the prediction reference. Static macroblocks reuse the corresponding pixels of the reference frame and skip residual coding. Micro-variable macroblocks only encode the residual incremental data with the reference frame. Dynamic macroblocks perform full coding. When the scene type is strong, weak, and static, the reference frame is used as the long-term reference frame. Each frame only encodes the overall differential increment with the reference frame. The quantization parameters are increased in static areas, while the standard image quality is maintained in dynamic areas. At the same time, the reference frame refresh interval is increased and the key frame insertion frequency is reduced.
4. The method according to claim 1, characterized in that, The step of adjusting encoding parameters, transmission parameters, and hardware power consumption parameters in conjunction with the scenario type includes: When the scene type is completely static, the encoding module enters a low-power standby mode to perform only integrity verification, the transmission module only transmits metadata, and the power consumption module causes the encoding module to reduce its frequency and the network module to reduce its power. When the scene type is weak static, the encoding module reserves some computing resources, and the transmission module transmits the layered incremental encoded bitstream and encapsulates and scrambles it through a private protocol. When the scene type is strong, weak, static, the encoding module maintains the normal operating level and performs a slight frequency reduction, while the transmission module transmits the complete incremental encoded bitstream and encapsulates and scrambles it with a private protocol, combined with frame header compression optimization. When the scene type is dynamic switching, the pre-cached full set of dynamic scene parameters are directly called, the incremental coding mode is exited and the standard inter-frame predictive coding is restored.
5. The method according to claim 1, characterized in that, The encapsulation and scrambling of the encoded bitstream using a private protocol frame structure includes: The encoded bitstream is re-encapsulated according to the privately defined frame header, frame payload, and frame trailer structure. The frame header contains a frame type identifier, timestamp, and length field; the frame payload contains scrambled encoded data; and the frame trailer contains an integrity check value. The frame payload data is scrambled using a preset scrambling algorithm, so that the encapsulated data cannot be directly parsed by the standard decoder.
6. The method according to claim 1, characterized in that, Also includes: After the device is powered on, it reads the computing power level of the main control chip and activates the corresponding recognition mode and compression strategy according to the computing power level.
7. The method according to claim 1, characterized in that, It also includes optimizations for storing and retrieving video data in static scenes: In a completely static scene, only the metadata of the base frame is stored; In weak static scenes, only the data of the changing pixel blocks are stored, and each incremental data is bound to the corresponding reference frame through a reference identifier to form a chain storage structure; When a dynamic switching event is detected, an event tag record is automatically generated and stored independently. Generate an independent integrity check value for each baseline index and incremental data entry, and bind and store it with the data. Physical storage media are managed by logical partitions, and the oldest static incremental data is overwritten first using a cyclic overwrite strategy.
8. A static scene video transmission optimization device, characterized in that, include: The scene recognition module is used to perform multi-level hierarchical static scene recognition on the acquired video frames. It outputs the scene type by filtering block-level pixel differences, comparing texture features and verifying motion vectors, combined with environmental adaptive threshold updates. The incremental coding module is used to encode based on a long-term reference frame and according to the scene type using a hierarchical incremental coding strategy. The linkage adjustment module is used to adjust the encoding parameters, transmission parameters, and hardware power consumption parameters in a linkage manner according to the scene type. The secure encapsulation module is used to encapsulate and scramble the encoded bitstream using a proprietary protocol frame structure, generating secure transmission data that is incompatible with standard decoders.
9. A computer device, characterized in that, Including processor and memory; The processor reads executable program code stored in the memory to run a program corresponding to the executable program code, so as to implement a static scene video transmission optimization method as described in any one of claims 1-7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements a static scene video transmission optimization method as described in any one of claims 1-7.