A camera network quality perception-based video code rate self-adaptive adjustment method

CN122824862APending Publication Date: 2026-09-25TOCODING TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610998235.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-06
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

这种恶性循环不仅无法有效修复画面质量,反而会将原本轻微的网络拥塞迅速演变为完全的传输瘫痪,最终导致视频流完全中断,用户体验急剧恶化

Benefits of technology

本发明能够在网络环境恶化的关键时刻避免雪上加霜的传输冲击,确保系统始终在可控的负载范围内运行。本发明实现了画面修复过程的平滑化和可控化,消除了传统方案中巨量数据突发传输对脆弱网络链路的打击,从而将原本可能导致系统崩溃的恶性循环转化为渐进式的良性恢复过程。这种处理方式不仅保证了视频传输的连续性和稳定性,更重要的是在网络资源极度稀缺的环境下实现了传输效率的最大化,使得系统能够在各种复杂网络条件下保持稳健的性能表现。此外,本发明还具备自适应调节能力,能够根据网络状况的实时变化动态优化传输策略,既避免了保守策略导致的资源浪费,也防止了激进策略引发的系统风险。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122824862A_ABST
    Figure CN122824862A_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of video coding transmission, and discloses a camera-based network quality sensing video code rate self-adaptive adjustment method; network state feedback data is acquired to calculate a congestion index and generate a code rate down-regulation instruction, and the concurrent situation of I frame request instructions is intelligently monitored during the execution of the code rate down-regulation process. When the I frame congestion collapse trap state is detected, the traditional key frame coding task is actively intercepted and converted into a distributed gradual refresh frame sequence. The sequence contains multiple transition P frames, each frame only contains part of the intra-frame coding macro block determined based on the congestion index, and the total data volume is constrained by the target code rate. The code rate shaping output of the sending queue is transmitted to achieve picture distortion repair without avoiding burst massive data transmission. The present application realizes smooth control of transmission load and gradual recovery of picture quality, and ensures the stability and continuity of video transmission in a complex network environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video encoding and transmission technology, and more specifically, to a method for adaptive adjustment of video bitrate based on network quality awareness of a camera. Background Technology

[0002] Existing technologies suffer from a "congestion-crash trap" problem during video transmission. The core of this problem lies in the logical conflict between network adaptive mechanisms and image quality restoration mechanisms. When the network environment deteriorates, the intelligent bitrate control system rationally reduces the video bitrate to adapt to limited bandwidth resources. However, network jitter and packet loss simultaneously cause image distortions such as pixelation and screen tearing at the decoding end. To quickly correct these visual defects, the system must request keyframes for image reconstruction. The problem is that keyframes, as independent decodeable frames containing complete image information, are often several times larger than ordinary frames. This is akin to forcibly inserting a super-large truck into an already congested highway. When the network link is already nearing saturation, this sudden influx of massive data packets can instantly overwhelm the network's capacity, leading to buffer overflows, a surge in transmission latency, and even a large-scale retransmission storm. After keyframe transmission failures or severe delays, the decoding end continuously detects image anomalies, thus issuing more keyframe requests, creating a cycle of "the worse the network, the more repair is needed; the more repair is done, the worse the network becomes." This vicious cycle not only fails to effectively restore image quality but also rapidly escalates minor network congestion into complete transmission paralysis, ultimately leading to a complete interruption of the video stream and a drastic deterioration in user experience. Current technology lacks the intelligent identification and coordinated handling capabilities for such concurrent conflict scenarios, making it impossible to ensure image restoration while avoiding further impact on network conditions.

[0003] In view of this, the present invention proposes a video bitrate adaptive adjustment method based on network quality awareness for cameras to solve the above problems. Summary of the Invention

[0004] To overcome the aforementioned deficiencies of the prior art and to achieve the above objectives, the present invention provides the following technical solution: a video bitrate adaptive adjustment method for a camera based on network quality awareness, comprising: The network status feedback data of the camera is acquired, the current network congestion index is calculated based on the network status feedback data, and a bit rate reduction instruction is generated based on the current network congestion index. The network status feedback data includes round-trip time, packet loss rate, and jitter variance. During the execution of the bitrate reduction instruction, it is monitored whether an I-frame request instruction for requesting keyframes is received. The I-frame request instruction is generated by the decoding end when it detects image distortion or bitstream error. If an I-frame request instruction is received during the execution of the bitrate downsampling instruction, it is identified that the current state is an I-frame congestion crash trap, and the traditional keyframe encoding task generated based on the I-frame request instruction is intercepted. Macroblock refresh quotas are determined based on the current network congestion index, and the traditional keyframe coding task is converted into a distributed progressive refresh frame sequence. The distributed progressive refresh frame sequence includes at least one transition P-frame. Each transition P-frame contains intra-coded macroblocks allocated based on the macroblock refresh quota, and the total data volume of each transition P-frame is constrained by the target bitrate corresponding to the bitrate downsampling instruction. The distributed progressive refresh frame sequence is written into the sending queue in a time sequence. The sending queue performs bitrate shaping output based on the target bitrate, so that the camera can repair image distortion while avoiding sending a sudden surge of keyframes.

[0005] The technical effects and advantages of the video bitrate adaptive adjustment method for cameras based on network quality awareness according to the present invention are as follows: This invention can prevent exacerbating transmission disruptions during critical moments of network degradation, ensuring the system always operates within a controllable load range. It smooths and controls the image restoration process, eliminating the impact of massive data bursts on fragile network links found in traditional solutions. This transforms a vicious cycle that could lead to system collapse into a gradual, benign recovery process. This approach not only guarantees the continuity and stability of video transmission but, more importantly, maximizes transmission efficiency in environments with extremely scarce network resources, enabling the system to maintain robust performance under various complex network conditions. Furthermore, this invention possesses adaptive adjustment capabilities, dynamically optimizing transmission strategies based on real-time changes in network conditions. This avoids resource waste caused by conservative strategies and prevents system risks arising from aggressive strategies. Attached Figure Description

[0006] Figure 1 This is a schematic diagram of a video bitrate adaptive adjustment method based on network quality awareness for a camera according to the present invention. Detailed Implementation

[0007] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0008] This application provides a video bitrate adaptive adjustment method for cameras based on network quality awareness. This method is specifically applied to intelligent bitrate control in real-time video transmission scenarios, enabling accurate prediction of network congestion and proactive avoidance of I-frame congestion traps. The implementing entities include, but are not limited to, network cameras, video encoding devices, streaming media transmission platforms, edge computing nodes, and other devices with video encoding and network transmission capabilities, all of which can be considered as the general control nodes in this application.

[0009] Please see Figure 1 In this embodiment of the invention, a specific implementation process of a video bitrate adaptive adjustment method based on network quality awareness for a camera includes: The system acquires network status feedback data from the camera, calculates the current network congestion index based on this data, and generates a bitrate reduction command based on the current network congestion index. The network status feedback data includes round-trip time (RTT), packet loss rate, and jitter variance. The acquisition process collects three key indicators in real time from the network transport layer: RTT, cumulative packet loss rate, and jitter variance. RTT reflects the transmission delay of data packets in the network, packet loss rate characterizes the proportion of packets lost due to network congestion, and jitter variance quantifies the stability of the delay. After normalizing this raw data, the current network congestion index is calculated according to a preset congestion weighting coefficient. This index comprehensively reflects the overall health of the network. When the congestion index exceeds a preset first congestion threshold, a bitrate reduction command is automatically generated. This command carries the target bitrate value and bitrate transition period parameters to guide the encoder to smoothly reduce the output bitrate and avoid exacerbating network congestion.

[0010] During the execution of bitrate downsampling instructions, the system monitors for the receipt of I-frame request commands for requesting keyframes. These I-frame request commands are generated by the decoder when it detects image distortion or bitstream errors. The monitoring process continuously scans the network control channel during bitrate downsampling to detect the receipt of I-frame request commands from the decoder. These commands are typically transmitted using RTCP feedback messages or dedicated control protocols, carrying a request timestamp and information about the distorted area. When the decoder detects pixelation, screen tearing, or decoding errors, it proactively initiates an I-frame request, hoping the encoder will send a complete keyframe to correct the error propagation caused by the corrupted reference frame. The received I-frame request commands undergo timestamp verification and validity checks to ensure the authenticity and timeliness of the request, providing accurate trigger signals for subsequent trap identification.

[0011] If an I-frame request command is received during the execution of a bitrate downsampling command, the system is identified as being in an I-frame congestion collapse trap state and the traditional keyframe encoding task generated based on the I-frame request command is intercepted. The identification process first obtains the generation timestamp of the I-frame request command, then queries the execution progress of the current bitrate downsampling command to determine whether the current bitrate has reached the target bitrate or is still within the bitrate transition period. When concurrent occurrence of I-frame requests and bitrate downsampling is detected, the system is determined to be in an I-frame congestion collapse trap state—a typical vicious cycle scenario: network congestion leads to packet loss, packet loss causes image distortion, the decoder requests I-frame repair, and the large size of the I-frame further exacerbates network congestion, forming a positive feedback collapse. Upon identifying the trap state, the encoding thread in the encoder used to generate complete instant decode refresh frames (IDR frames) is immediately intercepted, blocking the packet encapsulation operation of traditional I-frames and preventing large keyframes from entering the transmission queue during congestion.

[0012] Macroblock refresh quotas are determined based on the current network congestion index, and the traditional keyframe coding task is transformed into a distributed progressive refresh frame sequence. This sequence includes at least one transition P-frame, each containing intra-coded macroblocks allocated based on the macroblock refresh quota. The total data volume of each transition P-frame is constrained by the target bitrate corresponding to the bitrate downsampling instruction. The conversion process first reads the total number of macroblocks in the current image frame, which depends on the video resolution and macroblock size (typically 16×16 pixels). The refresh ratio of intra-coded macroblocks is determined based on the current network congestion index; the more severe the congestion, the smaller the refresh ratio, showing a negative correlation to ensure that the refresh load of a single frame does not exceed the network's capacity. The macroblock refresh quota is calculated based on the total number of macroblocks and the refresh ratio. This quota defines the maximum number of intra-coded macroblocks allowed in a single transition P-frame. The image frame is divided into unrefreshed macroblock regions and macroblock regions to be refreshed. For the regions to be refreshed, macroblocks are evenly distributed across N consecutive transition P-frames according to the macroblock refresh quota, where N is the integer part of the ratio of the total number of macroblocks to the macroblock refresh quota. Intra-prediction coding is applied to the macroblocks allocated in each transition P-frame, while inter-prediction coding is applied to the unallocated macroblocks, thus generating a distributed progressive refresh frame sequence. This method distributes the refresh load of a traditional I-frame across multiple P-frames, maintaining the size of each transition P-frame within the target bitrate constraint, achieving both image restoration and avoiding burst traffic.

[0013] The distributed progressive refresh frame sequence is written into the transmission queue in a time sequence. The transmission queue performs bitrate shaping based on the target bitrate to correct image distortion by avoiding the transmission of sudden large numbers of keyframes. During the output process, each transition P-frame in the distributed progressive refresh frame sequence is sequentially pushed into the transmission queue, which is equipped with a token bucket shaper. The token generation rate of the token bucket is directly controlled by the target bitrate. When the frame data volume of a transition P-frame exceeds the number of currently available tokens in the token bucket, the frame is split into multiple data packet fragments, and these fragments are sent sequentially to the network link according to the token replenishment rate, ensuring that the transmission rate does not exceed the target bitrate setting. Packet loss feedback during the transmission of the distributed progressive refresh frame sequence is monitored. If an intra-frame coded macroblock is lost in any transition P-frame, a new I-frame request response is not triggered. Instead, the macroblock address of the lost intra-frame coded macroblock is added back to the refresh queue of subsequent transition P-frames, and the refresh is completed through subsequent frames, maintaining the smoothness and continuity of the entire repair process.

[0014] In this embodiment of the invention, the detailed implementation steps of acquiring network status feedback data from the camera, calculating the current network congestion index based on the network status feedback data, and generating a bitrate reduction command based on the current network congestion index include: The network transport layer's round-trip time (RTT) gradient change, cumulative packet loss rate, and jitter variance are acquired according to a preset probing period. The probing process uses a fixed probing period (typically 1-5 seconds), and at the end of each period, the latest state data is collected from the network transport layer. The RTT gradient change is obtained by calculating the RTT difference between two consecutive periods, reflecting the trend of latency changes rather than absolute values, thus more sensitively capturing signals of network deterioration. The cumulative packet loss rate is the ratio of the total number of packets sent to the number of confirmed lost packets within the probing period, quantifying the severity of packet loss. Jitter variance is calculated by taking the variance of multiple RTT samples within the probing period, measuring the degree of latency fluctuation; high jitter usually indicates network instability. A sliding window smoothing technique is used to process this raw data, reducing the interference of instantaneous fluctuations on the judgment and extracting stable trend features.

[0015] The round-trip delay gradient, cumulative packet loss rate, and jitter variance are normalized respectively, and the current network congestion index is calculated based on preset congestion weight coefficients. The normalization process maps the three parameters with different dimensions to the [0, 1] interval, typically using max-min normalization or Z-score standardization. The statistical distribution of historical observation data is maintained to determine the normalization benchmark for each parameter. The congestion weight coefficients are preset according to the importance of each parameter in congestion assessment; a typical weight allocation is: round-trip delay gradient 0.4, cumulative packet loss rate 0.4, and jitter variance 0.2, reflecting the core role of delay variation and packet loss in congestion assessment. The current network congestion index is calculated by weighted summation: ; in This represents the current network congestion index. , , These are three congestion weighting coefficients. , , These represent the normalized round-trip delay gradient, packet loss rate, and jitter variance, respectively. The congestion index ranges from [0, 1], with a larger value indicating more severe network congestion.

[0016] The current network congestion index is compared with a preset first congestion threshold. When the current network congestion index exceeds the first congestion threshold, a target rate reduction ratio is calculated based on the difference between the current network congestion index and the first congestion threshold. The comparison process involves comparing the real-time calculated congestion index with the preset first congestion threshold (typically 0.6-0.7). When the congestion index exceeds the threshold, the network is considered to be in a congested state, requiring a bitrate reduction. The calculation of the target rate reduction ratio considers the degree to which the threshold is exceeded; the greater the exceedance, the larger the reduction, using a linear or non-linear mapping relationship. A maximum reduction ratio upper limit (usually 50%) is set to prevent a sudden drop in bitrate from affecting video quality, while a minimum reduction ratio lower limit (usually 10%) is set to ensure the effectiveness of the adjustment. The specific value of the reduction ratio is obtained by looking up a table or calculating it in a preset mapping function based on the exceedance difference. This mapping function comprehensively considers the balance between network recovery time and video quality maintenance.

[0017] A bitrate reduction command is generated based on a target reduction ratio. This command carries the target bitrate and a bitrate transition period, which indicates the duration of the smooth bitrate decrease. Specifically, the target bitrate value is first calculated based on the current encoded bitrate and the target reduction ratio. The target bitrate equals the current bitrate multiplied by (1 - reduction ratio). The bitrate transition period is set based on the severity of network congestion; the more severe the congestion, the shorter the transition period to speed up response. In cases of mild congestion, a longer transition period is set to ensure smooth bitrate adjustment and avoid abrupt changes in video quality. A typical transition period range is 3-10 seconds. The bitrate reduction command is generated as a structured message, containing fields such as the target bitrate, transition period, and priority identifier, and is sent to the encoder's bitrate control module for execution. Upon receiving the command, the encoder initiates the smooth bitrate decrease process, gradually adjusting the quantization parameter (QP) and encoding complexity within the transition period to achieve a smooth transition from the current bitrate to the target bitrate, avoiding abrupt changes in image quality caused by sudden bitrate drops.

[0018] In this embodiment of the invention, if an I-frame request instruction is received during the execution of a bitrate downsampling instruction, the detailed implementation steps for identifying the current state as an I-frame congestion collapse trap and intercepting the traditional keyframe encoding task generated based on the I-frame request instruction include: Upon receiving an I-frame request command, the system retrieves the timestamp of the command's generation and queries the execution progress of the current bitrate downscaling command. First, the timestamp field is parsed from the received I-frame request command message; this timestamp records the precise moment the decoder initiated the request. Simultaneously, the status register of the bitrate control module is queried to read the key parameters of the currently executing bitrate downscaling command: command initiation time, target bitrate, transition period, and current actual bitrate. Execution progress is determined by comparing the time difference between the current moment and the command initiation time with the transition period, while also checking whether the current actual bitrate has reached or is close to the target bitrate. A bitrate adjustment state machine is maintained to track bitrate changes in real time, providing accurate status information for progress queries.

[0019] If the execution progress indicates that the current bitrate has not yet reached the target bitrate, or if the current time is within the bitrate transition period, then the I-frame request command and the bitrate reduction command are considered concurrent. The determination process uses a dual-condition check: Condition one checks the difference between the current actual bitrate and the target bitrate; if the difference exceeds a threshold (usually 5% of the target bitrate), it is considered that the target bitrate has not yet been reached. Condition two checks whether the current time is within the time window of the bitrate transition period, determined by calculating (current time - command initiation time) to see if it is less than the transition period. Meeting either condition constitutes a concurrent state, because the network is still in the process of congestion management, and the bitrate has not yet stabilized at a lower, safe level. The identification of a concurrent state means that, before network congestion has been alleviated, the decoding end requests I-frame repair due to packet loss or errors, which is a typical precursor to a congestion trap.

[0020] When I-frame request commands and bitrate downscaling commands are executed concurrently, the current state is marked as an I-frame congestion trap state, and the encoding thread in the encoder used to generate complete real-time decoded refresh frames is suspended, blocking the packetization operation of complete real-time decoded refresh frames. Specifically, a "congestion trap state" flag is set in the global state register, which triggers the activation of a series of protection mechanisms. The encoder internally maintains multiple encoding threads, among which the thread responsible for I-frame generation normally responds to I-frame requests and generates complete IDR frames. The suspension operation puts this thread into a paused state through a thread synchronization mechanism, preventing it from executing encoding tasks. Simultaneously, an interceptor is set in the packetization module to prevent I-frame encoded data from being encapsulated into Network Abstraction Layer Units (NALUs) and written to the transmission buffer, even if I-frame encoded data is generated. This dual interception mechanism ensures that in the trap state, no large-volume complete I-frames enter the network transmission channel, fundamentally blocking the path to a vicious cycle.

[0021] In this embodiment of the invention, the detailed implementation steps for determining the macroblock refresh quota based on the current network congestion index and converting the traditional keyframe coding task into a distributed progressive refresh frame sequence include: The system reads the total number of macroblocks in the current image frame and determines the refresh rate of intra-frame coded macroblocks based on the current network congestion index, where the current network congestion index is negatively correlated with the refresh rate. Specifically, it obtains the current video resolution parameters from the encoder configuration information and calculates the total number of macroblocks based on the resolution and macroblock size (typically 16×16 pixels in the H.264 standard). For example, a 1920×1080 resolution video contains (1920 / 16)×(1080 / 16)=8160 macroblocks. The refresh rate is determined using an inverse function mapping of the congestion index; the more severe the congestion, the smaller the refresh rate, ensuring that the load per frame adapts to network capacity. The preset mapping relationship is: when the congestion index is in the range of 0.6-0.7, the refresh rate is 15%-20%; when it is in the range of 0.7-0.8, it is 10%-15%; and when it exceeds 0.8, it drops to 5%-10%. This negative correlation design embodies the adaptive principle of "becoming more conservative as the network gets worse," employing a more gradual refresh strategy during severe congestion to reduce the network burden.

[0022] The macroblock refresh quota is calculated based on the total number of macroblocks and the refresh rate. The macroblock refresh quota is the maximum number of intra-coded macroblocks allowed in a single transitional P-frame. The calculation uses a simple multiplication relationship: Macroblock refresh quota = Total number of macroblocks × Refresh rate, rounded down to obtain the integer number of macroblocks. For example, 8160 macroblocks have a quota of 1224 macroblocks at a 15% refresh rate. This quota value directly determines how many intra-coded macroblocks can be included in each transitional P-frame; the remaining macroblocks must be inter-coded. Since the data volume of intra-coded macroblocks is typically 5-10 times that of inter-coded macroblocks, strictly controlling the number of intra-coded macroblocks is an effective means of limiting the size of a single frame. After the quota calculation, a bitrate prediction verification is performed to ensure that the size of the P-frame encoded according to the quota does not exceed the maximum single-frame size corresponding to the target bitrate. If necessary, the quota value is further reduced.

[0023] The image frame is divided into unrefreshed macroblock regions and macroblock regions to be refreshed. For the macroblock regions to be refreshed, macroblocks in these regions are evenly distributed across N consecutive transition P-frames according to the macroblock refresh quota, where N is the rounded-up value of the ratio of the total number of macroblocks to the macroblock refresh quota. First, identify which macroblock regions in the current frame need to be refreshed due to lost or corrupted reference frames; these regions constitute the macroblock regions to be refreshed. The remaining regions that normally reference valid reference frames are unrefreshed macroblock regions. For the regions to be refreshed, calculate how many transition P-frames are needed to complete the refresh: N = ⌈Total number of macroblocks / Macroblock refresh quota⌉, rounded up to ensure that all macroblocks are eventually refreshed. Even distribution uses a round-robin scheduling strategy, allocating macroblocks in the regions to be refreshed sequentially to the N transition P-frames according to the scan order (usually raster scan), with the number of macroblocks allocated per frame not exceeding the quota limit. For example, 8160 macroblocks under a quota of 1224 require 7 transition P-frames, with approximately 1166 macroblocks allocated per frame. The allocation process also considers the spatial uniformity of macroblocks, trying to distribute the refreshed macroblocks of each frame throughout the screen and avoid concentrating them in a certain area, which would cause visual discontinuity.

[0024] Intra-frame predictive coding is used for macroblocks allocated in each transition P-frame, while inter-frame predictive coding is used for macroblocks not allocated, thus generating a distributed progressive refresh frame sequence. Specifically, when processing each transition P-frame, the encoder selectively switches the coding mode according to the macroblock allocation list. For macroblocks allocated to the current frame, the encoder forces the use of intra-frame predictive mode, predicting only by referring to neighboring macroblocks already encoded in the current frame to generate intra-coded macroblocks; for macroblocks not allocated to the current frame, the encoder uses the conventional inter-frame predictive mode, performing motion estimation and compensation by referring to the corresponding position in the previous frame to generate inter-coded macroblocks. Since intra-frame macroblocks can be decoded independently without depending on other frames, they gradually repair damaged areas in the image. After continuous transmission of N transition P-frames, all macroblocks to be refreshed have completed intra-frame coding refresh, the image quality is fully restored, and the size of each frame remains within a controllable range, avoiding congestion exacerbation caused by sudden large frames.

[0025] In this embodiment of the invention, after generating a distributed progressive refresh frame sequence by using intra-frame predictive coding for the macroblocks allocated in each transition P-frame and inter-frame predictive coding for the unallocated macroblocks, the method further includes: For each transition P-frame, the difference in the number of coded bits between its intra-frame coded macroblocks and inter-frame coded macroblocks is calculated. Specifically, after each transition P-frame is encoded, the total number of bits consumed by all intra-frame coded macroblocks and the total number of bits consumed by all inter-frame coded macroblocks in that frame are calculated, and the difference is obtained by subtracting the two. Since the compression efficiency of intra-frame coded macroblocks is lower than that of inter-frame coded macroblocks, the difference is usually positive and relatively large. This difference directly reflects the additional bit rate overhead introduced by the hybrid coding mode and is a key indicator for determining whether the size of a single frame may exceed the limit. A real-time coding feedback mechanism is adopted to continuously accumulate bit consumption during the encoding process, and the estimated difference can be obtained without waiting for the entire frame to be encoded, supporting early intervention.

[0026] If the difference in the number of encoded bits causes the total predicted bitrate of the transitional P-frame to exceed the target bitrate, a quantization parameter boosting operation is performed on the intra-frame coded macroblocks until the total predicted bitrate of the transitional P-frames is less than or equal to the target bitrate. The adjustment process first calculates the predicted bitrate for the frame based on the difference in the number of encoded bits and the frame rate, and determines whether it exceeds the target bitrate limit. When an exceedance is detected, a quantization parameter boosting operation is performed on the intra-frame coded macroblocks, i.e., the quantization parameter QP value is increased. A larger QP value results in coarser quantization, smaller encoded data volume, but greater distortion. An iterative adjustment strategy is used, gradually increasing the QP value of the intra-frame macroblocks (typically with a step size of 1-2), recalculating the predicted bitrate after each adjustment, until the predicted bitrate drops below the target bitrate. Quantization parameter boosting only applies to intra-frame coded macroblocks; the quantization parameters of inter-frame coded macroblocks maintain the original adjustment logic based on the bitrate downscaling instruction, ensuring the stability of video quality in unrefreshed areas. This selective quantization adjustment strategy, while meeting bitrate constraints, concentrates quality loss in the refresh area, which is already in a damaged state. Therefore, the quality degradation has a relatively small impact on the overall viewing experience, achieving an effective balance between bitrate control and quality maintenance.

[0027] In this embodiment of the invention, the detailed implementation steps of writing the distributed progressive refresh frame sequence into the transmission queue according to the time sequence, and having the transmission queue perform bitrate shaping output based on the target bitrate, so that the camera can repair image distortion while avoiding the transmission of sudden massive keyframes, include: Each transition P-frame in the distributed progressive refresh frame sequence is sequentially pushed into the transmission queue. The transmission queue is equipped with a token bucket shaper, and the token generation rate of the token bucket shaper is controlled by the target bitrate. The writing process follows the temporal order of the video frames, pushing the encoded transition P-frame data into the tail of the transmission queue. The queue uses a first-in, first-out (FIFO) strategy to manage frame data. The transmission queue integrates a token bucket shaper module, which maintains a virtual token pool. Tokens are generated at a constant rate, directly controlled by the target bitrate parameter, calculated as: Token generation rate (tokens / second) = Target bitrate (bits / second) / Token value (bits / token). Each token represents the right to send a certain number of bits of data; in a typical configuration, one token corresponds to 1000 bits. This ensures strict synchronization between the token generation rate and the target bitrate, achieving precise control of the transmission rate. The token bucket also has a maximum capacity limit, allowing a certain number of tokens to accumulate in a short period to cope with inter-frame volume fluctuations, but limiting the maximum burst traffic to prevent instantaneous peaks from exceeding network capacity.

[0028] When the frame data size of a transitional P-frame exceeds the current available tokens in the token bucket shaper, the transitional P-frame is split into multiple data packet fragments, which are then sent sequentially to the network link according to the token replenishment rate. First, the size of the transitional P-frame at the head of the queue is checked, and the number of tokens required to send the frame (frame size / token value) is calculated and compared with the current number of available tokens in the token bucket. When there are insufficient available tokens, the frame is not sent after token replenishment is completed; instead, it is split into multiple data packet fragments, each fragment's size matching the current number of available tokens. The splitting employs network layer fragmentation or application layer packetization techniques to ensure that fragments can be transmitted independently and correctly reassembled at the receiving end. Fragments are sent sequentially according to the available tokens: after sending a fragment, the corresponding tokens are consumed, and then the token bucket is replenished with new tokens before sending the next fragment. This rate-limiting sending mechanism ensures that the actual sending rate strictly adheres to the target bitrate setting, avoids burst transmission of frame data, smoothly distributes traffic across the timeline, effectively reduces instantaneous network load, and minimizes the risk of congestion.

[0029] The system monitors packet loss feedback during the transmission of the distributed progressive refresh frame sequence. If an intra-coded macroblock is lost within any transitional P-frame, the macroblock address of the lost intra-coded macroblock is added back to the refresh queue of subsequent transitional P-frames without triggering a new I-frame request response. Specifically, it continuously receives feedback messages from the decoding end, including ACK and NACK acknowledgments. When a NACK is received or no ACK is received within a timeout period, the corresponding data packet is determined to be lost. A macroblock mapping table is maintained for each transitional P-frame, recording which macroblocks use intra-coded encoding. When a data packet loss is detected, the system determines whether the lost data contains intra-coded macroblocks based on the packet sequence number and the mapping table. If so, the address information (macroblock index or spatial coordinates) of these lost macroblocks is extracted and added back to the refresh queue of subsequent transitional P-frames. During the encoding process of subsequent frames, these macroblocks will be preferentially assigned intra-coded modes to compensate for lost content. The key advantage of this re-flush mechanism is that it does not trigger the traditional I-frame request-response process, avoiding the exacerbation effect of sending a complete I-frame during congestion, and maintaining the consistency and effectiveness of the distributed progressive refresh strategy.

[0030] In this embodiment of the invention, after acquiring the network status feedback data from the camera and calculating the current network congestion index based on the network status feedback data, the method further includes: When the current network congestion index is less than or equal to a preset second congestion threshold, the system monitors whether an I-frame request command is received. The second congestion threshold is less than the first congestion threshold. Specifically, after calculating the current network congestion index, it is first compared with the second congestion threshold (typically 0.3-0.4) to determine if the network is in a good state. The second congestion threshold is significantly lower than the first congestion threshold (0.6-0.7), forming an intermediate region corresponding to three state levels: "good network," "normal network," and "congested network." When the congestion index is less than or equal to the second congestion threshold, it indicates a good network state with the ability to transmit larger data frames. In this state, I-frame request commands are still monitored, but the processing strategy is completely different from that in the congestion state. This dual-threshold design enables adaptive switching of the control strategy, employing the most suitable I-frame response method under different network conditions, balancing repair speed and network load.

[0031] If an I-frame request command is received and the current network congestion index is less than or equal to the second congestion threshold, the ratio of the current available bandwidth to the average volume of historical I-frames is calculated to obtain the I-frame carrying redundancy. The calculation process first obtains an estimate of the current available bandwidth, which can be obtained from a bandwidth estimation report fed back by the receiver or from the probe algorithm of the transmitter. Simultaneously, it queries the historical database to statistically analyze the average volume of I-frames generated over a past period (usually the most recent 50-100 frames). The I-frame carrying redundancy is calculated by dividing the two: Redundancy = Current Available Bandwidth / Average Volume of Historical I-Frames. This ratio reflects the network's carrying margin for I-frames; a larger ratio indicates more network capacity and greater safety in sending I-frames. Typical redundancy values ​​are between 1.5 and 3.0; less than 1 indicates insufficient bandwidth to carry a typical I-frame, while greater than 3 indicates very ample bandwidth.

[0032] When the redundancy of an I-frame exceeds a preset safety factor, traditional keyframes are allowed to be generated in response to I-frame request commands. The quantization parameters of these traditional keyframes are limited to ensure their actual encoded size does not exceed the maximum single-frame transmission threshold calculated based on the currently available bandwidth. First, the calculated redundancy is compared with a preset safety factor (typically 1.5-2.0). When the redundancy exceeds the safety factor, the network is deemed capable of safely sending I-frames, allowing the encoder to generate a complete traditional IDR keyframe. However, to prevent short-term congestion caused by I-frames exceeding expectations, restrictions are imposed on the keyframe's encoding parameters. First, the maximum single-frame transmission threshold is calculated based on the currently available bandwidth and frame rate. This threshold defines the maximum amount of data allowed per frame without causing congestion. Before generating an I-frame, the encoder inputs this threshold to the rate control module. The rate control module iteratively adjusts the quantization parameter QP to ensure the predicted encoded size of the I-frame does not exceed this threshold. The adjustment process employs a rate-distortion optimization algorithm to maintain I-frame quality as much as possible while meeting volume constraints. This restricted I-frame generation strategy utilizes the rapid repair capability under good network conditions (compared to distributed refresh, a complete I-frame can repair the entire screen at once), while avoiding the impact of excessively large I-frames on the network through size control, thus achieving an optimal balance between repair efficiency and network security.

[0033] In this embodiment of the invention, when the redundancy of an I-frame is greater than a preset safety factor, it is permissible to generate a traditional keyframe to respond to an I-frame request command, and the quantization parameters of the traditional keyframe are limited so that the actual encoded volume of the traditional keyframe does not exceed the maximum transmission threshold of a single frame calculated based on the currently available bandwidth. Detailed implementation steps include: The system obtains the smoothed available bandwidth of the current network link and calculates the average bandwidth margin per frame based on the frame rate. The acquisition process reads the smoothed available bandwidth value from the network status monitoring module. Smoothing typically employs an Exponentially Weighted Moving Average (EWMA) algorithm to reduce the impact of instantaneous fluctuations and obtain a stable bandwidth estimate. The average bandwidth margin per frame is calculated by dividing the smoothed available bandwidth by the frame rate: Average bandwidth margin per frame = Smoothed available bandwidth / Frame rate. For example, with an available bandwidth of 2 Mbps and a frame rate of 25 fps, the average margin per frame is 80 Kbit. This margin represents the bandwidth resources available per frame under uniform distribution and serves as a benchmark for setting single-frame size limits.

[0034] The maximum transmission threshold for a single frame is determined by multiplying the average bandwidth margin per frame by a preset multiplier, where the preset multiplier is dynamically adjusted based on the current network congestion index. Specifically, the maximum transmission threshold for a single frame is obtained by multiplying the average bandwidth margin per frame by the preset multiplier. The preset multiplier reflects the tolerance for instantaneous bandwidth bursts. This multiplier is not a fixed value but is dynamically adjusted based on the current network congestion index: when the congestion index is extremely low (close to 0), the multiplier can be set to 2.5-3.0, allowing for larger bursts; when the congestion index is close to the second congestion threshold, the multiplier drops to 1.5-2.0, limiting the burst intensity. This dynamic multiplier strategy allows the maximum transmission threshold for a single frame to flexibly adapt to subtle changes in the network, granting greater transmission freedom when network conditions are better and being more conservative when conditions are poor. After the maximum transmission threshold for a single frame is calculated, it serves as a hard constraint for I-frame encoding; the size of any I-frame must not exceed this threshold.

[0035] Before the encoder encodes traditional keyframes, the maximum transmission threshold for a single frame is input to the rate control module. The rate control module iteratively adjusts the quantization parameters of the traditional keyframes until the predicted coding volume of the traditional keyframes is less than or equal to the maximum transmission threshold for a single frame. During the control process, when the encoder is preparing to generate an I-frame, the maximum transmission threshold for a single frame is first passed to the rate control module as a target constraint. The rate control module employs an iterative method of predictive coding and parameter adjustment: initially, a default QP value is used to test-code some macroblocks, and the volume of the complete I-frame is predicted based on the test coding results; if the predicted volume exceeds the threshold, the QP value is increased (reducing quality but decreasing volume), and test coding and prediction are repeated; if the predicted volume is significantly lower than the threshold and there is room for quality improvement, the QP value is appropriately decreased; the iterative process continues until the predicted volume converges to a range slightly below the threshold (usually retaining a 5-10% safety margin). After iteration, the finally determined QP value is used to formally encode the complete I-frame, ensuring that the generated I-frame volume strictly meets the constraints. This iterative control method based on prediction and feedback maximizes I-frame quality while satisfying volume constraints, achieving an optimized balance between constraint compliance and coding efficiency.

[0036] After traditional keyframe encoding is completed, the keyframe is encapsulated into a Network Abstraction Layer (NAL) unit, and a high-priority identifier is assigned to each NAL unit. This allows the network transport layer to prioritize discarding P-frame packets with non-high-priority identifiers when queue overflow occurs. The encapsulation process encapsulates the encoded I-frame data into NALU format according to the H.264 / H.265 standard. The NALU header contains metadata such as frame type and timestamp. A high-priority identifier is set in the NALU header or the associated RTP header; this identifier is identifiable at both the network transport layer and the receiving end. When the transmission queue becomes congested due to sudden traffic spikes or network fluctuations, facing the risk of buffer overflow, the priority identifiers of each packet are checked. P-frame packets marked as low priority or without identifiers are discarded first, while high-priority I-frame packets are retained. This differentiated discarding strategy is based on the difference in importance between I-frames and P-frames: I-frames are the baseline for decoding, and their loss can lead to long-term error propagation, while the impact of losing a single P-frame is relatively localized and short-lived. Through this priority protection mechanism, even in extreme cases, the successful transmission of critical repair frames can be ensured, maximizing the image restoration effect.

[0037] In this embodiment of the invention, after dividing the image frame into an unrefreshed macroblock region and a macroblock region to be refreshed, and for the macroblock region to be refreshed, uniformly distributing the macroblocks in the macroblock region to be refreshed to N consecutive transitional P frames according to the macroblock refresh quota, the method further includes: During the generation of the distributed progressive refresh frame sequence, if a decrease in the current network congestion index is detected and a bitrate increase command is triggered, the macroblock refresh quota for subsequent transitional P-frames is dynamically increased based on the increased available bandwidth. Specifically, during progressive refresh execution, changes in the network congestion index are continuously monitored. When a sustained decrease in the congestion index is detected, falling below a certain improvement threshold, the network condition is deemed to have improved, triggering a bitrate increase command. The bitrate increase frees up additional bandwidth, which can be used to increase the refresh load per frame. The refresh ratio corresponding to the increased available bandwidth is recalculated, and this ratio is higher than the original setting. The increased macroblock refresh quota is calculated based on the new refresh ratio, and this quota is applied to subsequent transitional P-frames that have not yet been encoded. For example, if the original quota was 1000 macroblocks / frame, it increases to 1500 macroblocks / frame after network improvement, allowing subsequent frames to contain more intra-coded macroblocks and accelerating the refresh rate.

[0038] The remaining frame count N of the distributed progressive refresh frame sequence is shortened based on the increased macroblock refresh quota to accelerate the image distortion repair process. The shortening process recalculates the number of frames required to complete the remaining macroblocks to be refreshed: new remaining frame count = ⌈remaining macroblocks to be refreshed / increased macroblock refresh quota⌉. Because the denominator increases, the new remaining frame count decreases, meaning the entire refresh task can be completed in fewer frames. For example, if the original plan required 10 frames, the increased quota only requires 7 frames, saving 3 frames, which is equivalent to a 120-millisecond reduction in repair time at 25fps. The refresh scheduling plan is updated, redistributing the remaining macroblocks to be refreshed into the shortened frame count to ensure even distribution. This dynamic adjustment mechanism fully utilizes the resource increments brought about by improved network conditions, maximizing repair speed while ensuring network security.

[0039] After all macroblock regions to be refreshed in the distributed progressive refresh frame sequence have completed intra-frame encoding, a normal P-frame encoded stream is generated, and the I-frame congestion trap state is exited. Specifically, the refresh progress is continuously tracked, and the number of macroblocks that have completed intra-frame encoding is counted. When this number reaches the total number of macroblocks to be refreshed, the refresh task is considered complete. The "congestion trap state" flag is cleared in the global status register, notifying all modules to resume normal operation. The encoder releases the suspension of the I-frame encoding thread, restoring its responsiveness to handle possible future normal I-frame requests. Subsequent frames use the conventional P-frame encoding mode, and all macroblocks autonomously choose intra-frame or inter-frame prediction based on motion estimation results, no longer subject to the mandatory constraints of refresh quotas. Bitrate adjustment instructions continue to be executed, dynamically adjusting encoding parameters according to network status. The exit from the state signifies the successful resolution of the I-frame congestion trap, completing image restoration during network congestion, avoiding the vicious cycle in traditional solutions, and restoring stable video transmission.

[0040] In this embodiment of the invention, after the sending queue performs bitrate shaping and output based on the target bitrate, the method further includes: The system collects decoding status feedback information from the receiving end for the distributed progressive refresh frame sequence. This feedback includes a screen distortion repair confirmation flag or coordinates of remaining distorted areas. The collection process receives status feedback messages from the decoding end via the reverse control channel, using a dedicated protocol or RTCP extended format. After receiving and decoding the distributed progressive refresh frame sequence, the decoding end detects and analyzes the screen quality. If all damaged areas are repaired and the screen returns to normal, a screen distortion repair confirmation flag is sent. If unrepaired distorted or erroneous areas remain, their spatial coordinates (usually represented by macroblock addresses or pixel coordinates) are extracted, encapsulated as remaining distorted area coordinate information, and sent to the encoding end. The received feedback information is parsed and classified to distinguish between complete and partial repair, providing a basis for subsequent processing.

[0041] When the decoding status feedback includes coordinates of residual distorted areas, the set of macroblocks corresponding to these coordinates is extracted, and this set is remarked as forced-refresh macroblocks. First, the coordinates of the residual distorted areas in the feedback message are parsed, and the corresponding macroblock index set is calculated based on the mapping relationship between coordinates and macroblocks. It is then checked whether these macroblocks have already undergone intra-frame coding during the previous progressive refresh process, and the reasons for unsuccessful repair are analyzed (this could be due to packet loss in the macroblock or an incorrect reference frame selection at the decoding end). These macroblocks are remarked as forced-refresh macroblocks and given a higher refresh priority to ensure priority allocation of intra-frame coding resources in subsequent processing. The forced-refresh marking ensures that these macroblocks maintain intra-frame coding mode during the normal coding process until a successful repair confirmation is received.

[0042] A forced refresh macroblock is inserted into the intra-frame coding macroblock list of the currently encoded P-frame, and its bitrate overhead is included in the total bitrate constraint of the current P-frame during encoding output, until a picture distortion repair confirmation flag is received. Specifically, when processing subsequent P-frames, the existence of a forced refresh macroblock is first checked; if it exists, it is added to the intra-frame coding macroblock list of the current frame. A resource coordination mechanism is used to ensure that the encoding overhead of the forced refresh macroblock is included in the total bitrate constraint of the current frame. If necessary, the encoding parameters of other macroblocks are adjusted or the quality of the forced refresh macroblock itself is appropriately reduced to ensure that the overall frame size does not exceed the limit. The encoding output of the forced refresh macroblock adopts a reliable transmission strategy, and a retransmission mechanism or forward error correction protection can be set for it to improve the probability of successful transmission. The compensation encoding process continues, and after each transmission, feedback is awaited from the receiver to check whether a picture distortion repair confirmation flag has been received. When a confirmation flag is finally received, all forced refresh flags are cleared, and the fully normal encoding mode is restored. This feedback-based continuous compensation mechanism forms a complete repair closed loop, ensuring that the picture quality can be fully restored in the end.

[0043] This invention achieves intelligent adaptive bitrate adjustment and image distortion repair for camera video transmission under network congestion conditions through real-time network status monitoring, congestion index calculation, I-frame congestion trap identification, distributed progressive refresh conversion, and bitrate shaping output. The core innovation of this invention lies in identifying and avoiding the vicious cycle of I-frame congestion collapse traps. By converting sudden large I-frames into a distributed progressive refresh P-frame sequence, image damage is gradually repaired while strictly adhering to bitrate constraints, avoiding the avalanche effect of "congestion-packet loss-I-frame request-more severe congestion" in traditional solutions. The adaptive mechanism of this invention can dynamically adjust the refresh strategy according to network status. It allows for rapid full I-frame repair when the network is good, adopts a conservative progressive refresh when the network is congested, and accelerates the refresh process when the network improves, achieving an optimal balance between repair speed, video quality, and network load. A compensation mechanism based on decoding feedback further ensures the integrity and reliability of the repair, forming a complete control closed loop.

[0044] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

[0045] It should be noted that all formulas in this manual are calculated by removing dimensions and taking their numerical values. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters and thresholds in the formulas are set by those skilled in the art according to the actual situation.

[0046] Although embodiments of the invention have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the claims and their equivalents.

Claims

1. A method for adaptive adjustment of video bitrate based on network quality awareness for a camera, characterized in that, include: The network status feedback data of the camera is acquired, the current network congestion index is calculated based on the network status feedback data, and a bitrate reduction instruction is generated based on the current network congestion index. The network status feedback data includes round-trip time, packet loss rate, and jitter variance. During the execution of the bitrate reduction instruction, it is monitored whether an I-frame request instruction for requesting keyframes is received. The I-frame request instruction is generated by the decoding end when it detects image distortion or bitstream error. If the I-frame request instruction is received during the execution of the bitrate reduction instruction, it is identified that the current state is an I-frame congestion collapse trap, and the traditional keyframe encoding task generated based on the I-frame request instruction is intercepted. The macroblock refresh quota is determined based on the current network congestion index, and the traditional keyframe coding task is converted into a distributed progressive refresh frame sequence. The distributed progressive refresh frame sequence includes at least one transition P-frame. Each transition P-frame contains an intra-coded macroblock allocated based on the macroblock refresh quota, and the total data volume of each transition P-frame is constrained by the target bitrate corresponding to the bitrate reduction instruction. The distributed progressive refresh frame sequence is written into the transmission queue in a time sequence, and the transmission queue performs bitrate shaping output based on the target bitrate, so that the camera can repair image distortion while avoiding the transmission of a sudden large number of keyframes.

2. The method according to claim 1, characterized in that, The steps of acquiring network status feedback data from the camera, calculating the current network congestion index based on the network status feedback data, and generating a bitrate reduction instruction based on the current network congestion index include: The round-trip delay gradient change value, cumulative packet loss rate and jitter variance of the network transport layer are obtained according to the preset detection period. The round-trip delay gradient change value, the cumulative packet loss rate, and the jitter variance are normalized respectively, and the current network congestion index is calculated based on the preset congestion weight coefficient. The current network congestion index is compared with a preset first congestion threshold. When the current network congestion index is greater than the first congestion threshold, the target reduction ratio is calculated based on the difference between the current network congestion index and the first congestion threshold. The bitrate reduction instruction is generated based on the target reduction ratio, wherein the bitrate reduction instruction carries the target bitrate and the bitrate transition period.

3. The method according to claim 1, characterized in that, If the I-frame request instruction is received during the execution of the bitrate downscaling instruction, the system identifies that it is currently in an I-frame congestion collapse trap state and intercepts the traditional keyframe encoding task generated based on the I-frame request instruction, including: Upon receiving the I-frame request instruction, obtain the generation timestamp of the I-frame request instruction and query the execution progress of the bitrate reduction instruction at the current moment; If the execution progress indicates that the current bitrate has not yet reached the target bitrate, or the current moment is within the bitrate transition period, then it is determined that the I-frame request instruction and the bitrate reduction instruction are concurrent; When the I-frame request instruction and the bitrate reduction instruction are concurrent, the current state is marked as the I-frame congestion crash trap state, and the encoding thread in the encoder used to generate the complete real-time decoded refresh frame is suspended, blocking the packetization operation of the complete real-time decoded refresh frame.

4. The method according to claim 1, characterized in that, The step of determining the macroblock refresh quota based on the current network congestion index and converting the traditional keyframe coding task into a distributed progressive refresh frame sequence includes: Read the total number of macroblocks in the current image frame and determine the refresh ratio of intra-frame coded macroblocks based on the current network congestion index; The macroblock refresh quota is calculated based on the total number of macroblocks and the refresh ratio. The macroblock refresh quota is the maximum number of intra-coded macroblocks that a single frame transition P-frame is allowed to contain. The image frame is divided into an unrefreshed macroblock region and a macroblock region to be refreshed. For the macroblock region to be refreshed, the macroblocks in the macroblock region to be refreshed are evenly distributed to N consecutive transition P frames according to the macroblock refresh quota, where N is a value obtained by rounding up the ratio of the total number of macroblocks to the macroblock refresh quota. Intra-frame predictive coding is applied to the macroblocks allocated in each of the transition P frames, and inter-frame predictive coding is applied to the macroblocks not allocated, thereby generating the distributed progressive refresh frame sequence.

5. The method according to claim 4, characterized in that, After generating the distributed progressive refresh frame sequence by applying intra-frame predictive coding to the macroblocks allocated in each of the transition P frames and inter-frame predictive coding to the unallocated macroblocks, the method further includes: For each of the aforementioned transition P-frames, calculate the difference in the number of coded bits between its intra-frame coded macroblocks and inter-frame coded macroblocks; If the difference in the number of encoded bits causes the total predicted bit rate of the transition P-frame to exceed the target bit rate, then a quantization parameter lifting operation is performed on the intra-coded macroblock until the total predicted bit rate of the transition P-frame is less than or equal to the target bit rate. The quantization parameter lifting operation only applies to the intra-frame coded macroblocks, while the quantization parameters of the inter-frame coded macroblocks remain unchanged based on the original adjustment logic of the bitrate downsampling instruction.

6. The method according to claim 1, characterized in that, The step of writing the distributed progressive refresh frame sequence into the transmission queue in a time sequence, and having the transmission queue perform bitrate shaping output based on the target bitrate, so that the camera can repair image distortion while avoiding the transmission of sudden massive amounts of keyframes, includes: Each transition P-frame in the distributed progressive refresh frame sequence is sequentially pushed into the transmission queue, which is equipped with a token bucket shaper, and the token generation rate of the token bucket shaper is controlled by the target bit rate. When the frame data volume of the transition P-frame exceeds the current number of available tokens in the token bucket shaper, the transition P-frame is split into multiple data packet fragments, and the data packet fragments are sent to the network link in sequence according to the token replenishment rate. The system monitors packet loss feedback during the transmission of the distributed progressive refresh frame sequence. If an intra-coded macroblock is lost in any transition P-frame, the macroblock address of the lost intra-coded macroblock is added back to the refresh queue of the subsequent transition P-frame without triggering a new I-frame request response.

7. The method according to claim 1, characterized in that, After acquiring the network status feedback data from the camera and calculating the current network congestion index based on the network status feedback data, the method further includes: When the current network congestion index is less than or equal to a preset second congestion threshold, monitor whether the I-frame request instruction is received, wherein the second congestion threshold is less than the first congestion threshold; If the I-frame request instruction is received and the current network congestion index is less than or equal to the second congestion threshold, the ratio of the current available bandwidth to the average volume of historical I-frames is calculated to obtain the I-frame bearer redundancy. When the redundancy of the I-frame is greater than the preset safety factor, a traditional keyframe is allowed to be generated in response to the I-frame request instruction, and the quantization parameters of the traditional keyframe are restricted so that the actual encoding volume of the traditional keyframe does not exceed the maximum transmission threshold of a single frame calculated based on the currently available bandwidth.

8. The method according to claim 7, characterized in that, When the redundancy of the I-frame is greater than a preset safety factor, the generation of a traditional keyframe is permitted to respond to the I-frame request command, and the quantization parameters of the traditional keyframe are limited so that the actual encoded volume of the traditional keyframe does not exceed the maximum transmission threshold per frame calculated based on the currently available bandwidth, including: Obtain the smooth available bandwidth of the current network link and calculate the average bandwidth margin per frame based on the frame rate; The maximum transmission threshold for a single frame is determined by multiplying the average bandwidth margin of the single frame by a preset multiplier, wherein the preset multiplier is dynamically adjusted based on the current network congestion index. Before the encoder encodes the traditional keyframe, the maximum transmission threshold of a single frame is input to the bitrate control module, which iteratively adjusts the quantization parameters of the traditional keyframe until the predicted coding volume of the traditional keyframe is less than or equal to the maximum transmission threshold of a single frame. After the traditional keyframe is encoded, the traditional keyframe is encapsulated into a network abstraction layer unit, and the network abstraction layer unit is marked with a high priority identifier.

9. The method according to claim 4, characterized in that, After dividing the image frame into an unrefreshed macroblock region and a macroblock region to be refreshed, and for the macroblock region to be refreshed, evenly distributing the macroblocks in the macroblock region to be refreshed into N consecutive transitional P frames according to the macroblock refresh quota, the method further includes: During the generation of the distributed progressive refresh frame sequence, if the current network congestion index is detected to be decreasing and a bitrate increase instruction is triggered, the macroblock refresh quota of the subsequent transition P frames is dynamically increased based on the increased available bandwidth. The remaining number of frames N in the distributed progressive refresh frame sequence is shortened based on the increased macroblock refresh quota in order to accelerate the process of repairing image distortion. After all macroblock regions to be refreshed in the distributed progressive refresh frame sequence have completed intra-frame coding, a normal P-frame encoded stream is generated, and the I-frame congestion collapse trap state is exited.

10. The method according to claim 1, characterized in that, After the transmission queue performs bitrate shaping output based on the target bitrate, the method further includes: Collect decoding status feedback information from the receiving end for the distributed progressive refresh frame sequence. The decoding status feedback information includes a screen distortion repair confirmation flag or coordinates of the remaining screen distortion area. When the decoding status feedback information contains the coordinates of the remaining screen distortion area, the macroblock set corresponding to the coordinates of the remaining screen distortion area is extracted, and the macroblock set is re-marked as a forced refresh macroblock; The forced refresh macroblock is inserted into the intra-frame coding macroblock list of the currently encoded P-frame, and encoded and output within the bitrate overhead of the forced refresh macroblock, which is included in the total bitrate constraint of the current P-frame, until the image distortion repair confirmation flag is received.