Traffic optimization system and method for cloud edge cooperative moving target detection task based on adaptive coding

By adopting an adaptive encoding cloud-edge collaborative traffic optimization system in mobile object detection tasks, the video encoding parameters are dynamically adjusted, and the delay and quality problems of real-time video transmission in bandwidth-constrained environments are solved, and the accuracy and real-time detection are improved.

CN120017839AActive Publication Date: 2025-05-16NANJING UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510053263.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-16
Estimated Expiration
2045-01-14

AI Technical Summary

Technical Problem

In an environment with limited bandwidth, high network latency and frequent network fluctuations, real-time video transmission faces problems of high latency, low video quality and high packet loss rate, which affects the accuracy of mobile target detection and the system's real-time response capabilities.

Method used

The traffic optimization system based on adaptive encoding is adopted for cloud-edge collaborative mobile object detection tasks. Through modules such as video frame acquisition, network state perception, foreground detection, integration of areas of interest and quantitative parameter adjustment, video encoding parameters are dynamically adjusted, and high-quality encoding of areas of interest is given priority to reduce the encoding quality of background areas.

Benefits of technology

It effectively reduces cloud-edge collaborative traffic, while maintaining the quality of mobile target areas, improving the accuracy of target detection and real-time system, and reducing the processing pressure on the edge computing end and the cloud.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120017839A_ABST
    Figure CN120017839A_ABST
Patent Text Reader

Abstract

The invention discloses a flow optimization system and method for a cloud edge cooperative moving target detection task based on adaptive coding. The method comprises the following steps: collecting video frames; measuring a network state; identifying a region of interest through foreground detection; collecting historical reasoning results to obtain a current data packet region of interest; integrating the regions of interest; calculating a global quantization parameter and an area-of-interest quantization parameter offset according to the video frame content complexity, the area-of-interest content complexity and the current network state; and encoding and transmitting the data packet. According to the method, the quality of the target area can be effectively maintained while the cloud edge collaborative traffic is reduced, so that the accuracy of target detection is improved. In addition, according to the method, the calculation burden is effectively relieved, and higher performance is provided in real-time video processing scenes such as edge calculation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of edge computing real-time video stream analysis, and specifically relates to a traffic optimization system and method for cloud-edge collaborative mobile target detection tasks based on adaptive coding. Background Art

[0002] With the rapid development of the Internet of Things (IoT), video surveillance, and edge computing, real-time video transmission plays a vital role in key areas such as autonomous driving, security monitoring, and telemedicine. In mobile target detection tasks, such as identifying pedestrians and vehicles in dynamic traffic scenes or analyzing abnormal activities in security scenes, real-time transmitted video data needs to take into account both high quality and low latency to ensure the accuracy of detection results and the real-time response of the system. However, in an environment with limited bandwidth, high network latency, and frequent network fluctuations, real-time video transmission faces challenges such as high latency, low video quality, and high packet loss rate, which directly affect the accuracy of target detection and the real-time response capability of the system.

[0003] In this context, how to optimize real-time video traffic, reduce the burden of network transmission, and ensure the accuracy and timeliness of mobile target detection has become an urgent problem to be solved. As the core means to solve the traffic optimization problem, video coding technology, especially adaptive coding strategies for bandwidth-constrained environments, has attracted much attention in recent years. However, the current mainstream video coding methods still have the following significant limitations:

[0004] (1) Traditional video encoding methods (such as H.264, H.265, and AV1) dynamically adjust the quantization parameter (QP) of the video frame to change the video compression ratio when the network bandwidth fluctuates to reduce transmission traffic; the traditional method is based on a global QP adjustment mechanism and uses a consistent quantization standard for all areas of the image frame. Although this method can effectively reduce bandwidth usage, it fails to distinguish the priorities of key areas and non-key areas, resulting in key details in the video such as license plates, faces, pedestrians and other areas of interest being over-compressed, thereby affecting the accuracy of subsequent target detection.

[0005] (2) Video encoding methods based on deep learning use deep learning models such as convolutional neural networks to model the semantic features of video content, identify areas of interest, and perform high-quality encoding on them, while using low-quality compression on other areas. This method has high computational complexity and requires running deep learning models on edge devices, which places high demands on the computing power and memory capacity of edge devices. Edge devices are usually resource-constrained, and running deep learning models will significantly increase the computing latency and energy consumption of the devices, making it difficult to meet the needs of scenarios with high real-time requirements such as autonomous driving and security monitoring. Summary of the invention

[0006] In view of the above-mentioned deficiencies in the prior art, the purpose of the present invention is to provide a traffic optimization system and method for cloud-edge collaborative mobile target detection tasks based on adaptive coding. While reducing the cloud-edge collaborative traffic, the present invention can effectively maintain the quality of the mobile target area, thereby improving the accuracy of target detection. In addition, the method of the present invention effectively reduces the computational burden by reducing the coding quality of the background area, and provides higher performance in real-time video processing scenarios such as edge computing.

[0007] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0008] The present invention discloses a traffic optimization system for cloud-edge collaborative mobile target detection tasks based on adaptive coding, which is applied to edge terminals and cloud terminals. The edge terminal includes: a video frame acquisition module, a network status perception module, a historical data packet reasoning result collection module, a foreground detection module, an area of ​​interest integration module, a quantization parameter adjustment module and a coding module; the cloud terminal includes a mobile target detection module;

[0009] The video frame acquisition module is used to acquire continuous frame data in the real-time video stream and pack the acquired video frames of batch size batch_size into a data packet;

[0010] The network status perception module is used to monitor the network bandwidth BW in real time. t , network delay D t and packet loss rate loss_rate t , measure the current network status S t ;

[0011] The historical data packet reasoning result collection module is used to obtain the reasoning results of the historical data packet from the cloud And the inference results are based on the motion vector of the historical target Perform calibration to obtain the set of regions of interest for each frame of the current data packet

[0012] The foreground detection module is used to perform foreground detection on the collected video frames, separate the foreground information, and identify the dynamically changing set of regions of interest.

[0013] The region of interest integration module is used to integrate all regions of interest and Integrate and divide into sets and collection

[0014] The quantization parameter adjustment module is used to calculate the video frame content complexity com tot and Collection and The content complexity of each area of ​​interest in com roi , combined with the current network status S t Calculate the global quantization parameter QP of the video encoding and the quantization parameter offset ΔQP of the region of interest;

[0015] The encoding module is used to perform low-quality encoding on non-interest areas in the video using a global quantization parameter QP, and to perform high-quality encoding on the interest area using QP+ΔQP;

[0016] The mobile target detection module is used to perform mobile target detection on the data transmitted by the edge end, package the detection results into data packets, and send the inference results of the historical data packets to the edge end.

[0017] A traffic optimization method for cloud-edge collaborative mobile target detection tasks based on adaptive coding of the present invention is based on the above system, and the steps are as follows:

[0018] 1) Collect continuous frame data from the real-time video stream, and pack the collected video frames with a batch size of batch_size into a data packet;

[0019] 2) Real-time monitoring of network bandwidth BW t , network delay D t and packet loss rate loss_rate t , measure the current network status S t ;

[0020] 3) Perform foreground detection on the video frame in step 1) to separate the foreground information and identify the dynamically changing set of regions of interest

[0021] 4) Obtain inference results of historical data packets from the cloud Motion vector pair reasoning results based on historical targets Calibrate to get the set of regions of interest for each frame of the current data packet

[0022] 5) For all regions of interest and Integrate and divide into sets and collection

[0023] 6) Calculate the video frame content complexity com tot and Collection and The content complexity of each area of ​​interest in com roi , combined with the current network status S tCalculate the global quantization parameter QP of the video encoding and the quantization parameter offset ΔQP of the region of interest;

[0024] 7) Encode the non-interesting areas in the video at low quality using the global quantization parameter QP. and Each region of interest in the image is encoded with high quality using QP+ΔQP, and the encoded video stream is transmitted to the cloud.

[0025] Furthermore, the step 1) specifically includes: collecting continuous frame data I from the real-time video stream and storing it in a buffer; setting a fixed batch size batch_size, and when batch_size video frames are collected, packaging the video frames with a batch size of batch_size into a data packet.

[0026] Furthermore, the step 2) specifically includes:

[0027] 21) Real-time monitoring of network bandwidth BW t , network delay D t and packet loss rate loss_rate t ;

[0028] 22) According to the network bandwidth BW t , network delay D t and packet loss rate loss_rate t Perform network status measurement, expressed as S t , Among them, α, β, and γ are weight coefficients, which are used to adjust the influence of each indicator on the network status.

[0029] Furthermore, the step 3) specifically includes:

[0030] 31) Model the video background and obtain the background model M, which is used to represent the pixel x at a given time t. t The probability distribution M(x) belongs to the background t ), M(x t )=p(x t |θ t ), where p(x t |θ t ) is the conditional probability, indicating that given the background parameter θ t Next, pixels x t Probability of belonging to the background;

[0031] 32) The pixels x in each video frame t Compare with the background model M to determine the pixel x t Whether it belongs to the background, if M(x t) represents a background model that has a high degree of fit with the current pixel value, then the pixel belongs to the background, that is, x t ∈B t , B t Represents the background area in the current frame; if M(x t ) represents a low fit with the current pixel value, then the pixel does not belong to the background, that is, x t ∈F t , F t Represents the current foreground area; the fit between the pixel and the background model is measured by the significance I t (x t ) to indicate: in is a metric function used to measure the pixel x t The difference between the background model M and the background model M is t (x t ) is greater than the preset threshold θ s When x t was judged as a prospect;

[0032] 33) For the current foreground area F t The dilation operation is used to remove noise points, and the erosion operation is used to fill the gaps in the foreground area. The foreground area after the dilation operation is The foreground area after the erosion operation is in is a structural element, which is used to determine the shape and size of the dilation and erosion operations; the contour of the foreground image after dilation and erosion operations is extracted to obtain the boundary of the foreground area, which is expressed as in Contours(·) is a contour extraction function used to extract the region boundary. After the contour is extracted, the boundary Screen and select valid target areas that meet certain conditions as the set of areas of interest Condition is the condition used to filter the valid target area. represents the set of regions of interest in the i-th frame obtained by foreground detection;

[0033] 34) Update the background model to adapt to changes in lighting and slow-moving objects; update the background model, the expression is: in is the update function, according to the pixel x t For the background model parameter θ tto update.

[0034] Furthermore, the step 4) specifically includes:

[0035] 41) Obtain the inference results of historical data packets completed by the target detection model inference on the cloud Contains the inference results of batch size batch_size frames, and the inference results of each frame The expression is as follows:

[0036]

[0037] Among them, i∈[1,batch_size], represents the i-th frame in the data packet, and m is the number of regions of interest in the i-th frame; Represents the set of all regions of interest in a frame, each region of interest includes the coordinates of the upper left corner of the bounding box Width of the region of interest high and confidence

[0038] 42) Set the jth bounding box in the i-th frame In the i+1th frame The location has changed from Move to The motion vector is in

[0039] 43) According to the motion vector and time offset, the bounding box in each historical frame is offset to align the inference result in the historical data packet to the current data packet to be encoded; the offset of each bounding box is calculated as follows:

[0040] For the jth bounding box of the i-th frame in the historical data packet The offset bounding box of the current frame Adjusted by the following formula: in is the coordinate of the upper left corner of the jth bounding box in the i-th frame in the inference result, is the motion vector component of the i-th frame, Δt is the time offset, which represents the difference between the frame time and the current time, are the width and height of the jth bounding box, respectively, and α is the decay rate; for the region of interest of batch_size frames in the historical data packet, the offset region of interest set is expressed as in represents the set of regions of interest in the i-th frame, M is the number of regions of interest in the i-th frame.

[0041] Furthermore, the step 5) specifically includes:

[0042] 51) A set of regions of interest obtained according to foreground detection in step 3) And the set of regions of interest obtained based on historical reasoning results in step 4) The union of the two is regarded as the candidate region, and all candidate regions are integrated and divided; specifically: the bounding box set of the i-th frame obtained by foreground detection is The bounding box set of the i-th frame obtained by motion vector calibration of historical inference results is Where N represents the number of foreground bounding boxes in the current frame, and M represents the number of bounding boxes in the historical reasoning results; each bounding box is composed of a D-dimensional feature vector, which includes the coordinates of the upper left corner of the bounding box (x, y), the width w and height h of the bounding box, and the confidence c;

[0043] 52) Calculate the bounding box intersection over union (IoU) for each pair of foreground detection bounding boxes And the bounding box obtained from the historical data packet results Calculate its IoU value to determine the degree of overlap, where

[0044] 53) Set an IoU threshold. When the IoU value is greater than the IoU threshold, and is the detection result of the same target. and Merge, expressed as Add to Area of ​​Interest Collection Where p is the counting number, Merge(·) is the region merging function, including taking the intersection and union of two regions; when the IoU value is less than the IoU threshold, and is the detection result of different targets, skip the current match; wait for each pair and After all the comparisons are completed, and Not added to the collection Add areas of interest to the collection middle.

[0045] Furthermore, the step 6) specifically includes:

[0046] 61) Calculate the video frame content complexity com tot, where the calculation formula for the video frame content complexity of the i-th frame is F i For data packets The i-th video frame; CPX(·) is the content complexity calculation function, which is measured by the texture complexity metric based on discrete cosine transform or the pixel change rate; the content complexity of the video frame with batch size batch_size is com tot ,

[0047] 62) According to the current network status S t and video frame content complexitycom tot , calculate and obtain the global quantization parameter QP, QP = f(S t ,com tot ), where f(·) represents a multivariate regression function, which is fitted by linear regression or neural network regression model;

[0048] 63) Calculate the set of regions of interest and Content complexity of com roi , Where batch_size is the batch size, m is a set The number of regions of interest in the i-th frame, n is the set The number of regions of interest in the i-th frame, cx k,l =CPX(Region k,l ), used to calculate the content complexity of the region of interest, Region k,l for In the collection Corresponding image area; according to the set of regions of interest and Content complexity of com roi and video frame content complexitycom tot , calculate the quantization parameter offset of the region of interest, motion is the motion change of the target in the region of interest; g(·) is a multivariable mapping function, using a weighted linear model or a neural network model. The implementation of the weighted linear model is: Among them, λ1,λ2,λ3,λ4 are weights, δ is the offset, The indicator function takes the value 1 when the condition is true, otherwise it takes the value 0. i,j is the confidence level of this region.

[0049] Furthermore, in step 7), the multimedia API is used to perform low-quality encoding on the non-interesting areas in the video with a global quantization parameter QP, and high-quality encoding on the interesting areas with QP+ΔQP.

[0050] Beneficial effects of the present invention:

[0051] 1. The present invention improves the accuracy and real-time performance of target detection under bandwidth-constrained conditions. Through an adaptive coding strategy, low quantization parameters are preferentially used for high-quality coding in areas of interest, effectively ensuring the accuracy of target detection in key areas. At the same time, the coding quality of non-interested areas is reduced, reducing the overall amount of coded data and ensuring real-time performance and detection effect under limited bandwidth.

[0052] 2. The dynamic adaptive mechanism of the present invention enhances the robustness and flexibility of the system, uses real-time bandwidth and network conditions to adjust coding parameters, and dynamically optimizes coding strategies in combination with prospect detection and historical reasoning results, so that the system can flexibly respond to network fluctuations and improve its adaptability to complex network environments.

[0053] 3. The present invention optimizes the collaborative traffic between the cloud and the edge under the task of mobile target detection, effectively reduces the data transmission traffic, and saves system resources; it can flexibly adjust the encoding strategy in different network environments. High quantization parameter encoding is used for non-interested areas, which significantly reduces the data traffic of video encoding and transmission while ensuring the accuracy of target detection, thereby reducing the processing pressure of the edge computing end and the cloud, and saving precious computing and communication resources for application scenarios such as video surveillance and intelligent transportation. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 It is a functional block diagram of the system of the present invention.

[0055] Figure 2 Flow chart of the method of the present invention.

[0056] Figure 3 Schematic diagram of the method of the present invention. DETAILED DESCRIPTION

[0057] In order to facilitate the understanding of those skilled in the art, the present invention is further described below in conjunction with embodiments and drawings. The contents mentioned in the implementation modes are not intended to limit the present invention.

[0058] Reference Figure 1As shown, a traffic optimization system for cloud-edge collaborative mobile target detection tasks based on adaptive coding of the present invention is applied to the edge and cloud ends, and the edge end includes: a video frame acquisition module, a network status perception module, a historical data packet reasoning result collection module, a foreground detection module, an area of ​​interest integration module, a quantization parameter adjustment module and an encoding module; the cloud end includes a mobile target detection module;

[0059] The video frame acquisition module is used to acquire continuous frame data in the real-time video stream and pack the acquired video frames of batch size batch_size into a data packet;

[0060] The network status perception module is used to monitor the network bandwidth BW in real time. t 、Network delay D t and packet loss rate loss_rate t , measure the current network status S t ;

[0061] The historical data packet reasoning result collection module is used to obtain the reasoning results of the historical data packet from the cloud And the inference results are based on the motion vector of the historical target Perform calibration to obtain the set of regions of interest for each frame of the current data packet

[0062] The foreground detection module is used to perform foreground detection on the collected video frames, separate the foreground information, and identify the dynamically changing set of regions of interest.

[0063] The region of interest integration module is used to integrate all regions of interest and Integrate and divide into sets and collection

[0064] The quantization parameter adjustment module is used to calculate the video frame content complexity com tot and Collection and The content complexity of each area of ​​interest in com roi , combined with the current network status S t Calculate the global quantization parameter QP of the video encoding and the quantization parameter offset ΔQP of the region of interest;

[0065] The encoding module is used to perform low-quality encoding on non-interest areas in the video using a global quantization parameter QP, and to perform high-quality encoding on the interest area using QP+ΔQP;

[0066] The mobile target detection module is used to perform mobile target detection on data transmitted by the edge end using a deep learning-based target detection algorithm, package the detection results into data packets, and send the inference results of the historical data packets to the edge end.

[0067] Reference Figure 2 , Figure 3 As shown, a traffic optimization method for cloud-edge collaborative mobile target detection tasks based on adaptive coding of the present invention is based on the above system, and the steps are as follows:

[0068] 1) Collect continuous frame data from the real-time video stream, and pack the collected video frames with a batch size of batch_size into a data packet; specifically, the following steps are performed: collect continuous frame data I from the real-time video stream and store it in a buffer; set a fixed batch size batch_size, and when batch_size video frames are collected, pack the video frames with a batch size of batch_size into a data packet, which is used in step 3) to use the input of foreground detection to reduce the collaborative traffic between the cloud and the edge. The video frames in each batch are represented as

[0069] 2) Real-time monitoring of network bandwidth BW t 、Network delay D t and packet loss rate loss_rate t , measure the current network status S t ; Specifically include:

[0070] 21) Real-time monitoring of network bandwidth BW t 、Network delay D t and packet loss rate loss_rate t ;

[0071] 22) According to the network bandwidth BW t 、Network delay D t and packet loss rate loss_rate t Perform network status measurement, expressed as S t , Among them, α, β, and γ are weight coefficients, which are used to adjust the influence of each indicator on the network status.

[0072] 3) Perform foreground detection on the video frame in step 1) to separate the foreground information and identify the dynamically changing set of regions of interest The step 3) specifically includes:

[0073] 31) Model the video background and obtain the background model M, which is used to represent the pixel x at a given time t. t The probability distribution M(x) belongs to the background t), M(x t )=p(x t |θ t ), where p(x t |θ t ) is the conditional probability, indicating that given the background parameter θ t Next, pixels x t Probability of belonging to the background;

[0074] 32) The pixels x in each video frame t Compare with the background model M to determine the pixel x t Whether it belongs to the background, if M(x t ) represents a background model that has a high degree of fit with the current pixel value, then the pixel belongs to the background, that is, x t ∈B t , B t Represents the background area in the current frame; if M(x t ) represents a low fit with the current pixel value, then the pixel does not belong to the background, that is, x t ∈F t , F t Represents the current foreground area; the fit between the pixel and the background model is measured by the significance I t (x t ) to indicate: in is a metric function used to measure the pixel x t The difference between the background model M and the background model M is t (x t ) is greater than the preset threshold θ S When x t was judged as a prospect;

[0075] 33) For the current foreground area F t The dilation operation is used to remove noise points, and the erosion operation is used to fill the gaps in the foreground area. The foreground area after the dilation operation is The foreground area after the erosion operation is in is a structural element, which is used to determine the shape and size of the dilation and erosion operations; the contour of the foreground image after dilation and erosion operations is extracted to obtain the boundary of the foreground area, which is expressed as in Contours(·) is a contour extraction function used to extract the region boundary. After the contour is extracted, the boundary Screen and select valid target areas that meet certain conditions (determine whether the contour belongs to a valid region of interest, usually including area threshold, shape, position, etc.) as the region of interest set Condition is the condition used to filter the valid target area. represents the set of regions of interest in the i-th frame obtained by foreground detection;

[0076] 34) Update the background model to adapt to changes in lighting and slow-moving objects; update the background model, the expression is: in is the update function, according to the pixel x t For the background model parameter θ t to update.

[0077] 4) Obtain inference results of historical data packets from the cloud Motion vector pair reasoning results based on historical targets Calibrate to get the set of regions of interest for each frame of the current data packet Specifically include:

[0078] 41) Obtain the inference results of historical data packets completed by the target detection model inference on the cloud Contains the inference results of batch size batch_size frames, and the inference results of each frame The expression is as follows:

[0079]

[0080] Among them, i∈[1,batch_size], represents the i-th frame in the data packet, and m is the number of regions of interest in the i-th frame; Represents the set of all regions of interest in a frame, each region of interest includes the coordinates of the upper left corner of the bounding box Width of the region of interest high and confidence

[0081] 42) Set the jth bounding box in the i-th frame In the i+1th frame The location has changed from Move to The motion vector is in

[0082] 43) According to the motion vector and time offset, the bounding box in each historical frame is offset to align the inference result in the historical data packet to the current data packet to be encoded; the offset of each bounding box is calculated as follows:

[0083] For the jth bounding box of the i-th frame in the historical data packet The current frame bounding box after offset Adjusted by the following formula: in is the coordinate of the upper left corner of the jth bounding box in the i-th frame in the inference result, is the motion vector component of the i-th frame, Δt is the time offset, which represents the difference between the frame time and the current time, are the width and height of the jth bounding box, respectively, and α is the decay rate; for the region of interest of batch_size frames in the historical data packet, the offset region of interest set is expressed as in represents the set of regions of interest in the i-th frame, M is the number of regions of interest in the i-th frame.

[0084] 5) For all regions of interest and Integrate and divide into sets and collection Specifically include:

[0085] 51) A set of regions of interest obtained according to foreground detection in step 3) And the set of regions of interest obtained based on historical reasoning results in step 4) The union of the two is regarded as the candidate region, and all candidate regions are integrated and divided; specifically: the bounding box set of the i-th frame obtained by foreground detection is The bounding box set of the i-th frame obtained by motion vector calibration of historical inference results is Where N represents the number of foreground bounding boxes in the current frame, and M represents the number of bounding boxes in the historical reasoning results; each bounding box is composed of a D-dimensional feature vector, which includes the coordinates of the upper left corner of the bounding box (x, y), the width w and height h of the bounding box, and the confidence c;

[0086] 52) Calculate the bounding box intersection over union (IoU) for each pair of foreground detection bounding boxes And the bounding box obtained from the historical data packet results Calculate its IoU value to determine the degree of overlap, where

[0087] 53) Set an IoU threshold. When the IoU value is greater than the IoU threshold, and is the detection result of the same target. and Merge, expressed as Add to Area of ​​Interest Collection Where p is the counting number, Merge(·) is the region merging function, including taking the intersection and union of two regions; when the IoU value is less than the IoU threshold, and is the detection result of different targets, skip the current match; wait for each pair and After all the comparisons are completed, and Not added to the collection Add areas of interest to the collection middle.

[0088] 6) Calculate the video frame content complexity com tot and Collection and The content complexity of each area of ​​interest in com roi , combined with the current network status S t Calculating the global quantization parameter QP of video coding and the quantization parameter offset ΔQP of the region of interest; specifically including:

[0089] 61) Calculate the video frame content complexity com tot , where the calculation formula for the video frame content complexity of the i-th frame is F i For data packets The i-th video frame; CPX(·) is the content complexity calculation function, which is measured by the texture complexity metric based on discrete cosine transform or the pixel change rate; the content complexity of the video frame with batch size batch_size is com tot ,

[0090] 62) According to the current network status S t and video frame content complexitycom tot , calculate and obtain the global quantization parameter QP, QP = f(S t ,com tot), where f(·) represents a multivariate regression function, which is fitted by linear regression or neural network regression model;

[0091] 63) Calculate the set of regions of interest and Content complexity of com roi , Where batch_size is the batch size, m is a set The number of regions of interest in the i-th frame, n is the set The number of regions of interest in the i-th frame, cx k,l =CPX(Region k,l ), used to calculate the content complexity of the region of interest, Region k,l for In the collection Corresponding image area; according to the set of regions of interest and Content complexity of com roi and video frame content complexitycom tot , calculate the quantization parameter offset of the region of interest, motion is the motion change of the target in the region of interest; g(·) is a multivariable mapping function, using a weighted linear model or a neural network model. The implementation of the weighted linear model is: Among them, λ1,λ2,λ3,λ4 are weights, δ is the offset, which can be obtained through regression training. The indicator function takes the value 1 when the condition is true, otherwise it takes the value 0. i,j is the confidence level of this region.

[0092] 7) Encode the non-interesting areas in the video at low quality using the global quantization parameter QP, and encode the interesting areas at high quality using QP+ΔQP, and transmit the encoded video stream to the cloud;

[0093] Among them, the multimedia API is used to encode the non-interesting areas in the video with low quality using the global quantization parameter QP, and the set and Each region of interest in the image is encoded with high quality at QP+ΔQP.

[0094] The present invention has many specific application paths. The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements can be made without departing from the principle of the present invention. These improvements should also be regarded as the protection scope of the present invention.

Claims

1. A traffic optimization system for cloud-edge collaborative mobile target detection tasks based on adaptive coding, characterized in that: Applied to the edge and cloud, the edge includes: video frame acquisition module, network status perception module, historical data packet reasoning result collection module, foreground detection module, region of interest integration module, quantization parameter adjustment module and encoding module; the cloud includes mobile target detection module; The video frame acquisition module is used to acquire continuous frame data in the real-time video stream and pack the acquired video frames of batch size batch_size into a data packet; The network status perception module is used to monitor the network bandwidth BW in real time. t 、Network delay D t and packet loss rate loss_rate t , measure the current network status S t ; The historical data packet reasoning result collection module is used to obtain the reasoning results of the historical data packet from the cloud And the inference results are based on the motion vector of the historical target Perform calibration to obtain the set of regions of interest for each frame of the current data packet The foreground detection module is used to perform foreground detection on the collected video frames, separate the foreground information, and identify the dynamically changing set of regions of interest. The region of interest integration module is used to integrate all regions of interest and Integrate and divide into sets and collection The quantization parameter adjustment module is used to calculate the video frame content complexity com tot and Collection and The content complexity of each area of ​​interest in com roi , combined with the current network status S t Calculate the global quantization parameter QP of the video encoding and the quantization parameter offset ΔQP of the region of interest; The encoding module is used to perform low-quality encoding on non-interest areas in the video using a global quantization parameter QP, and to perform high-quality encoding on the interest area using QP+ΔQP; The mobile target detection module is used to perform mobile target detection on the data transmitted by the edge end, package the detection results into data packets, and send the inference results of the historical data packets to the edge end.

2. A traffic optimization method for cloud-edge collaborative mobile target detection tasks based on adaptive coding, based on the system of claim 1, characterized in that: The steps are as follows: 1) Collect continuous frame data from the real-time video stream, and pack the collected video frames with a batch size of batch_size into a data packet; 2) Real-time monitoring of network bandwidth BW t 、Network delay D t and packet loss rate loss_rate t , measure the current network status S t ; 3) Perform foreground detection on the video frame in step 1) to separate the foreground information to identify the dynamically changing set of regions of interest 4) Obtain inference results of historical data packets from the cloud Motion vector pair reasoning results based on historical targets Calibrate to get the set of regions of interest for each frame of the current data packet 5) For all regions of interest and Integrate and divide into sets and collection 6) Calculate the video frame content complexity com tot and Collection and The content complexity of each area of ​​interest in com roi , combined with the current network status S t Calculate the global quantization parameter QP of the video encoding and the quantization parameter offset ΔQP of the region of interest; 7) Encode the non-interesting areas in the video at low quality using the global quantization parameter QP. and Each region of interest in the image is encoded with high quality using QP+ΔQP, and the encoded video stream is transmitted to the cloud.

3. The traffic optimization method for cloud-edge collaborative mobile target detection tasks based on adaptive coding according to claim 2 is characterized in that: The step 1) specifically includes: collecting continuous frame data I from the real-time video stream and storing it in a buffer; setting a fixed batch size batch_size, and when batch_size video frames are collected, packaging the video frames with a batch size of batch_size into a data packet.

4. The traffic optimization method for cloud-edge collaborative mobile target detection tasks based on adaptive coding according to claim 2 is characterized in that: The step 2) specifically includes: 21) Real-time monitoring of network bandwidth BW t , network delay D t and packet loss rate loss_rate t ; 22) According to the network bandwidth BW t , network delay D t and packet loss rate loss_rate t Perform network status measurement, expressed as S t , Among them, α, β, and γ are weight coefficients, which are used to adjust the influence of each indicator on the network status.

5. The traffic optimization method for cloud-edge collaborative mobile target detection tasks based on adaptive coding according to claim 2 is characterized in that: The step 3) specifically includes: 31) Model the video background and obtain the background model M, which is used to represent the pixel x at a given time t. t The probability distribution M(x) belongs to the background t ), M(x t )=p(x t |θ t ), where p(x t |θ t ) is the conditional probability, indicating that given the background parameter θ t Next, pixels x t Probability of belonging to the background; 32) The pixels x in each video frame t Compare with the background model M to determine the pixel x t Whether it belongs to the background, if M(x t ) represents a background model that has a high degree of fit with the current pixel value, then the pixel belongs to the background, that is, x t ∈B t , B t Represents the background area in the current frame; if M(x t ) represents a low fit with the current pixel value, then the pixel does not belong to the background, that is, x t ∈F t , F t Represents the current foreground area; the fit between the pixel and the background model is measured by the significance I t (x t ) to indicate: in is a metric function used to measure the pixel x t The difference between the background model M and the background model M is t (x t ) is greater than the preset threshold θ S When x t was judged as a prospect; 33) For the current foreground area F t The dilation operation is used to remove noise points, and the erosion operation is used to fill the gaps in the foreground area. The foreground area after the dilation operation is The foreground area after the erosion operation is in is a structural element, which is used to determine the shape and size of the dilation and erosion operations; the foreground image after dilation and erosion operations is contour extracted to obtain the boundary of the foreground area, which is expressed as in Contours(·) is a contour extraction function used to extract the region boundary. After the contour is extracted, the boundary Screen and select valid target areas that meet certain conditions as the set of areas of interest Condition is the condition used to filter the valid target area. represents the set of regions of interest in the i-th frame obtained by foreground detection; 34) Update the background model to adapt to changes in lighting and slow-moving objects; update the background model, the expression is: in is the update function, according to the pixel x t For the background model parameter θ t to update.

6. The traffic optimization method for cloud-edge collaborative mobile target detection tasks based on adaptive coding according to claim 2 is characterized in that: The step 4) specifically includes: 41) Obtain the inference results of historical data packets completed by the target detection model inference on the cloud Contains the inference results of batch size batch_size frames, and the inference results of each frame The expression is as follows: Among them, i∈[1,batch_size], represents the i-th frame in the data packet, and m is the number of regions of interest in the i-th frame; Represents the set of all regions of interest in a frame, each region of interest includes the coordinates of the upper left corner of the bounding box Width of the region of interest high and confidence 42) Set the jth bounding box in the i-th frame In the i+1th frame The location has changed from Move to The motion vector is in 43) According to the motion vector and time offset, the bounding box in each historical frame is offset to align the inference result in the historical data packet to the current data packet to be encoded; the offset of each bounding box is calculated as follows: For the jth bounding box of the i-th frame in the historical data packet The current frame bounding box after offset Adjusted by the following formula: in is the coordinate of the upper left corner of the jth bounding box in the i-th frame in the inference result, is the motion vector component of the i-th frame, Δt is the time offset, which represents the difference between the frame time and the current time, are the width and height of the jth bounding box, respectively, and α is the decay rate; for the region of interest of batch_size frames in the historical data packet, the offset region of interest set is expressed as in represents the set of regions of interest in the i-th frame, M is the number of regions of interest in the i-th frame.

7. The traffic optimization method for cloud-edge collaborative mobile target detection tasks based on adaptive coding according to claim 6 is characterized in that: The step 5) specifically includes: 51) A set of regions of interest obtained according to foreground detection in step 3) And the set of regions of interest obtained based on historical reasoning results in step 4) The union of the two is regarded as the candidate region, and all candidate regions are integrated and divided; specifically: the bounding box set of the i-th frame obtained by foreground detection is The bounding box set of the i-th frame obtained by motion vector calibration of historical inference results is Where N represents the number of foreground bounding boxes in the current frame, and M represents the number of bounding boxes in the historical reasoning results; each bounding box is composed of a D-dimensional feature vector, which includes the coordinates of the upper left corner of the bounding box (x, y), the width w and height h of the bounding box, and the confidence c; 52) Calculate the bounding box intersection over union (IoU) for each pair of foreground detection bounding boxes And the bounding box obtained from the historical data packet results Calculate its IoU value to determine the degree of overlap, where 53) Set an IoU threshold. When the IoU value is greater than the IoU threshold, and is the detection result of the same target. and Merge, expressed as Add to Area of ​​Interest Collection Where p is the counting number, Merge(·) is the region merging function, including taking the intersection and union of two regions; when the IoU value is less than the IoU threshold, and is the detection result of different targets, skip the current match; wait for each pair and After all the comparisons are completed, and Not added to the collection Add areas of interest to the collection middle.

8. The traffic optimization method for cloud-edge collaborative mobile target detection tasks based on adaptive coding according to claim 2 is characterized in that: The step 6) specifically includes: 61) Calculate the video frame content complexity com tot , where the calculation formula for the video frame content complexity of the i-th frame is F i For data packets The i-th video frame; CPX(·) is the content complexity calculation function, which is measured by the texture complexity metric based on discrete cosine transform or the pixel change rate; the content complexity of the video frame with batch size batch_size is com tot , 62) According to the current network status S t and video frame content complexitycom tot , calculate and obtain the global quantization parameter QP, QP = f(S t ,com tot ), where f(·) represents a multivariate regression function, which is fitted by linear regression or neural network regression model; 63) Calculate the set of regions of interest and Content complexity of com roi , Where batch_size is the batch size, m is a set The number of regions of interest in the i-th frame, n is the set The number of regions of interest in the i-th frame, cx k,l =CPX(Region k,l ), used to calculate the content complexity of the region of interest, Region k,l for In the collection Corresponding image area; according to the set of regions of interest and Content complexity of com roi and video frame content complexitycom tot , calculate the quantization parameter offset of the region of interest, motion is the motion change of the target in the region of interest; g(·) is a multivariable mapping function, using a weighted linear model or a neural network model. The implementation of the weighted linear model is: Among them, λ1,λ2,λ3,λ4 are weights, δ is the offset, The indicator function takes the value 1 when the condition is true, otherwise it takes the value 0. i,j is the confidence level of this region.

9. The traffic optimization method for cloud-edge collaborative mobile target detection tasks based on adaptive coding according to claim 2 is characterized in that: In the step 7), the multimedia API is used to perform low-quality encoding on the non-interesting areas in the video with the global quantization parameter QP, and high-quality encoding on the interesting areas with QP+ΔQP.

Citation Information

Patent Citations

  • Video coding method and device, equipment and storage medium

    CN111479112A

  • Video hierarchical coding and decoding system and method based on region perception

    CN119299775A

  • Foreground Analysis Based on Tracking Information

    US20120027248A1

  • Method, apparatus and system for decoding tensors for content from a bitstream

    WO2024239042A1