A cloud-edge collaborative video stream analysis method and apparatus based on depth estimation
By using a cloud-edge collaborative video stream analysis method, depth estimation is used to generate partitioning rules and perform different quality encodings. Combined with the aggregation of edge and cloud inference results, the problem of limited camera computing resources is solved, and low-latency and high-accuracy video analysis is achieved.
Patent Information
- Application Number
- CN202310477632.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-28
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2043-04-28
AI Technical Summary
Existing video stream analysis systems struggle to achieve the optimal balance between low latency and high accuracy when processing camera computing resources are limited, and the offloading method guided by server feedback information increases inference latency.
A high-complexity DNN model is deployed in the cloud, while a low-complexity DNN model is deployed at the edge. Video frame partitioning rules are generated through depth estimation. The acquisition end encodes the video frames according to the partitioning rules at different quality levels. The inference results are aggregated at the edge and in the cloud, and the results are tracked using an LSTM module.
It significantly reduces latency while maintaining analytical accuracy, and solves the problem of inconsistent ROI definitions among different DNN models, thus improving the system's flexibility.
Smart Images

Figure CN117079108B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of video analytics, specifically a cloud-edge collaborative video stream analysis method and apparatus based on depth estimation. Background Technology
[0002] With the widespread adoption of deep neural networks and cameras, video stream inference has found broad application scenarios (e.g., urban traffic analysis and security anomaly detection). In these scenarios, cameras continuously collect video information and stream it to a remote server. The remote server receives the information, runs a deep neural network (DNN) model to analyze it, and returns the analysis results to the camera. This is the prototype of a video stream analysis system. However, as these video stream analysis systems are gradually applied, various problems have begun to emerge. The first problem is excessive analysis latency. To reduce video analysis latency, many researchers have reduced the amount of data transmitted over the network by only offloading the Region of Interest (ROI), thereby reducing analysis latency. The focus of these studies is on ROI selection. Since the purpose of selecting an ROI is to reduce network transmission latency, in video stream analysis systems, the ROI selection module can only be performed on the video source, i.e., the camera. Therefore, researchers have proposed various heuristic algorithms to utilize the limited computing resources on the camera for RIO selection. Video stream analysis systems have flourished, such as Glimpse (Chen, YH, et al. "Glimpse: Continuous, Real-Time Object Recognition on Mobile Devices." the 13th ACM Conference ACM, 2015.) and Reducto (Li, Yuanqi, et al. "Reducto: On-camera filtering for resource-efficient real-time video analytics." Proceedings of the Annual conference of the ACM Special Interest Group on Data Communication on the applications, technologies, architectures, and protocols for computer communication. 2020.). These video stream analysis systems all offer low latency while maintaining analytical accuracy without significantly reducing it. However, as research has deepened, some researchers have pointed out that only a small fraction of cameras currently in use have high computing resources; the vast majority of cameras lack sufficient computing resources and are primarily used for video acquisition and transmission. Therefore, continuing to apply the aforementioned rules would severely limit the application scenarios of video stream analysis systems.
[0003] To address the limitations of camera computing resources, researchers have turned their attention to servers. These researchers believe there's another way to reduce network transmission latency: changing the video encoding method. Traditional video streaming analysis systems compress and encode video based on Quality of Experience (QoE) during streaming. However, in video streaming analysis systems, the server receives video for DNN models to use for inference, and doesn't need to consider whether the QoE has been reduced during transmission. For example, in object detection tasks, compressing or even cropping the background in a video frame doesn't affect the DNN model's recognition accuracy, but humans will perceive a significant decrease in video quality. Therefore, researchers began to adjust the quality of video transmitted over the network, either actively or passively, to balance latency and accuracy in video analysis. A series of video streaming analysis systems focusing on this approach have emerged. For example, AWStream proposes dynamically adjusting the encoding quality of the next video segment to cope with bandwidth fluctuations. (Zhang, B., et al. "AWStream: adaptive wide-area streaming analytics." ACM Special Interest Group on Data Communication ACM, 2018.) DDS divides the offloading process into two parts. The first part encodes the video with a high quality loss and transmits it to the server to receive feedback from the server. The second part, based on the server's feedback, performs lossless compression on only a portion of the video and transmits it to the server for the DNN model to perform inference. (Du, K., et al. "Server-Driven Video Streaming for Deep Learning Inference." SIGCOMM '20: Annual conference of the ACM Special Interest Group on Data Communication on the applications, technologies, architectures, and protocols for computer communication ACM, 2020.) These offloading methods that utilize server feedback improve accuracy, but at the same time, they greatly increase inference latency, failing to achieve the optimal trade-off between accuracy and latency. Summary of the Invention
[0004] The purpose of this invention is to overcome the problems existing in the prior art and provide a cloud-edge collaborative video stream analysis method and device based on depth estimation.
[0005] This invention adopts the following technical solution: a cloud-edge collaborative video stream analysis method based on depth estimation, comprising:
[0006] Deploy the Server DNN model in the cloud and the Edge DNN model at the edge;
[0007] In the cloud, video frame partitioning rules are generated based on the depth of the video frame;
[0008] The acquisition end encodes different blocks with different qualities according to the cloud partitioning rules and then transmits them to the edge end;
[0009] The edge device segments the received video into regions, transmits the high-quality encoded blocks to the cloud server for inference, and simultaneously performs inference on the received video locally.
[0010] The cloud performs inference on the received video and returns the inference results to the edge device;
[0011] The edge device aggregates its own inference results and cloud inference results as the final result, and tracks the aggregated inference results.
[0012] Server DNN models are high-complexity models, while Edge DNN models are low-complexity models. The model complexity is distinguished by the number of layers: models with 100 or more layers are high-complexity models, while those with fewer than 100 layers are low-complexity models.
[0013] The partitioning rules for generating video frames include:
[0014] The camera at the acquisition end continuously captures video, and after encoding the captured video, it maintains high-quality encoding and streams it to the cloud;
[0015] After decoding the received data, the cloud performs depth estimation and DNN inference, then generates the camera partitioning rules and transmits them to the acquisition end.
[0016] The partitioning rules assign different quality encodings to different blocks, including:
[0017] 1) Divide the video frame into multiple tiles, where each tile is a square macroblock with a width that is the greatest common factor of the width and height of the video frame;
[0018] 2) The cloud performs inference on the decoded image using ServerDNN and EdgeDNN respectively, and finds the tile regions with different inference results;
[0019] 3) Perform depth estimation on the video frame to obtain the depth value of each pixel in the frame and calculate the average depth value of the pixels contained in each tile of the video frame;
[0020] 4) Calculate the average depth value of the different tiles obtained in step 3), which is the unloading threshold;
[0021] 5) Based on the unloading threshold in step 4), mark the completion of the tiles partitioning rule generation, and return the partitioning rule to the camera in the form of a hash value as the basis for its partitioning encoding.
[0022] Region segmentation includes:
[0023] 1) Decode the received video data;
[0024] 2) Perform inference on the decoded video;
[0025] 3) Traverse different blocks of the video frame and cover all low-quality encoded blocks at the acquisition end with the same pixel value;
[0026] 4) Send the modified video frames to the cloud for inference.
[0027] The aggregation inference results are tracked using an attention-based LSTM module, which is based on an attention-based LSTM network structure.
[0028] An apparatus includes a data acquisition end, an edge end, and a cloud end, wherein a complex DNN model is deployed in the cloud end and a simple DNN model is deployed in the edge end, and the cloud-edge collaborative video stream analysis method based on depth estimation is run between the data acquisition end, the edge end, and the cloud end.
[0029] Compared with the prior art, the present invention has the following beneficial effects:
[0030] 1. This invention proposes selecting Region of Interest (ROI) regions for video frames based on different image depths, solving the problem of different DNN models having different ROI definitions. Different DNN models can share the same depth estimation result. Changing the DNN model during system inference requires no additional operations.
[0031] 2. This invention proposes using LSTM to dynamically predict inference results from historical inference results as a supplement to video analysis, so as to greatly reduce latency while ensuring accuracy. Attached Figure Description
[0032] Figure 1 This is a framework diagram of the method of the present invention;
[0033] Figure 2 A flowchart illustrating the workflow for generating partitioning rules in the cloud. Detailed Implementation
[0034] To make the objectives, technical solutions, and advantages of this application clearer, the application is described and illustrated below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the application. All other embodiments obtained by those skilled in the art based on the embodiments provided in this application without inventive effort are within the scope of protection of this application. Furthermore, it is understood that although the efforts made in such a development process may be complex and lengthy, for those skilled in the art related to the content disclosed in this application, modifications to design, manufacturing, or production based on the technical content disclosed in this application are merely conventional technical means and should not be construed as insufficient disclosure of the content of this application.
[0035] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment that is mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described in this application may be combined with other embodiments without conflict.
[0036] Unless otherwise defined, the technical or scientific terms used in this application shall have the ordinary meaning understood by one of ordinary skill in the art to which this application pertains. The terms "a," "an," "an," "the," and similar words used in this application do not indicate quantity limitation and may indicate singular or plural. The terms "comprising," "including," "having," and any variations thereof used in this application are intended to cover non-exclusive inclusion; "multiple" used in this application means two or more. "And / or" describes the relationship between related objects, indicating that three relationships may exist. For example, "A and / or B" can represent: A alone, A and B simultaneously, and B alone. The terms "first," "second," "third," etc., used in this application are merely to distinguish similar objects and do not represent a specific ordering of objects.
[0037] In summary, the method of this invention includes: deploying a high-complexity DNN inference model in the cloud and a low-complexity DNN inference model at the edge. Then, by introducing depth estimation technology, video frame partitioning rules are generated in the cloud based on the depth of the video frames. The acquisition end encodes different blocks of different quality according to the cloud partitioning rules and transmits them to the edge server. While performing DNN inference on the video, the edge server streams the high-quality encoded blocks back to the cloud for DNN inference. Finally, the edge server aggregates its own inference results and the cloud inference results as the final result, and uses an attention-based LSTM module to track the aggregated inference result.
[0038] To achieve the objectives of this invention, the technical solution adopted is summarized as follows:
[0039] A cloud-edge collaborative video stream analysis method based on depth estimation includes the following steps:
[0040] (1) Deploying a complex DNN model in the cloud is called Server DNN in this paper, while deploying a simple DNN model at the edge is called Edge DNN in this paper.
[0041] (2) The camera at the acquisition end continuously captures video, and after encoding the captured video, it maintains high-quality encoding and streams it to the cloud.
[0042] (3) After the cloud decodes the received data, it performs depth estimation and DNN inference respectively, and then generates the camera partitioning rules and transmits them to the acquisition end.
[0043] The partitioning rule encoding in step (3) includes the following steps:
[0044] 1) Divide the video frame into multiple tiles, where each tile is a square macroblock with a width that is the greatest common factor of the width and height of the video frame;
[0045] 2) The decoded image is inferred using Server DNN and Edge DNN in the cloud to find the tile regions with different inference results.
[0046] 3) Perform depth estimation on the video frame to obtain the depth value of each pixel in the frame and calculate the average depth value of the pixels contained in each tile of the video frame.
[0047] 4) Calculate the average depth value of different tiles obtained in step 3), which is the unloading threshold.
[0048] 5) Based on the unloading threshold in step 4), mark the tiles partitioning rules as complete. Return the partitioning rules to the camera in the form of hash values as the basis for its partitioning encoding.
[0049] (4) The acquisition end divides the acquired video into partitions according to the rules generated in the cloud and encodes them with different qualities. After the encoding is completed, the video is streamed to the edge server.
[0050] (5) The edge end will perform regional segmentation on the received video, transmit the high-quality encoded blocks to the cloud server for inference, and at the same time perform inference on the received video locally.
[0051] The region cutting in step (5) includes the following steps:
[0052] 1) Decode the received video data.
[0053] 2) Perform inference on the decoded video.
[0054] 3) Traverse different blocks of the video frame and cover all low-quality encoded blocks at the acquisition end with the same pixel value.
[0055] 4) Send the modified video frames to the cloud for inference.
[0056] (6) The cloud performs inference on the received video and returns the inference results to the edge.
[0057] (7) The edge end aggregates the local inference results and the cloud inference results as the final inference result of the video frame and tracks it using the classic attention-based LSTM module.
[0058] This embodiment provides an example of a cloud-edge collaborative video stream analysis system based on depth estimation.
[0059] In this embodiment, as Figure 1 The diagram illustrates the overall workflow of video stream analysis in a cloud-edge network environment. The specific steps are as follows:
[0060] Pre-start phase: The acquisition end performs lossless encoding on the acquired video and transmits it to the cloud. The cloud decodes the received video and sends it to the partitioning rule generation module. This module uses a partitioning rule generation algorithm to generate a partitioning scheme and transmits the result back to the acquisition end. For example... Figure 2 The diagram illustrates the workflow for generating partitioning rules, with the specific steps as follows:
[0061] Suppose that the video frame is divided into 5x8 blocks, with a length of 8 and a width of 5.
[0062] The cloud performs inference on the decoded video frames using both Server DNN and Edge DNN to find blocks where the inference results differ. Let the set of blocks be denoted as . Each number represents a block number where the inference results from the cloud and the edge differ.
[0063] Depth estimation is performed on the video frames to obtain the average depth set of each divided block. Here, each 'a' represents the height of a block.
[0064] The average depth (AVD) of each block in the block set is calculated using the following formula:
[0065]
[0066] Using the obtained AVD as the offloading threshold, each block in AveDep is marked. Blocks with a depth greater than AVD are marked as HQ, indicating high-quality coding, while blocks with a depth less than AVD are marked as LQ, indicating low-quality coding.
[0067] Partitioned Transmission Phase: The acquisition end encodes the video according to the partitioning scheme of the first phase and then streams it to the edge end. After decoding the video, the edge end calls the region segmentation algorithm in the region segmentation module to segment the video frame by frame, and then transmits the segmented results to the cloud server. The region segmentation steps are as follows:
[0068] The received video data is decoded.
[0069] The decoded video is used for inference to obtain the inference result LD.
[0070] Traverse different blocks of the video frame, and re-encode all low-quality encoded blocks from the acquisition end with the same pixel values.
[0071] The modified video frames are sent to the cloud for inference.
[0072] Real-time inference stage: Server DNN and Edge DNN are used for inference on the decoded video at the cloud and edge respectively. The cloud inference result is returned to the edge, where the inference results from both the cloud and edge are aggregated to obtain the final inference result. An attention-based LSTM module is used to track the inference results of historical frames. The aggregation inference steps are as follows:
[0073] The edge device receives the inference result (HD) returned from the cloud.
[0074] At the edge, compare the inference results of HD and LD obtained from region segmentation in the first block. If they are the same, no processing is performed; otherwise, the inference result of HD is used to overwrite the inference result of LD.
[0075] The final result JD is obtained by iteratively processing the 40 blocks according to the steps above.
[0076] The inference steps for historical frames by the attention-based LSTM module are as follows:
[0077] Input the aggregated inference results of the first N frames into the LSTM module .
[0078] The LSTM module outputs the processing result of the new frame. .
[0079] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0080] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art could make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.
Claims
1. A cloud-edge collaborative video stream analysis method based on depth estimation, characterized in that, include: Deploy the Server DNN model in the cloud and the Edge DNN model at the edge; In the cloud, video frame partitioning rules are generated based on the depth of the video frame; The acquisition end encodes different blocks with different qualities according to the cloud partitioning rules and then transmits them to the edge end; The partitioning rules assign different quality encodings to different blocks, including: 1) Divide the video frame into multiple tiles, where each tile is a square macroblock with a width that is the greatest common factor of the width and height of the video frame; 2) The cloud uses both complex and simple DNN structures to infer the decoded image and find the tile regions with different inference results; 3) Perform depth estimation on the video frame to obtain the depth value of each pixel in the frame and calculate the average depth value of the pixels contained in each tile of the video frame; 4) Calculate the average depth value of the different tiles obtained in step 3), which is the unloading threshold; 5) Based on the unloading threshold in step 4), mark the completion of the tiles partitioning rule generation, and return the partitioning rule to the camera in the form of a hash value as the basis for its partitioning encoding; The edge device segments the received video into regions, transmits the high-quality encoded blocks to the cloud server for inference, and simultaneously performs inference on the received video locally. The cloud performs inference on the received video and returns the inference results to the edge device; The edge device aggregates its own inference results and cloud inference results as the final result, and tracks the aggregated inference results.
2. The cloud-edge collaborative video stream analysis method based on depth estimation according to claim 1, characterized in that, The Server DNN model is a high-complexity model, and the Edge DNN model is a low-complexity model. The model complexity is distinguished by the number of layers. Models with 100 or more layers are considered high-complexity models, and those with fewer than 100 layers are considered low-complexity models.
3. The cloud-edge collaborative video stream analysis method based on a low-complexity model of depth estimation according to claim 1, characterized in that, The partitioning rules for generating video frames include: The camera at the acquisition end continuously captures video, and after encoding the captured video, it maintains high-quality encoding and streams it to the cloud; After decoding the received data, the cloud performs depth estimation and DNN inference, then generates the camera partitioning rules and transmits them to the acquisition end.
4. The cloud-edge collaborative video stream analysis method based on depth estimation according to claim 1, characterized in that, The region cutting includes: 1) Decode the received video data; 2) Perform inference on the decoded video; 3) Traverse different blocks of the video frame and cover all low-quality encoded blocks at the acquisition end with the same pixel value; 4) Send the modified video frames to the cloud for inference.
5. The cloud-edge collaborative video stream analysis method based on depth estimation according to claim 1, characterized in that, The aggregation inference results are tracked using an attention-based LSTM module, which is based on an attention-based LSTM network structure.
6. An apparatus, characterized in that: It includes a data acquisition terminal, an edge terminal, and a cloud terminal. A complex DNN model is deployed in the cloud terminal, and a simple DNN model is deployed in the edge terminal. The data acquisition terminal, the edge terminal, and the cloud terminal run the cloud-edge collaborative video stream analysis method based on depth estimation as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Ship online monitoring method and terminal
CN113792619A
Cloud edge collaborative video stream processing method and system for unmanned system
CN114154018A
Dynamic frequency and deep learning model unloading joint adjustment method and system based on deep reinforcement learning
CN115827239A