Intelligent video stream processing method based on edge cloud collaboration

Through the intelligent video streaming processing method of edge-cloud collaboration, the lightweight object detection model L-Yolov4 and feature map difference comparison method are used to solve the problems of high latency and high bandwidth cost of traditional cloud monitoring methods, and realize efficient and reliable video streaming processing in edge environments.

CN120075478AInactive Publication Date: 2025-05-30CHINA UNIV OF PETROLEUM (EAST CHINA)
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510133581.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-05-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional cloud video surveillance methods have problems with high transmission latency and high bandwidth costs, and the computing power and memory resources of edge computing are limited, making it difficult to achieve high-precision security monitoring.

Method used

A smart video stream processing method with edge and cloud collaboration is proposed. By designing a lightweight object detection model L-Yolov4, the model is split and deployed on edge devices and cloud centers, and the feature map difference comparison method and compression algorithm are used to optimize transmission bandwidth and computing resource usage.

Benefits of technology

It realizes an effective balance between detection delay and accuracy, reduces transmission bandwidth consumption, improves the efficiency and reliability of video stream processing, and is suitable for resource-constrained edge environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120075478A_ABST
    Figure CN120075478A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of big data processing, and particularly relates to an intelligent video stream processing method based on edge cloud collaboration. The invention provides a feature difference comparison method, which is used for comparing the features output by different images at a segmentation layer, transmitting the images with smaller feature difference to cloud for detection, detecting the images with larger feature difference at the edge, and not detecting the images without difference basically, so that batch and hierarchical processing is realized, and repeated detection of target-free images is avoided. According to the method, the features to be transmitted to the cloud are quantified and compressed, the bandwidth loss is reduced, and the feature difference parameters are continuously updated in the edge-cloud cooperative detection process, so that the number of the edge and cloud processing images is controlled, rapid detection of the edge and accurate detection of the cloud are fully exerted, and the detection efficiency is improved. And relatively high detection precision is achieved on the premise of ensuring real-time performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of big data processing, and particularly relates to an intelligent processing method for video streams with edge-cloud collaboration. Background Art

[0002] With the rapid development of China's economy and technology, people's living standards have been continuously improved, and the public's attention to life health and production safety has increased day by day. Traditional safety management relies on a large amount of manpower. However, due to the complex factory environment and limited manpower, it is difficult for manual labor to comprehensively cover potential safety hazards. A slight oversight may lead to serious property losses and even threaten the lives of workers. Therefore, the current safety management mode urgently needs a more meticulous and efficient solution. The progress of computer vision technology provides new means and technical support for safety production management.

[0003] In modern factories, although a large number of surveillance cameras are installed, the traditional safety management mode still relies on manual monitoring and management of videos, resulting in a large amount of manpower consumption. The safety management mode based on computer vision can use cameras distributed in every corner of the factory for 24-hour all-weather safety monitoring, significantly reducing the labor and material costs, and even achieving unattended automated management. However, image recognition technology has high requirements for computing resources, and traditional factory servers often struggle to meet these needs. The introduction of cloud computing provides remote computing and storage services for factories, effectively solving the problem of insufficient computing resources. Currently, many factories upload video data to the cloud for hazard source identification. However, due to the huge amount of video data, the uploading process consumes a large amount of network resources, and coupled with the high latency of cloud processing, it is impossible to achieve timely response and difficult to quickly handle potential safety hazards.

[0004] With the popularization of video surveillance and the increase in safety management tasks, the amount of factory video data has increased significantly. At the same time, safety production management requires real-time monitoring and rapid disposal of potential safety hazards, but traditional single computing modes (such as cloud computing or edge computing) cannot meet the requirements of big data processing and high timeliness at the same time. Although cloud computing can process a large amount of data, due to its high latency and long-distance transmission, it is difficult to achieve real-time response to sudden safety hazards. While edge computing can provide low-latency processing, its computing and storage resources are limited and cannot independently support the safety monitoring needs of a large-scale factory area. Summary of the Invention

[0005] The present invention proposes an edge-cloud collaborative intelligent video stream processing method, which solves the problems of high transmission delay and high bandwidth cost existing in traditional cloud monitoring methods, overcomes the limitations of limited computing power and memory resources at the edge, and at the same time, this method overcomes the bottleneck of limited computing power and memory resources at the edge, and solves the problem that it is difficult to deploy high-precision complex models at the edge, resulting in insufficient detection accuracy.

[0006] The technical solution of the present invention is implemented as follows:

[0007] An edge-cloud collaborative intelligent video stream processing method includes the following steps:

[0008] S1. Design a lightweight object detection model L-Yolov4 to obtain an object detection model suitable for resource-constrained edge devices;

[0009] S2. Modularly split the object detection model Yolov4, where the feature aggregation module and the detection head part are deployed in the cloud center, while the lightweight object detection model L-Yolov4 is deployed on the edge device;

[0010] S3. Propose a feature map difference comparison method to find video frames with a higher probability of the existence of an object by comparing the feature differences between two frames of images;

[0011] S4. Compress the output of the i-th layer through a compression algorithm to reduce the consumption of transmission bandwidth;

[0012] S5. According to the detection results, update and adjust the threshold parameters of the feature map in real time to make the edge-cloud collaborative intelligent video stream processing architecture system more adaptable to the dynamic environment;

[0013] S6. Output the detection results.

[0014] Optionally, in step S1, the L-Yolov4 model reduces the output of the three-scale feature maps C 3 、C 4 and C 5 in the original Yolov4 to two-scale feature maps C 4 and C 5 . In the model architecture, the feature map C 5 output by the backbone network is input into the SPP module to generate a fused feature map M 5 , M 5 generates a feature map P ′ 5 after a series of convolutional operations, and generates corresponding prediction results through the Yolo detection head. At the same time, the feature map M 5 adjusts the number of channels and resolution to those of the feature map C 4 through convolutional and upsampling operations.are the same, and then, through the concatenation operation in the channel dimension, C 4 and M 5 are fused, and a feature map P ′ 4 is generated through a convolution operation. Finally, the model outputs prediction results P 4 and P 5 .

[0015] Optionally, in step S3, the basic idea of the feature map difference comparison method is as follows: The detection program in the edge server extracts consecutive frames from the real-time monitoring video stream at a set regular time interval t, and then uses the detection program to extract the feature information of the f j frames on channel k of the output of the i-th layer network to obtain a feature matrix f j The sum of the feature elements of the frames and the calculation methods of the difference d between f j and the original image are as follows:

[0016]

[0017] where n ij is the element included in the feature matrix N, and the values of x and y are both 52;

[0018]

[0019] where F represents the sum of the feature elements obtained from the original frame. Compare the size of the difference d with the threshold α to determine whether to continue detecting this frame. If continue to detect, then compare d with the threshold β to determine whether to continue detecting at the edge or upload it to the cloud server for detection.

[0020] Optionally, in step S4, 8-bit is selected to quantize the output feature values as follows:

[0021]

[0022] where V ∈ R N×M×C , is the quantized feature tensor, representing the feature value on channel C at the N-th row and M-th column in the feature tensor. V min and V max represent the maximum and minimum feature values of the current image at the output of the i-th layer network;

[0023] Upload V min and V max to the cloud center, and at the cloud center, the restoration of the quantized feature is realized by inverse quantization. The inverse quantization is expressed as follows:

[0024]

[0025] Among them, in the formula, V min and V max are the maximum and minimum feature values uploaded from the edge server to the cloud server.

[0026] Optionally, in step S5, after the image detection is completed in the present invention, the current detection time is recorded and stored in the edge server. The model for image detection is jointly composed of a partial model deployed in the edge device and a partial model in the cloud center. These two partial models cooperate to complete the image detection task. Subsequently, the detection delay of the current frame is compared with the average detection delay t of the cooperative model. If the detection delay is greater than t, the interval length of [α, β] is reduced, that is: β is decreased or α is increased. If the detection delay is less than t, the interval length of [α, β] is increased, that is: β is increased or α is decreased. The value ranges of α and β are α = [a 1 , a 2 , β = [b 1 , b 2 ;

[0027] The update formula for β is as follows:

[0028] β new = max{min{β old − γ 1 × (Q d × t d − s), b 2}, b 1};

[0029] Among them, γ 1 is the weight parameter, with a value range of [0, 1], Q d is the number of video frame sequences to be detected stored in the edge server, t d is the current average detection delay of the model, s is the extraction interval time of the video frames, b 1 represents the minimum value of the parameter β, and b 2 represents the maximum value of the parameter β;

[0030] The update formula for α is,

[0031] α new = max{min{γ 2 × (0.1 − β new ), a 2}, a 1};

[0032] Among them, γ 2 is the weight parameter, with a value range of [0, 1], a 1 represents the minimum value of the parameter α, and a 2 represents the maximum value of the parameter α.

[0033] After adopting the above technical solution, the beneficial effects of the present invention are as follows:

[0034] The present invention proposes an edge-cloud collaborative intelligent video stream processing method, aiming to overcome the deficiencies of pure edge computing and pure cloud computing. This method combines the advantages of edge computing and cloud computing, uses the lightweight object detection model L-Yolov4 to perform preliminary processing of images on edge devices, and transfers some tasks to the cloud. This method not only solves the problem of detection accuracy caused by limited computing power in edge computing but also avoids the challenges of high transmission latency and excessive bandwidth costs in cloud computing. Specifically, edge devices are responsible for real-time extraction of video stream features and use L-Yolov4 for preliminary processing, thereby improving processing efficiency in an environment with limited resources. For tasks with large feature differences, the system uploads them to the cloud for further processing through the edge-cloud collaboration mechanism. Through this collaborative method, the present invention effectively balances detection latency and accuracy, ensuring both real-time response and high detection accuracy, and significantly improving the efficiency and reliability of video stream processing.

[0035] The present invention proposes a feature difference comparison method to compare the features output by different images at the segmentation layer, transmit the features with small differences to the cloud for detection, detect the features with large differences at the edge, and do not detect the features with almost no differences, realizing batch and hierarchical processing and avoiding repeated detection of targetless images.

[0036] The present invention quantizes and compresses the features to be transmitted to the cloud to reduce bandwidth loss. During the edge-cloud collaborative detection process, the feature difference parameters are continuously updated, thereby controlling the number of images processed by the edge and the cloud, giving full play to the fast detection of the edge and the precise detection of the cloud, and achieving high detection accuracy on the premise of ensuring real-time performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings without creative efforts.

[0038] Figure 1 is the network structure of lightweight L-Yolov4;

[0039] Figure 2 is the segmentation method of the object detection model Yolov4;

[0040] Figure 3 is the training loss curve of the smoking detection model;

[0041] Figure 4 is the training loss curve of the phone call detection model;

[0042] Figure 5 is the training loss curve of the pedestrian detection model;

[0043] Figure 6 are the detection results of pedestrians, smoking, and phone calls of L-Yolov4;

[0044] Figure 7 are the detection results of pedestrians, smoking, and phone calls of Yolov4;

[0045] Figure 8 is the image extracted from the video image;

[0046] Figure 9 is the feature difference value between different images;

[0047] Figure 10 The change curves of parameters α and β. Specific implementation manners

[0048] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0049] The embodiments of the present application disclose an intelligent processing method for video streams with edge-cloud collaboration.

[0050] 1. Method

[0051] According to Figures 1 to 10 As shown, an intelligent processing method for video streams with edge-cloud collaboration includes the following steps:

[0052] S1. Design a lightweight object detection model L-Yolov4 to obtain an object detection model suitable for resource-constrained edge devices;

[0053] S2. Split the object detection model Yolov4, and deploy the feature aggregation module and the detection head part in the Yolov4 model after splitting in the cloud center, and deploy the lightweight object detection model L-Yolov4 on the edge device;

[0054] S3. Propose a feature map difference comparison method to find video frames with a greater possibility of the existence of the target by comparing the feature differences between two frames of images, avoid the detection of redundant frames, and reduce the communication bandwidth;

[0055] S4. Compress the output of the i-th layer through a compression algorithm to reduce the consumption of transmission bandwidth;

[0056] S5. According to the detection results, update and adjust the threshold parameters of the feature map in real time to make the edge-cloud collaborative video stream intelligent processing architecture system more adaptable to the dynamic environment;

[0057] S6. Output the detection results.

[0058] In step S1, the present invention improves on the basis of the Yolov4 network structure (as Figure 1 shown), Figure 1 The content shown in is the improved content. The present invention compresses the network structure to reduce the model complexity. The L-Yolov4 model reduces the output of the three-scale feature maps C 3 , C 4 and C 5 in the original Yolov4 to two-scale feature maps C 4 and C 5 to effectively reduce the computational complexity. In the model architecture, the feature map C 5 output by the backbone network is input into the SPP module for fusing local features and global features to generate the fused feature map M 5 . M 5 generates the feature map P ′ 5 after a series of convolutional operations, and generates the corresponding prediction results through the Yolo detection head. At the same time, the feature map M 5 is adjusted in terms of the number of channels and resolution to be the same as that of the feature map C 4 through convolutional and upsampling operations. Then, through the splicing operation in the channel dimension, C 4 and M 5 are fused and generate the feature map P ′ 4 . Finally, the model outputs the prediction results P 4 and P 5 at two scales.

[0059] In step S2, the present invention splits the Yolov4 network structure into two parts (as Figure 2 shown), and deploys them on the edge server and the cloud server respectively. The original Yolov4 contains a total of 161 layers of networks, and the number of parameters output by each layer and the memory size occupied by the output parameters are different. Considering the transmission bandwidth loss between the edge server and the cloud server and the memory occupancy rate of the edge server, the network output at the split should be as small as possible to reduce the transmission bandwidth loss; after the feature extraction of the original Yolov4 by CSPDarkNet53, C 3 , C 4 and C5 The outputs of three scales, and subsequently, upsampling and feature aggregation processes are mainly performed on the three scales, where C 3 has a parameter quantity of 52 * 52 * 256, and C 4 is 26 * 26 * 512, and C 5 is 13 * 13 * 1024, and C 5 has a parameter quantity much smaller than that of C 3 and C 4 . If the transmission bandwidth is to be minimized, the network layer that outputs C 5 should be selected as the segmentation layer. However, in the original Yolov4, C 3 and C 4 need to participate in subsequent feature fusion, and two feature maps need to be uploaded to the cloud server, which greatly increases the transmission bandwidth. To minimize the transmission bandwidth, the network layer that outputs C 3 is used as the final segmentation layer. Since the feature extraction part of the lightweight L-Yolov4 has the same structure as that of the original Yolov4, the lightweight L-Yolov4 model is deployed on the edge server, and the network structure after the Yolov4 segmentation layer is deployed on the cloud server.

[0060] In step S3, the present invention proposes a method for comparing feature map differences. By comparing the feature differences between two frames of images, video frames with a greater possibility of the presence of the target are found, redundant frame detection is avoided, and the communication bandwidth is reduced. The basic idea of this method is that the detection program in the edge server extracts consecutive frames from the real-time monitoring video stream at a set regular time interval t, and then the detection program extracts the feature information of the f j frame on channel k of the output of the i-th layer network to obtain the feature matrix Calculate the sum of the feature elements of the f j frame using formula (1) Obtain the difference d between the f j frame and the original image using formula (2).

[0061]

[0062] In formula (1), n ij is the element contained in the feature matrix N. The present invention uses an image with a size of 416 * 416 as the input, and the segmentation layer is the network layer where C5 is located. Therefore, the values of x and y are 52.

[0063]

[0064] Among them, F represents the sum of the feature elements obtained from the original frame. Compare the size of the difference d with the threshold α to determine whether to continue detecting this frame. If continue to detect, then compare d with the threshold β to determine whether to continue detecting at the edge or upload it to the cloud server for higher-precision detection.

[0065] In step S4, the present invention compresses the output of the i-th layer through a compression algorithm, thereby reducing the consumption of transmission bandwidth. Most existing deep convolutional neural networks use 32-bit floating-point representation, wasting a large amount of communication resources during network transmission. According to the lossless network feature compression method in [reference], the output features at the splitting point i-th layer are compressed and transmitted to the cloud server. Before compression, an n-bit uniform quantizer is used to quantize the feature values. The present invention selects 8-bit to quantize the output feature values, as shown in formula (3).

[0066]

[0067] Among them, V ∈ R N×M×C , is the quantized feature tensor, representing the feature value on the channel C at the N-th row and M-th column in the feature tensor, V min and V max represent the maximum and minimum feature values of the current image output at the i-th layer of the network. A large number of studies have shown that when n > 6, the impact of the n-bit uniform quantizer on the accuracy of image classification and object detection can be ignored.

[0068] In the edge-cloud collaboration framework of the present invention, the features are transmitted to the cloud server only when the feature difference α < d ≤ β. Therefore, when this condition is met, it is necessary to calculate the maximum and minimum feature values V min and V max of the output features of the current image, and use formula (3) to quantize the original feature matrix, representing the quantized feature values with fewer bits to save the transmission bandwidth. Transmit V min and V max to the cloud center, and at the cloud center, the inverse quantization of formula (4) is used to restore the quantized features .

[0069]

[0070] Among them, in the formula, V min and V max are the maximum and minimum feature values of the features uploaded from the edge server to the cloud server.

[0071] In step S5, after the image detection is completed in the present invention, the current detection time is recorded and stored in the edge server. The model used for image detection is jointly composed of a partial model deployed in the edge device and a partial model in the cloud center. These two partial models cooperate to complete the image detection task. Subsequently, the detection delay of the current frame is compared with the average detection delay t of the collaborative model. If the detection delay is greater than t, the interval length of [α, β] is reduced (β is decreased or α is increased), thereby reducing the probability of transmission to the cloud server, enabling more images to be detected in the edge server, and shortening the detection time. If the detection delay is less than t, the interval length of [α, β] is increased (β is increased or α is decreased), thereby increasing the probability of video frame transmission to the server, improving the detection accuracy, reducing the memory occupancy rate of the edge server, and shortening the detection time of other abnormal events at the same time. The value ranges of α and β are α = [a 1 , a 2 , β = [b 1 , b 2 ,

[0072] The update formula for β is as shown in (5):

[0073] β new = max{min{β old ― γ 1 × (Q d × t d ― s), b 2}, b 1} (5)

[0074] Among them, γ 1 is the weight parameter, and its value range is [0, 1]. Q d is the number of video frame sequences to be detected stored in the edge server. t d is the current average detection delay of the model. s is the extraction interval time of the video frame. b 1 represents the minimum value of the parameter β and serves as the lower limit. b 2 represents the maximum value of the parameter β and serves as the upper limit. b 1 and b 2 jointly define the value range of the parameter β to ensure that β always changes within a reasonable interval to guarantee system stability and performance.

[0075] The update formula for α is:

[0076] α new = max{min{γ 2 × (0.1 ― β new ), a 2}, a 1} (6)

[0077] Among them, γ 2is a weight parameter, and its value range is [0, 1], a 1 represents the minimum value of parameter α, the lower limit constraint, a 2 represents the maximum value of parameter α, the upper limit constraint, a 1 and a 2 together define the value range of parameter α, ensuring that α always varies within a reasonable interval so as to maintain stability and reliability during system operation. a in the formula 1 , a 2 , b 1 , b 2 can be calculated in two cases, and the calculation basis for the two cases is as follows:

[0078] When the edge server has a high load, increase the proportion of cloud processing and reduce the amount of tasks at the edge. When the edge server has a low load, improve the task processing ability at the edge and reduce the dependence on the cloud.

[0079] (1) When the load of the edge server is high (the number of video frames to be detected stored is large or the detection delay t is large), more detection tasks need to be transferred to the cloud to reduce the pressure on the edge server. a 1 represents the minimum value of parameter α, which is used to ensure that the load of the edge server will not be too high (control the minimum lower limit). a 2 represents the maximum value of parameter α, which is used to prevent the edge server from completely stopping the detection task (control the maximum upper limit). b 1 represents the minimum value of parameter β, which limits the probability of tasks being transferred to the cloud and ensures that the edge server will not be idle due to too few tasks. b 2 represents the maximum value of parameter β, which avoids excessive tasks being transferred to the cloud resulting in excessive pressure on bandwidth and computing resources.

[0080] (2) When the load of the edge server is low (for example, the number of video frames to be detected stored is small or the detection delay is small), more tasks can be processed on the edge device, thereby reducing the cost of data transmission to the cloud. a 1 In this case, it is adjusted to ensure that α will not be too close to the maximum value to avoid insufficient utilization of edge device resources. a 2 is adjusted to allow a larger value so that more tasks can be completed at the edge. b 1 is adjusted to a larger value to reduce the probability of tasks being transferred to the cloud and further reduce the network transmission overhead. b 2 is adjusted to a smaller value to ensure that the edge server completes most of the detection tasks instead of relying too much on the cloud.

[0081] Given a priori, take a historical video sequence, extract features from the video frame images, and calculate the difference between the image features when there is a target and when there is no target. a 2The difference value when the target first appears, a 1 takes values in the range (0, a 2 ), and b 1 takes the minimum feature difference when there is a target (the difference value is greater than a 2 ), and b 2 is the mean of the minimum feature difference and the maximum feature difference when there is a target.

[0082] Posteriorly, it is obtained that a 1 , a 2 , b 1 , and b 2 are all set to 0, that is, at the beginning, all images are detected on the cloud server, and the feature matrix of the first image without a target is saved. When the target is first detected, the feature difference value at this time is recorded, which is a 2 , then the value range of a 1 is (0, a 2 ). When the target is detected 50 times in total, the minimum feature difference value of the 50 times is calculated as b 1 (b 1 > a 2 ), and the mean of the minimum feature difference value and the maximum feature difference value of the 50 times is b 2 .

[0083] The edge-cloud collaboration in the present invention refers to the collaboration between edge computing and cloud computing.

[0084] 2. Experimental part

[0085] First, an experimental platform for the edge-cloud collaborative intelligent video surveillance system is built. NVIDIA Jetson TX2 is used as the edge server, and NVIDIA GTX1080Ti is used as the cloud server and model training server. The Pytorch framework is adopted to test and analyze the performance of the two detection models and the edge-cloud collaborative video surveillance system.

[0086] 2.1 Model performance analysis

[0087] (1) Experimental settings

[0088] Assume that a scenario (camera) needs to detect three types of targets: pedestrians, phone call behaviors, and smoking behaviors. The priorities of the three types of targets are set as pedestrians > smoking > phone call. The present invention collects and organizes three groups of datasets: smoking dataset (2400 images), phone call dataset (2500 images), and pedestrian dataset (3120 images), and trains the datasets through the Yolov4 model and the lightweight L-Yolov4 model respectively to obtain a smoking detection model, a phone call behavior detection model, and a pedestrian detection model.

[0089] Parameter settings: Before the training starts, set the initial learning rate to 0.01, and use the cosine annealing algorithm to continuously adjust the learning rate during the training process. Set the number of training epochs to 200, and train based on the pre-trained model (the model obtained by training on the COCO dataset). To accelerate the model convergence speed, when using Yolov4 to train the detection model, freeze the feature extraction network parameters in the first 100 training epochs, train the remaining network structure, and set the batch size to 4; in the last 100 training epochs, train the entire network structure and set the batch size to 2. When using the lightweight L-Yolov4 model to train the detection model, use the weights of the Yolov4 feature extraction network as the feature extraction parameters, freeze the network parameters of the feature extraction part, and only train the remaining network parameters. During the training process, divide each group of datasets into a training set, a validation set, and a test set according to the ratio of 7:2:1.

[0090] (2) Model comparison and analysis

[0091] Figures 3 - 5 is the loss curve during the training process on the three groups of datasets, where Figure 3 is the training loss curve of the smoking detection model, Figure 4 is the loss generated during the training process of the phone call detection model, Figure 5 is the loss curve of the pedestrian detection model.

[0092] Since the present invention retrains based on the weights of the pre-trained model, the initial loss is relatively low, and the model can converge rapidly within 25 training epochs. The present invention selects the model with the lowest validation set loss as the final detection model, and verifies the performance of the model on the test set, as shown in Table 1.

[0093] Table 1 Detection results of different models on the test set

[0094]

[0095] Use the trained Yolov4 detection model and the lightweight L-Yolov4 detection model to detect and visualize the images containing the targets, as shown in Figure 6 and Figure 7 shown.

[0096] As can be seen from the figure, Yolov4 is superior to the lightweight L-Yolov4 in terms of object localization and recognition accuracy. This is mainly because Yolov4 fully integrates semantic information and location information through multi-scale feature aggregation, resulting in a richer feature map. It also makes predictions at three scales, significantly improving the detection accuracy. Yolov4 has more feature scales and a more complex feature aggregation process. Its network structure is refined and has a larger number of parameters, so its overall performance is better than that of L-Yolov4. However, the relatively large number of parameters makes it difficult to deploy Yolov4 on edge devices with limited computing resources. In contrast, as a lightweight version of Yolov4, L-Yolov4 sacrifices some accuracy but significantly reduces the computational complexity and the number of parameters, making it suitable for deployment on edge devices. The lightweight design of L-Yolov4 reduces power consumption and computational burden while increasing the processing speed, making it perform better in application scenarios that require fast response and real-time detection. Therefore, L-Yolov4 is more suitable for edge computing environments, especially when dealing with images with simple features or small feature differences, where it can maintain high real-time performance without pursuing extremely high accuracy.

[0097] In summary, the advantage of the improved L-Yolov4 lies not in improving the detection accuracy, but in optimizing the use of computing resources, especially on edge devices. By reducing the number of parameters in the model, L-Yolov4 increases the processing speed and is particularly suitable for environments with limited computing resources. Although its accuracy is slightly lower than that of Yolov4, in some simple scenarios, the lightweight model is already effective enough to provide real-time detection capabilities. Specifically, by reducing the number of parameters and complexity, L-Yolov4 improves the real-time performance, enabling it to respond quickly on edge devices. For edge devices, real-time performance and fast processing are more important than absolute accuracy, especially in application scenarios that require quick decision-making and response. In addition, the deployment of L-Yolov4 can be combined with Yolov4 in the cloud center to form a collaborative working architecture. In this architecture, the edge device first extracts features and calculates differences for the image. If the feature differences in the image are large, indicating that there may be objects in the image, the features of this image are uploaded to the cloud center for further processing by Yolov4. For images with small feature differences, L-Yolov4 directly processes them on the edge device to obtain the prediction results. Such a design not only avoids uploading all images to the cloud center, saving bandwidth and computing resources, but also ensures that some images with small feature differences that still need to be detected can be further screened to avoid missing potential targets.

[0098] Therefore, L-Yolov4 is suitable for deployment on edge devices with limited computing resources to process images with small feature differences, while Yolov4 is more suitable for deployment in cloud centers with sufficient computing resources to process complex images with large feature differences. The collaborative work of the two can ensure the detection effect while guaranteeing real-time performance.

[0099] 2.2 System Simulation

[0100] Extract a video, frame the video at 10 frames per second, and extract 180 video images. As Figure 8 shown, the target appears in the 8th image and disappears in the 169th image. Use the feature extraction network to obtain the feature information of each image. Take the first image without the target as the original image, save the feature matrix output by the first channel of the network disassembly layer for the original image, and use the proposed feature difference map comparison method to obtain the difference d for the subsequent images. As Figure 9 shown.

[0101] Figure 9 The minimum difference in is obtained based on the average difference when the target disappears and appears. At this time, the value is 0.046. It can be seen that the image feature difference when the target appears is basically greater than the minimum difference. Although there is a target in the image with a feature difference of 0.02 in the figure, it is less than the minimum difference value and will not be detected for this image, resulting in missed detection. However, the missed detection probability value is only 0.018 and can be ignored. According to Figure 9 and the proposed feature difference method, the value ranges of α and β are set as α ∈ [0.02, 0.05], β ∈ [0.1, 0.2]. Set the initial value of α to 0.075 and β to 0.15. The model priority is: pedestrian detection model > smoking detection model > phone call detection model. Input 1000 images containing smoking, phone call, and pedestrian targets into the system, set the batch processing number to 50 (i.e., the detection sequence Q is 50), fix the interval for extracting video frames at 0.5 s, and record the average detection delay and detection accuracy generated during the detection process. Figure 10 is the change curve of the values of α and β during the system detection process. By setting the values of α and β, images are detected only at the edge, only in the cloud, and in edge-cloud collaborative detection, and the data in Table 2 are obtained.

[0102] Table 2 Video Processing Time and Accuracy

[0103]

[0104] Table 2 shows the video processing results under only edge servers, only cloud servers, and edge-cloud collaboration. In the present invention, detection latency and accuracy are used as metrics to evaluate the performance of the above three schemes. The detection time and detection accuracy of each batch of images are statistically analyzed, and the average detection time of 20 batches is used as the final detection latency, and the average detection accuracy is used as the final detection precision. The experimental results show that the edge-cloud collaboration-based scheme proposed in the present invention can balance detection latency and detection accuracy and improve the detection speed while ensuring detection accuracy.

[0105] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the technical solutions of the present invention shall be included within the protection scope of the present invention.

Claims

1. A video stream intelligent processing method for edge-cloud collaboration, characterized in that: The following steps are involved: S1. Design a lightweight target detection model L-Yolov4 to obtain a target detection model suitable for resource-constrained edge devices; S2. The target detection model Yolov4 is modularly split, where the feature aggregation module and the detection head are deployed in the cloud center, while the lightweight target detection model L-Yolov4 is deployed on the edge device; S3. A feature map difference comparison method is proposed to find the video frame where the target is more likely to exist by comparing the feature differences between two frames of images; S4, compressing the output of the i-th layer through a compression algorithm, thereby reducing transmission bandwidth consumption; S5. Update and adjust the threshold parameters of the feature map in real time according to the detection results, so that the edge-cloud collaborative video stream intelligent processing architecture system can be more adaptable to dynamic environments; S6. Output the detection results.

2. According to the edge-cloud collaborative video stream intelligent processing method of claim 1, it is characterized in that: In step S1, the L-Yolov4 model reduces the three scale feature map outputs C3, C4 and C5 in the original Yolov4 to two scale feature maps C4 and C5. In the model architecture, the feature map C5 output by the backbone network is input to the SPP module to generate the fused feature map M5. After a series of convolution operations, M5 generates the feature map P ′ 5, and the corresponding prediction results are generated through the Yolo detection head. At the same time, the feature map M5 is convolved and up-sampled to adjust the number of channels and resolution to the same as the feature map C4. Then, C4 and M5 are fused through the splicing operation of the channel dimension, and the feature map P is generated through the convolution operation. ′ 4. Finally, the model outputs prediction results P4 and P5 at two scales.

3. According to the edge-cloud collaborative video stream intelligent processing method of claim 2, it is characterized in that: In step S3, the basic idea of ​​the feature map difference comparison method is: the detection program in the edge server extracts continuous frames from the real-time monitoring video stream according to the set regular time interval t, and then uses the detection program to extract f j The feature information of the frame on channel k output by the i-th layer network is used to obtain the feature matrix f j The sum of the characteristic elements of the frame and f j The difference d with the original image is calculated as follows: Among them, n ij is the element contained in the feature matrix N, the values ​​of x and y are both 52; Among them, F represents the sum of the feature elements obtained from the original frame. The difference d is compared with the threshold α to determine whether to continue detecting the frame. If so, d is compared with the threshold β to determine whether to continue detecting at the edge or upload it to the cloud server for detection.

4. According to claim 3, the method for intelligent processing of edge-cloud-coordinated video streams is characterized in that: In step S4, 8-bit is selected to quantize the output feature value, as shown below: Where V∈R N×M×C , is the quantized feature tensor, which represents the eigenvalue of the Nth row and the Mth column in the feature tensor on channel C, V min and V max Represents the maximum and minimum eigenvalues ​​of the current image output in the i-th layer network; V min With V max Upload to the cloud center, and realize quantization features by inverse quantization in the cloud center The restoration and inverse quantization of is expressed as follows: Among them, V min and V max It is the maximum and minimum value of the feature uploaded from the edge server to the cloud server.

5. According to the edge-cloud collaborative video stream intelligent processing method of claim 4, it is characterized in that: In step S5, after completing the image detection, the present invention records the current detection time and stores it in the edge server. The model used for image detection is composed of a part of the model deployed in the edge device and a part of the model in the cloud center. The two parts of the model work together to complete the image detection task. Subsequently, the detection delay of the current frame is compared with the average detection delay t of the collaborative model. If the detection delay is greater than t, the interval length of [α, β] is reduced, that is, β is reduced or α is increased. If the detection delay is less than t, the interval length of [α, β] is expanded, that is, β is increased or α is reduced. The value range of α and β is α = [a1, a2], β = [b1, b2]; The β update formula is as follows: b new =max{min{β old ―γ1×(Q d ×t d ―s),b2},b1}; Among them, γ1 is the weight parameter, the value is [0,1], Q d is the number of video frame sequences to be detected stored in the edge server, t d is the average detection delay of the current model, s is the interval time of the extracted video frames, b1 represents the minimum value of parameter β, and b2 represents the maximum value of parameter β; The update formula of α is: a new =max{min{γ2×(0.1―β new ),a2},a1}; Among them, γ2 is the weight parameter, and its value range is [0,1]. a1 represents the minimum value of parameter α, and a2 represents the maximum value of parameter α.

Citation Information

Patent Citations

  • Monitoring video target real-time query method based on edge cloud convolutional neural network cascading

    CN112241719A

  • Unmanned aerial vehicle target detection method based on edge intelligence

    CN112488043A

  • Cloud edge cooperative training and deployment system and method of deep learning model

    CN115033253A

  • Edge cloud decoupling-based neural network reasoning method and device, equipment and medium

    CN116612371A

  • Method for high-precision multi-target tracking against complex background

    WO2022217840A1