Video analysis method, device, electronic device and medium based on edge-cloud collaboration

Through the video analysis method of edge-cloud collaboration, similarity comparison is used to filter similar frames in the video stream, solving the problems of missed and false alarms in the video surveillance system, and improving the accuracy of early warning and personalized configuration capabilities.

CN115103157BActive Publication Date: 2025-05-16ZHONGKEHONGYUN TECH (BEIJING) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210676243.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-15
Publication Date
2025-05-16
Estimated Expiration
2042-06-15

AI Technical Summary

Technical Problem

When existing video surveillance systems identify security violations, they are prone to missed and false alarm problems, and it is difficult to achieve personalized configuration of smart cameras.

Method used

Using a video analysis method based on edge cloud collaboration, the monitoring video stream is obtained, frames are extracted and similarity comparison is performed with the false warning picture set, similar frames are filtered, false alarms are reduced, and non-similar frames are input into the intelligent algorithm model for identification.

Benefits of technology

It improves the early warning accuracy of the video surveillance system, reduces the false alarm rate, enhances the system's personalized configuration capabilities, and can more effectively identify security violations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115103157B_ABST
    Figure CN115103157B_ABST
Patent Text Reader

Abstract

The present application relates to the fields of computer and artificial intelligence technology, and in particular to methods, devices, electronic devices and media for video analysis based on edge-cloud collaboration. The method includes: obtaining a surveillance video stream collected by a target camera; extracting frames from the surveillance video stream based on the acquired algorithm configuration information corresponding to the target camera to obtain multiple current frames; comparing the similarity of each current frame with the acquired false warning picture set to generate a similarity value corresponding to each current frame; comparing each similarity value with the false warning threshold, and treating pictures with similarity values ​​greater than or equal to the false warning threshold as similar frames, and treating pictures with similarity values ​​less than the false warning threshold as non-similar frames; inputting non-similar frames into the intelligent algorithm model corresponding to the algorithm configuration information, generating recognition results, and filtering similar frames. The present application has the effect of improving the accuracy of warnings of the early warning system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the fields of computers and artificial intelligence technologies, and in particular to video analysis methods, devices, electronic devices and media based on edge-cloud collaboration. Background Art

[0002] Currently, video surveillance systems are usually deployed in key areas of industrial production such as energy, electricity, communications, and chemicals, and the corresponding violations are screened manually, which is prone to underreporting and consumes a lot of manpower. Therefore, people hope to use artificial intelligence algorithms to identify security violations in real time to reduce manpower consumption and reduce the occurrence of underreporting of violations.

[0003] In order to solve the above problems, the following methods can be adopted: 1. Purchase cameras with built-in security violation algorithms (such as not wearing a helmet, personnel intrusion, etc.) and redeploy them on site. However, this wastes the existing camera equipment acquisition resources and costs more money to purchase smart cameras. In addition, the smart cameras on the market can only perform general security rule detection and cannot control and customize their own security rule algorithms. 2. The video stream of the monitoring system is directly forwarded to a third-party artificial intelligence cloud platform or a private intelligent platform for analysis. A variety of artificial intelligence algorithms can be freely configured on the intelligent platform; however, the video stream occupies a huge bandwidth. Even if the intelligent platform is privately deployed, it is almost impossible to carry the transmission of video stream data of all devices.

[0004] For the second method mentioned above, edge devices are added to the monitoring system in the relevant technology to perform edge computing. Edge computing is a technology that pushes intelligence and computing closer to reality. Service computing (artificial intelligence recognition algorithm) is deployed close to the device side to improve data processing efficiency and reduce data processing delay. Through the coordination of edge computing and cloud servers, the video surveillance system is made more intelligent.

[0005] At present, in the applied artificial intelligence algorithm model (AI model), it is necessary to set a suitable recognition result confidence threshold for the AI ​​model as a basis for determining whether there is a target to be identified in the scene. The traditional approach is usually based solely on experience and tolerance for misidentification. However, since the actual recognition scenes are mostly complex backgrounds, if the preset confidence threshold is too low, the positive result missed detection rate can be reduced but the negative result misjudgment rate will increase, which will lead to an increase in false alarms. If the preset confidence threshold is set too high, although the negative result misjudgment rate can be reduced, the positive result missed detection rate will increase, which will cause some alarms to be unable to be issued. Therefore, the warning accuracy of the artificial intelligence algorithm model applied in the relevant technology needs to be improved. Summary of the invention

[0006] In order to improve the warning accuracy of the early warning management system, the present application provides a video analysis method, device, electronic device and medium based on edge-cloud collaboration.

[0007] In the first aspect, the present application provides a video analysis method based on edge-cloud collaboration, which adopts the following technical solution:

[0008] A video analysis method based on edge-cloud collaboration, comprising:

[0009] Get the surveillance video stream collected by the target camera;

[0010] Extract frames from the surveillance video stream based on the acquired algorithm configuration information corresponding to the target camera to obtain multiple current frames;

[0011] Compare each current frame with the acquired false warning picture set for similarity, and generate a similarity value corresponding to each current frame;

[0012] Compare each similarity value with the false alarm threshold, take pictures with similarity values ​​greater than or equal to the false alarm threshold as similar frames, and take pictures with similarity values ​​less than the false alarm threshold as non-similar frames;

[0013] The non-similar frames are input into the intelligent algorithm model corresponding to the algorithm configuration information to generate recognition results, and the similar frames are filtered.

[0014] By adopting the above technical solution, similar frames are filtered to avoid the situation where the intelligent algorithm model identifies the image as a warning image after similar frames are input into the intelligent algorithm model, causing a false warning. Non-similar frames can be input into the intelligent algorithm model for further identification. Each frame extracted from the current monitoring video stream is filtered by performing a similarity comparison based on the false warning images stored in the past, and images that may cause false warnings are removed. This can compensate for the false warning rate caused by the empirical confidence threshold, thereby improving the warning accuracy of the warning system.

[0015] In a possible implementation, performing a similarity comparison between any current frame and the acquired false alarm picture set to generate a similarity value corresponding to any current frame includes:

[0016] The false alarm picture set includes at least one picture group, and the picture group includes marked core false alarm pictures;

[0017] Compare the similarity between any current frame and each of the picture groups respectively to generate respective first similarity values, and use the first similarity value with the largest value as the similarity value corresponding to any current frame;

[0018] The step of comparing the similarity between any current frame and any picture group to generate a first similarity value includes:

[0019] If the picture group includes associated pictures related to the core false alarm picture, generating an inference picture group according to any current frame, wherein the inference picture group includes derivative pictures corresponding to each of the associated pictures;

[0020] The inference picture group is compared with any one of the picture groups in terms of similarity to generate a first similarity value between any one of the current frames and any one of the picture groups.

[0021] In a possible implementation manner, generating an inference picture group according to any current frame includes:

[0022] Determine a first timestamp of the core false alarm picture and a second timestamp of any current frame;

[0023] determining a time difference between the first timestamp and the second timestamp;

[0024] Obtaining the timestamp of each of the associated pictures;

[0025] Determine the acquisition time point corresponding to each associated picture according to the timestamp of each associated picture and the time difference;

[0026] Extracting each picture in the monitoring video stream as each derivative picture according to each acquisition time point;

[0027] The inference picture group is generated according to any one current frame and each of the derived pictures.

[0028] In a possible implementation, the comparing the inference picture group with the any picture group in similarity to generate a first similarity value between the any current frame and the any picture group includes:

[0029] Determine a first similarity between any current frame and the core false positive picture;

[0030] Determine a second similarity between each of the associated pictures and the derived pictures corresponding to each of the associated pictures;

[0031] A first similarity value between any current frame and any picture group is generated according to the first similarity and each of the second similarities.

[0032] In a possible implementation, extracting frames from the surveillance video stream based on the acquired algorithm configuration information corresponding to the target camera to obtain multiple current frames includes:

[0033] Decapsulate the video transmission protocol to generate the monitoring video stream in h264 or h265 format;

[0034] Decoding the surveillance video stream to obtain picture frame information in RGB color space or YUV color space;

[0035] Extract frame data regularly according to the acquired frame extraction interval;

[0036] Performing a scaling operation on the frame data to obtain a set resolution image;

[0037] The image after the scaling operation is encoded to obtain multiple current frames.

[0038] In a possible implementation, the non-similar frame is input into a target algorithm model corresponding to the algorithm configuration information to generate a recognition result, and then the method further includes:

[0039] If the recognition result indicates an abnormality, the abnormal recognition result is sent to a cloud server and stored;

[0040] If the number of the abnormal identification results is greater than or equal to one, searching for a terminal to be communicated with, the terminal to be communicated with being a terminal capable of communicating with an edge device;

[0041] If the terminal to be communicated is detected within the set sensing range, the abnormal identification result is sent to the terminal to be communicated.

[0042] In a possible implementation manner, the sending of the abnormality identification result to the terminal to be communicated further includes:

[0043] If a reply instruction based on any of the abnormal identification results is obtained from the communication terminal, any of the stored abnormal identification results is deleted and a processing identifier is generated;

[0044] The processing identifier is sent to the cloud server, and any abnormal identification result in the cloud server is marked according to the processing identifier.

[0045] In the second aspect, the present application provides a video analysis device based on edge-cloud collaboration, which adopts the following technical solution:

[0046] A video analysis device based on edge-cloud collaboration, the device comprising:

[0047] An acquisition module is used to acquire the surveillance video stream collected by the target camera;

[0048] A frame extraction module is used to extract frames from the monitoring video stream based on the acquired algorithm configuration information corresponding to the target camera to obtain multiple current frames;

[0049] A comparison module, used to compare the similarity of each current frame with the acquired false warning picture set, and generate a similarity value corresponding to each current frame;

[0050] A screening module, used to compare each similarity value with a false alarm threshold, and to take pictures with similarity values ​​greater than or equal to the false alarm threshold as similar frames, and to take pictures with similarity values ​​less than the false alarm threshold as non-similar frames;

[0051] The recognition module is used to input the non-similar frames into the intelligent algorithm model corresponding to the algorithm configuration information, generate recognition results, and filter the similar frames.

[0052] In a possible implementation, when the comparison module compares the similarity between any current frame and the acquired false warning picture set and generates a similarity value corresponding to any current frame, it is specifically used to:

[0053] The false alarm picture set includes at least one picture group, and the picture group includes marked core false alarm pictures;

[0054] Compare the similarity between any current frame and each of the picture groups respectively to generate respective first similarity values, and use the first similarity value with the largest value as the similarity value corresponding to any current frame;

[0055] The step of comparing the similarity between any current frame and any picture group to generate a first similarity value includes:

[0056] If the picture group includes associated pictures related to the core false alarm picture, generating an inference picture group according to any current frame, wherein the inference picture group includes derivative pictures corresponding to each of the associated pictures;

[0057] The inference picture group is compared with any one of the picture groups in terms of similarity to generate a first similarity value between any one of the current frames and any one of the picture groups.

[0058] In a possible implementation, when the comparison module generates the inference picture group according to any current frame, it is specifically used to:

[0059] Determine a first timestamp of the core false alarm picture and a second timestamp of any current frame;

[0060] determining a time difference between the first timestamp and the second timestamp;

[0061] Obtaining the timestamp of each of the associated pictures;

[0062] Determine the acquisition time point corresponding to each associated picture according to the timestamp of each associated picture and the time difference;

[0063] Extracting each picture in the monitoring video stream as each derivative picture according to each acquisition time point;

[0064] The inference picture group is generated according to any one current frame and each of the derived pictures.

[0065] In a possible implementation, when the comparison module compares the inference picture group with the any picture group for similarity and generates a first similarity value between the any current frame and the any picture group, it is specifically used to:

[0066] Determine a first similarity between any current frame and the core false positive picture;

[0067] Determine a second similarity between each of the associated pictures and the derived pictures corresponding to each of the associated pictures;

[0068] A first similarity value between any current frame and any picture group is generated according to the first similarity and each of the second similarities.

[0069] In a possible implementation, when the frame extraction module extracts frames from the monitoring video stream based on the acquired algorithm configuration information corresponding to the target camera to obtain multiple current frames, it is specifically used to:

[0070] Decapsulate the video transmission protocol to generate the monitoring video stream in h264 or h265 format;

[0071] Decoding the surveillance video stream to obtain picture frame information in RGB color space or YUV color space;

[0072] Extract frame data regularly according to the acquired frame extraction interval;

[0073] Performing a scaling operation on the frame data to obtain a set resolution image;

[0074] The image after the scaling operation is encoded to obtain multiple current frames.

[0075] In a possible implementation, the analysis device further includes a connection module, which is used to input the non-similar frame into a target algorithm model corresponding to the algorithm configuration information to generate a recognition result, and when the recognition result represents an abnormality, send the abnormal recognition result to a cloud server and store the abnormal recognition result;

[0076] If the number of the abnormal identification results is greater than or equal to one, searching for a terminal to be communicated with, the terminal to be communicated with being a terminal capable of communicating with an edge device;

[0077] If the terminal to be communicated is detected within the set sensing range, the abnormal identification result is sent to the terminal to be communicated.

[0078] In a possible implementation, after sending the abnormality identification result to the terminal to be communicated, the connection module is specifically used to:

[0079] If a reply instruction based on any of the abnormal identification results is obtained from the communication terminal, any of the stored abnormal identification results is deleted and a processing identifier is generated;

[0080] The processing identifier is sent to the cloud server, and any abnormal identification result in the cloud server is marked according to the processing identifier.

[0081] In a third aspect, the present application provides an electronic device, which adopts the following technical solution:

[0082] An electronic device, comprising:

[0083] at least one processor;

[0084] Memory;

[0085] At least one application, wherein at least one application is stored in a memory and configured to be executed by at least one processor, and the at least one application is configured to: execute the above-mentioned video analysis method based on edge-cloud collaboration.

[0086] In a fourth aspect, the present application provides a computer-readable storage medium, which adopts the following technical solution:

[0087] A computer-readable storage medium, comprising: storing a computer program that can be loaded by a processor and execute the above-mentioned video analysis method based on edge-cloud collaboration.

[0088] In summary, this application includes the following beneficial technical effects:

[0089] Similar frames are filtered to prevent the intelligent algorithm model from identifying the image as a warning image after similar frames are input into the intelligent algorithm model, causing a false warning. Non-similar frames can be input into the intelligent algorithm model for further identification. Each frame extracted from the current monitoring video stream is filtered by similarity comparison based on the false warning images stored in the past, and images that may cause false warnings are removed. This can compensate for the false warning rate caused by the empirical confidence threshold, thereby improving the warning accuracy of the early warning system. BRIEF DESCRIPTION OF THE DRAWINGS

[0090] Figure 1 It is a hardware schematic diagram of a monitoring video stream analysis system based on edge-cloud collaboration according to an embodiment of the present application;

[0091] Figure 2 It is a hardware schematic diagram of a monitoring video stream analysis system based on edge-cloud collaboration according to an embodiment of the present application;

[0092] Figure 3 It is a hardware schematic diagram of a monitoring video stream analysis system based on edge-cloud collaboration according to an embodiment of the present application;

[0093] Figure 4 It is a flowchart of a video analysis method based on edge-cloud collaboration in an embodiment of the present application;

[0094] Figure 5 It is a block diagram of a video analysis device based on edge-cloud collaboration according to an embodiment of the present application;

[0095] Figure 6 It is a schematic diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0096] The following is combined with Figure 1-6 This application is described in further detail.

[0097] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0098] Reference Figure 1, the embodiment of the present application provides a monitoring video stream analysis system based on edge-cloud collaboration, including a cloud, multiple edge devices and cameras corresponding to the edge devices. The edge devices support mainstream video stream access protocols such as RTSP, are not picky about camera brands and resolutions, and can connect to cameras with one click; users can configure corresponding algorithms for each edge device through a cloud server according to scene requirements to implement the configuration of camera execution algorithms. Each camera can be configured with multiple algorithms at the same time, and each algorithm can be configured to multiple cameras at the same time; the edge device is a device that provides a core network entry point to the model service, which is used to collect different pictures according to the algorithm of the early warning management system, and perform model reasoning, filter out unconcerned information, submit concerned information to the cloud service of the early warning system layer, and display the early warning information through the cloud server. Since there is no need to use large GPU servers and other equipment, the platform has a wider range of applicable scenarios, and has high efficiency in some harsh environments where large servers cannot be maintained, and reduces network bandwidth limitations and improves system processing capabilities.

[0099] Reference Figure 1 In order to facilitate the management of algorithm models, cameras, edge devices and warning events, an early warning management system is configured on the cloud server provided in the embodiment of the present application. The early warning management system includes an early warning management module, a configuration management module and a system management module.

[0100] The early warning management module supports viewing the processing status of all early warning information. According to the processing status, it is divided into three states: unconfirmed, confirmed and false warning; the early warning information includes time, organization, event type, device name, snapshot, status, operation and other information. Click the snapshot to view the details, and support operations such as zooming in, zooming out and rotating the picture. Click the operation button to process the early warning information. The status can be divided into confirmed and false warning, and batch operation of early warning information is supported. The early warning management module provides retrieval functions according to the warning time, organization, event type and device name. The early warning management module also provides statistical functions, provides a carousel function for early warning events, supports rapid processing of early warning information, supports querying early warning records by time series, and can output early warning statistical charts, which is convenient for users to manage long-term early warning logs, analyze dangerous events, and achieve early prevention and disposal.

[0101] Reference Figure 1 The configuration management module includes: edge device management and camera management. Among them, the camera information is managed through the camera management list, which supports camera addition, editing, and algorithm configuration operations. Before configuring the algorithm, the camera needs to be associated with a specific edge device so that the camera can be enabled through the edge device; the basic information of the camera is managed through the camera list, which displays the organization where the camera is located, the camera name, the installation location, and the associated edge device.

[0102] The basic information of the camera can be found in the following table, Table 1:

[0103] Table 1

[0104] parameter illustrate organize Camera Organization Device Name Camera Name Video stream address Support RTSP, HTTP, RTMP, etc. Installation location Camera installation location Edge Devices Edge devices associated with cameras Image size The resolution of the image after encoding Frame extraction interval The time interval for extracting frames from the video stream Frame extraction method Edge device frame extraction mode selection

[0105] The edge device management function supports the addition, deletion, query, modification, and startup and disable operations of edge devices, and supports the upgrade and startup of models. When configuring the model for the edge device, you can configure the added general algorithm model for it through the edge device list operation column. The following table 2 shows several algorithm models:

[0106] Table 2

[0107] Algorithm model name Algorithm Description Face Recognition It can be used to detect facial information appearing in the scene. When a face not in the whitelist is detected, an alarm will be triggered; when a face in the blacklist is detected, an alarm will be triggered; helmet It can be used to detect whether people in the scene are wearing helmets; Smoking It can be used to detect whether people in the scene are smoking; smoke Can be used to detect whether smoke appears in the scene; Leaving the job It can be used to detect whether someone has left their work station in the scene; flame It can be used to detect whether there is flame in the scene; Detention It can be used to detect whether there is stagnation in the scene; Break-in It can be used to detect whether there is a person breaking into the scene; Gathering of people It can be used to detect whether there are people gathering in the scene; Make a call Can be used to detect whether there is a call in the scene

[0108] After the algorithm model is sent to the edge device, the algorithm model corresponding to each camera can be configured specifically for the application scenario of each camera. The camera algorithm configuration information is shown in Table 3 below:

[0109] Table 3:

[0110] parameter illustrate algorithm Select the model algorithm that has been started in the associated edge device for configuration Start time period Set the time period for edge device analysis. Multiple time periods can be configured. Warning interval Warning interval Confidence Threshold Determine whether the detection result is a warning event, and adaptively adjust the optimal warning threshold in the specific environment Analysis locale Use polygonal annotation to mark the analysis area of ​​interest to the user

[0111] Through the above configuration, after cameras in different areas are configured with the same algorithm model, the security level can be adjusted through different confidence thresholds, and targeted attention can be given to the area of ​​each camera.

[0112] Reference Figure 1 and Figure 2 ,The management module of the system early warning management system provides ,submodules such as users, organizations, models, data, and services, realizing the ,management capabilities of personnel at different security levels. ,This module provides monitoring logs for platform services to facilitate ,rapid location of system problems and service recovery.

[0113] (1) User management supports user creation, approval, user disabling / enabling, department setting and other user management functions, and provides administrator and operator users. Administrators can delete operators, but operators do not have the function of adding new users.

[0114] (2) Organization management supports adding organizations for users and viewing all personnel in the organization.

[0115] (3) Model management supports the addition, configuration, and editing of models, and supports online editing of model information. It supports online configuration of basic information of the algorithm, including the warning content that needs to be displayed on the picture, and the status of the algorithm.

[0116] (4) Data management supports users to configure online the warning images and non-warning images that need to be cleaned up in the system, as well as the monitoring service logs and collection service logs transmitted by smart devices.

[0117] (5) The service version supports online upgrading of the software version of smart devices and updating of the device's collection service and model service.

[0118] Specifically, refer to Figure 3 ,The analysis process of the surveillance video stream analysis system based on edge-cloud collaboration is as follows Figure 3 As shown:

[0119] (1) The user logs in to the intelligent early warning system to view the configuration of edge devices, equipment, models, etc.

[0120] (2) After the acquisition service on the edge device is running, it requests the configuration of the camera on the intelligent analysis system at regular intervals and caches the configuration information locally on the edge device.

[0121] (3) The model service on the edge device starts the corresponding model based on the model configuration information.

[0122] (4) The acquisition service on the edge device periodically captures images from the video surveillance camera device according to the frame extraction interval of the device configuration information. According to the detection interval of the algorithm configuration information, the captured images are periodically called to the configured model service for inference. Before inference, it is necessary to perform image similarity detection based on different algorithm configurations and false warning images, and determine whether algorithm inference is needed based on the detection results (the algorithms currently supported by the system include helmet detection, smoking detection, smoke detection, flame detection, intrusion detection, leaving the post detection, face recognition, making phone calls, detention, and crowd gathering); upload the inferred warning information to the warning service.

[0123] (5) After the warning service receives the image and related information uploaded by the edge device acquisition service, it saves the image to the server and stores the warning information in the database.

[0124] (6) Users can query and process relevant warning information on the warning service platform.

[0125] After the hardware system is deployed, in order to further solve the problem of reducing the false alarm rate of setting the confidence threshold based on experience, the embodiment of the present application further provides a video analysis method based on edge-cloud collaboration, which is executed by any edge device in the above-mentioned monitoring video stream analysis system based on edge-cloud collaboration, and the method includes:

[0126] Step S10: Acquire the monitoring video stream collected by the target camera.

[0127] Specifically, each edge device is connected to at least one camera, each camera corresponds to its own monitoring area, and each monitoring area corresponds to a monitoring video stream; the target camera is any camera connected to the edge device.

[0128] Step S20: extract frames from the monitoring video stream based on the acquired algorithm configuration information corresponding to the target camera to obtain multiple current frames.

[0129] Specifically, each edge device can be configured with multiple intelligent algorithm models. For the same intelligent algorithm model, when it is configured in different cameras, the corresponding algorithm configuration information may be different. The algorithm configuration information includes the configuration information in Table 3.

[0130] Each intelligent algorithm model is trained on the cloud server. Then, the administrator downloads the trained intelligent algorithm model into an offline model and sends it to the edge device through the early warning management system on the cloud server. The edge device configured with the offline model can perform collection service (i), monitoring service (ii), model service (iii) and log service (iv).

[0131] (i) The acquisition service is a software service responsible for acquiring the video stream of the network camera and managing the logic of image reasoning, warning information push, etc. The acquisition service is divided into a frame extraction module and an inference call module.

[0132] The frame extraction module extracts frames from the monitoring video stream based on the algorithm configuration information corresponding to the acquired target camera to obtain multiple current frames, including: step Sa1 (not shown in the figure), decapsulating the video transmission protocol to generate a monitoring video stream in h264 or h265 format; decoding the monitoring video stream to obtain image frame information in RGB color space or YUV color space; step Sa2 (not shown in the figure), extracting frame data at regular intervals according to the acquired frame extraction interval; step Sa3 (not shown in the figure), performing a scaling operation on the frame data to obtain a set resolution image; step Sa4 (not shown in the figure), encoding the image after the scaling operation to obtain multiple current frames.

[0133] Specifically, first decapsulate the rtsp or http video transmission protocol to obtain the h264 or h265 encoded video stream, then decode the video stream to obtain 25 to 30 frames per second of image frame information in the RGB color space or YUV color space, then extract the frame data according to the frame extraction interval, first perform a resize operation, select a suitable interpolation method to obtain a specified resolution image, and then perform JPEG encoding on the resized image to obtain a JPEG lossy compressed image. The generated image file can be controlled by controlling the JPEG compression level (greater than 0 and less than or equal to 100) to control the balance between bandwidth occupancy and inference effect.

[0134] When the edge device has different manufacturers / models, the configuration process of the frame extraction module can be debugged according to the needs, specifically:

[0135] Based on the jetson edge device, use ffmpeg to decapsulate the video stream to obtain h264 or h265 image data frame by frame, then use NVDEC hardware to decode the compressed current frame data to obtain NV12 format image information, then resize it according to the specified resolution in the NV12 format space, and finally use NVJPG hardware to generate JPEG image information.

[0136] Based on the atlas compilation device, ffmpeg is used to decapsulate the video stream to obtain h264 or h265 image data frame by frame, and then the dvp hardware video decoding module is used to decode the compressed current frame data to obtain the image information in yuv420 format, and then resize it according to the specified resolution in the yuv format space, and finally use the dvpp hardware jpeg image encoding module to generate jpeg image information. Therefore, when the traditional monitoring system is intelligently transformed, the model of the edge device is first identified, and then the device model is configured separately.

[0137] The frame extraction module will store all jpeg image information in the memory in an updated manner and map it to the camera ID. The main purpose of storing it in the memory is to reduce IO consumption, increase image processing speed, and reduce latency. However, it will also increase some memory usage, and the main consideration is to trade space for time.

[0138] The inference call module obtains the latest jpeg image information of the specified camera, and then calls the corresponding model according to the algorithm configured for the camera and the frame interval period. The relationship between the model and the algorithm is configured in the model configuration management. All target detection information returned after the call is stored in the memory, and each algorithm can filter out effective warning information from the detection result information according to its own algorithm logic. This method can reduce the consumption of computing power, and each algorithm does not need to be called once, thereby consuming more computing power.

[0139] (ii) The monitoring service is responsible for collecting device status (GPU, CPU utilization and temperature, memory occupancy), monitoring the running status of the collection service and model service, and is responsible for the software version update of the collection service and model service. Therefore, it is divided into three modules: collection device status module, service status monitoring module, and service software update module.

[0140] The underlying layer of the device status acquisition module depends on the monitoring tools provided by the specific device. Jetson uses the jtop tool to periodically acquire device status, and atlas200 uses the npu-smi tool to periodically acquire device status.

[0141] The service status monitoring model will periodically detect and collect service and model service heartbeat packets based on TCP communication. If no heartbeat packet is detected for 3 beats, the corresponding service is considered abnormal. If the heartbeat packet is received again after a short abnormality, the corresponding service is considered to have returned to normal.

[0142] The service software update module uses the SCP protocol to pull the update package that has been placed in the cloud. After verifying the update package through md5 to ensure that the downloaded data is correct, it will be unzipped and installed, and the collection service and model service will be restarted.

[0143] (iii) Model service is responsible for managing the software service of starting, stopping and upgrading the algorithm model. Each algorithm model provides an interface for receiving image data and returns the prediction results to the caller. Each algorithm model provides HTTP service based on the mongoose library. The interface routing is only provided to the collection service for inference calls and the monitoring service for monitoring model status. No service interface is provided to the outside world.

[0144] Based on the Jetson edge device model, TensorRT model reasoning acceleration optimization is provided. By using layer fusion and quantization technology, the plan model under fp16 precision will have a 2 to 3 times faster reasoning speed than the original model, which greatly improves the throughput of model reasoning and reduces latency.

[0145] The edge device model based on Atlas 200 provides the ATC tool to convert other framework models into OM offline format models, and uses the aipp tool to modify the model input into YUV format so that the DVPP hardware can be used for accelerated processing in the image preprocessing stage.

[0146] During the model upgrade process, use scp to obtain the corresponding version of the algorithm model file from the cloud, verify the integrity of the model file through md5, and then restart the corresponding model service.

[0147] (iv) The log service is responsible for uploading the log files generated by the edge device collection service, monitoring service, and model service, so that when the device fails, the time and cause of the problem can be checked and analyzed. The collection service, monitoring service, and model service will periodically update their own log files, and the log service will send the generated log files to the cloud without retaining them. Since the log file is a 24-hour operation log, it takes up a lot of space for long-term operation. The disk of the edge device is generally small, so long-term logs are not retained on the edge device disk.

[0148] All services use libcurl library as HTTP client, and periodically submit access requests to the warning system layer in POST mode, and obtain relevant configuration information, such as model configuration, algorithm configuration, camera configuration, etc. One-way data flow ensures data security and privacy issues of edge devices, and edge devices do not provide active access methods. All services are self-started based on systemctld mode to ensure that our edge analysis layer software services can automatically recover after the service stops due to abnormal device hardware or system status.

[0149] When setting the algorithm configuration information in Table 3 for each camera, the confidence threshold of the intelligent algorithm model is set according to the experience of the management personnel, so there may be errors and constant debugging is required. In order to reduce the false alarms caused by the confidence threshold, the edge device set in the embodiment of the present application can record the events marked as "false alarms" by the management personnel and sent by itself, that is, the false alarm pictures that may exist in the subsequent monitoring video stream are screened according to the false alarm pictures corresponding to the false alarms to reduce the false alarm rate.

[0150] Step S30: perform a similarity comparison between each current frame and the acquired false alarm picture set to generate a similarity value corresponding to each current frame.

[0151] Among them, the acquired false alarm picture set is the false alarm picture set corresponding to the image sent by the edge device to the cloud server and the result of the processing of the image by the cloud server is a false alarm. Therefore, the false alarm picture set on each edge device comes from the images that the edge device has sent to the cloud server in history. In addition, the false alarm picture set contains at least one frame of false alarm image marked by the management personnel. When the user marks a warning event as a false alarm on the warning management system of the cloud server, one or more frames of images corresponding to the warning event are automatically added to the false alarm picture set.

[0152] Specifically, content-based similarity calculation between images is used in many scenarios, such as image clustering, image retrieval, and image-based personalized recommendation. The pHash algorithm can be used to obtain the hash strings of two images respectively, and then the similarity of the hash strings of the two images can be compared to determine whether the two images are similar.

[0153] Step S40: Compare each similarity value with the false alarm threshold, and regard pictures with similarity values ​​greater than or equal to the false alarm threshold as similar frames, and regard pictures with similarity values ​​less than the false alarm threshold as non-similar frames.

[0154] Step S50: input the non-similar frames into the intelligent algorithm model corresponding to the algorithm configuration information, generate recognition results, and filter the similar frames.

[0155] An embodiment of the present application provides a video analysis method based on edge-cloud collaboration, which filters similar frames to prevent the intelligent algorithm model from identifying the image as a warning image after similar frames are input into the intelligent algorithm model, causing a false warning. As for non-similar frames, they can be input into the intelligent algorithm model for further identification. Each frame of the image extracted from the current monitoring video stream is filtered by performing a similarity comparison based on the false warning images stored in the past, and images that may cause false warnings are removed. This can compensate for the false warning rate caused by the empirical confidence threshold, thereby improving the warning accuracy of the warning system.

[0156] Furthermore, according to the scene requirements, managers can select intelligent algorithm models with different functions to identify abnormal events. When the abnormal event is associated with a dynamic target, it is necessary to select a model that analyzes multiple consecutive frames of images to identify the dynamic target. For example, when monitoring the sorting process of express delivery, it is very important to monitor in real time whether the sorting personnel are performing illegal actions or violent sorting. During the analysis process, the express delivery itself can be used as a dynamic target, and the moving path of the express delivery can be analyzed to indirectly identify whether the staff is performing violent sorting. The posture changes of the sorting personnel in each frame can also be analyzed to directly identify whether the sorting personnel are performing violent sorting. In this case, the sorting personnel are used as dynamic targets.

[0157] When the abnormal event is associated with a static target, you can choose an intelligent algorithm model that uses a single-frame image for identification, such as identifying the target object / target posture in a single-frame image. For example, to monitor illegal operations, including whether smoking is illegal in the venue, whether not wearing a helmet is illegal in the venue, whether making phone calls is illegal in the venue, etc., you can judge whether there is an illegal operation by judging the posture of the target object in a certain frame of the image, or the relative position relationship between the target object (human body) and the target object (object).

[0158] Identifying static targets can meet the needs of most monitoring scenarios, but in some complex scenarios, the results obtained from analyzing a single frame of an image may still have a high false alarm rate; for example: Scenario 1: In a specific area, you cannot make or receive phone calls. At this time, when training the intelligent algorithm model, the input sample set may contain sample sets of people making or receiving phone calls at various angles captured by the camera. In the sample set, a person may put the phone to his or her ear to answer a call, or put the phone in front of his or her face to make a video call. At this time, there may be a false alarm: the user puts the phone in front of him or her to check the time or perform other operations on the phone. At this time, the single frame of the image captured is similar to the sample set corresponding to the intelligent algorithm model. At this time, although the user did not make or receive a call, it caused a false alarm of the system. For example, in scenario 2, in a restaurant kitchen, the target object (person) is not allowed to eat, but there may be a false alarm of eating when the target object is too close to the food when observing the color and smell of the dish.

[0159] Take scenario 2 as an example to illustrate the specific process of false alarms: In scenario 2, select a set of false alarm pictures a that are identified as non-violation operations by the intelligent algorithm model, as shown in Table 4:

[0160] Table 4

[0161] Image number (in chronological order) Image Content Image a0 The target person and the target food are both present in the same picture; Image a1 The distance between the target person and the target food is less than the first distance value; Image a2 The distance between the target person and the target food is less than the second distance value (the second distance value is greater than the first distance value); Image a3 The distance between the target person and the target food is greater than the second distance value, and the target person's mouth movement is v1; Image a4 The distance between the target person and the target food is greater than the second distance value, and the target person's mouth movement is v2; Image a5 The distance between the target person and the target food is greater than the second distance value, and the target person's mouth movement is v3;

[0162] In scenario 2, a set of pictures b corresponding to illegal operations identified by the intelligent algorithm model is selected, as shown in Table 5 below:

[0163] Table 5

[0164] Image number (in chronological order) Image Content Image b0 The target person and the target food are both present in the same picture; Image b1 The distance between the target person and the target food is less than the first distance value; Image b2 The distance between the target person and the target food is less than the second distance value (the second distance value is greater than the first distance value); Image b3 The distance between the target person and the target food is greater than the second distance value, and the target person's mouth movement is v1; Image b4 The distance between the target person and the target food is greater than the second distance value, and the target person's mouth movement is v2; Image b5 The distance between the target person and the target food is greater than the second distance value, and the target person's mouth movement is v3;

[0165] Referring to Tables 4 and 5, the similarity between picture a3 in picture set a and picture b3 in picture set b is greater than the second threshold value. At this time, whether it is identified as a warning or a non-warning mainly depends on whether the previous and next frames of picture a3 are analyzed, and whether the previous and next frames of picture b3 are analyzed; specifically, if the previous and next frames of picture a3 are analyzed, the result corresponding to picture set a is a routine operation, and no warning is given; if the previous and next frames of picture b3 are analyzed, the result corresponding to picture set b is an unconventional operation, and a warning is given; however, if only picture a3 in picture set a and picture b3 in picture set b are analyzed, then a warning may be given or no warning may be given, which may easily lead to false warnings.

[0166] Therefore, in a more complex scene, if any current frame extracted in step S20 is similar to picture a3, it is still possible to be similar to picture b3, so further judgment is required. If a method of real-time tracking of dynamic targets is adopted, the recognition accuracy may be improved to a certain extent, but when there are multiple dynamic targets in the scene, the analysis of multiple frames of pictures continuously brings a huge amount of calculation, and it is difficult to ensure the real-time nature of the warning.

[0167] Therefore, in order to improve the accuracy of target recognition in complex scene recognition while reducing the amount of calculation and improving the efficiency of calculation; in one embodiment of the present application, in step S30: a similarity comparison is performed between any current frame and the acquired false warning picture set, and a similarity value corresponding to each current frame is generated, including:

[0168] Step S301 (not shown in the figure), the false alarm picture set includes at least one picture group, and the picture group includes a marked core false alarm picture.

[0169] Among them, each image group corresponds to different false alarm events identified by the same intelligent algorithm model on the same edge device. Each false alarm event includes at least one core false alarm image marked by the management personnel. The core false alarm image represents that after the image is input into the intelligent algorithm model, the output result of the intelligent algorithm model is a warning / abnormal image.

[0170] Step S302 (not shown in the figure), compare the similarity between any current frame and each picture group respectively, generate respective first similarity values, and take the first similarity value with the largest value as the similarity value corresponding to any current frame.

[0171] Among them, the content in any current frame is compared with each picture group (historical false warning events) one by one, and the first similarity value with the largest value is selected to determine the maximum probability of a false warning, and the maximum probability is used as the similarity between any current frame and the false warning event.

[0172] In one embodiment of the present application, in step S302, any current frame is compared with any picture group for similarity, including: step S3021 (not shown in the figure), if the picture group contains associated pictures related to the core false alarm picture, then an inference picture group is generated according to any current frame, and the inference picture group contains derivative pictures corresponding to each associated picture. Step S3022 (not shown in the figure), the inference picture group is compared with any picture group for similarity, and a first similarity value between any current frame and any picture group is generated.

[0173] The associated images are several frames of images marked by the management personnel that are associated with the core false alarm images in time series. Specifically, the operation page of the management system prompts the management personnel to mark an image as a core false alarm image, and the core false alarm image represents the key image identified as a warning event.

[0174] When marking false alarm events, managers can analyze the causes of false alarms. If managers determine that the cause of the false alarm is only because of occlusion and other factors in the single-frame image, the managers only mark the core false alarm image of the single frame, and there is no need to mark the associated images; if managers determine that the cause of the false alarm is that the intelligent algorithm model does not conduct a comprehensive analysis of the frames before / after the core false alarm image, multiple frames before / after the core false alarm image can be extracted as associated images according to the time period, or multiple frames before / after the core false alarm image can be extracted as associated images based on experience; therefore, each image group contains at least the core false alarm image, and may contain associated images associated with the core false alarm image.

[0175] Among them, step S3021: if the picture group contains associated pictures related to the core false alarm picture, the method of generating the inference picture group according to any current frame can be: (1) generating the inference picture group according to the number of frames in the picture group and the position of the core false alarm picture in the picture group, specifically including: according to the correspondence between the core false alarm picture and any current frame, each associated picture corresponds to each derived picture, and generating the inference picture group, for example: picture group: {associated picture 1, associated picture 2, associated picture 3, core false alarm picture, associated picture 4, associated picture 5}; inference picture group: {derivative picture 1, derived picture 2, derived picture 3, any current frame, derived picture 4, derived picture 5}. (2) It can also be extracted according to the time period. Generally, for the same intelligent algorithm model and the same camera, the frame extraction interval is the same. At this time, the method of extracting according to the time period is equivalent to method 1; if the frame extraction interval corresponding to the same intelligent algorithm model and the same camera changes, method 2 is different from method 1.

[0176] Therefore, in one embodiment of the present application, step S3021 (not shown in the figure), generating an inference picture group according to any current frame, includes:

[0177] Step Sb1 (not shown in the figure), determining the first timestamp of the core false alarm picture and the second timestamp of any current frame.

[0178] Step Sb2 (not shown in the figure), determining the time difference between the first timestamp and the second timestamp.

[0179] Step Sb3 (not shown in the figure), obtaining the timestamp of each associated picture.

[0180] Step Sb4 (not shown in the figure), determining the collection time point corresponding to each associated picture according to the timestamp and time difference of each associated picture.

[0181] Step Sb5 (not shown in the figure), extracting each picture in the monitoring video stream as each derivative picture according to each acquisition time point.

[0182] Step Sb6 (not shown in the figure): generate an inference picture group according to any current frame and each derived picture.

[0183] Specifically, each frame of the extracted image carries a unique timestamp. If the timestamp of the core false alarm image marked by the management personnel is 00:10:05 on XX day (00:10:05 on XX day), continuing with the above example, the image group with timestamp is: image group {associated image 1 (00:09:50 on XX day), associated image 2 (00:09:55 on XX day), associated image 3 (00:10:00 on XX day), core false alarm image (00:10:05 on XX day), associated image 4 (00:10:10 on XX day), associated image 5 (00:10:15 on XX day)};

[0184] If the time of any current frame is XX+1 day 00:10:00 (XX+1 day 00:10:5 seconds), then the time difference = XX+1 day - XX day, and the corresponding inference picture group with timestamp is: inference picture group {derivative picture 1 (XX+1 day 00:09:50 seconds), derivative picture 2 (XX+1 day 00:09:55 seconds), derivative picture 3 (XX+1 day 00:10:0 seconds), any current frame (XX+1 day 00:10:5 seconds), derivative picture 4 (XX+1 day 00:10:10 seconds), derivative picture 5 (XX+1 day 00:10:15 seconds)}. It can be seen that the time difference between the timestamp of each associated picture and the timestamp of its corresponding derivative picture is equal to the time difference.

[0185] Furthermore, in one embodiment of the present application, it also includes: if the identification object corresponding to the intelligent algorithm model is a dynamic target and the picture group contains core pictures and associated pictures, then identifying the first moving speed of the dynamic object in the picture group; identifying the second moving speed of the dynamic object in the inference picture group; and adjusting the number of frames of the inference picture group according to the first moving speed and the second moving speed.

[0186] Specifically, the second moving speed / the first moving speed=the total number of frames of the picture group / the total number of frames of the inference picture group (wherein the number of frames is a positive integer obtained by a rounding function).

[0187] Among them, the first moving speed corresponds to the first dynamic object in the picture group, and the second moving speed corresponds to the second dynamic object in the inference picture group. The first dynamic object and the second dynamic object are the same object that the intelligent algorithm model needs to identify. The first / second is just for distinguishing in different video streams, that is, the dynamic object is a person, and is not distinguished based on the first person or the second person.

[0188] When different people perform the same set of actions (action a+action b+action c) in different scenes, if the first moving speed is greater than the second moving speed, only the first frame a, the second frame a, and the third frame a of the first dynamic object need to be extracted; while for the second dynamic object, in order to capture and complete the same set of actions, the first frame b, the second frame b, the third frame c, the fourth frame d, and the fifth frame d need to be extracted. Therefore, by adjusting the parameter of the target object speed, the generated inference picture group can include the false warning pictures in the picture group as much as possible according to the actual situation, thereby improving the accuracy of false warning recognition.

[0189] In one embodiment of the present application, in step S3022: the inference picture group is compared with any picture group for similarity, and a similarity value corresponding to any current frame is generated, including: step Sd1 (not shown in the figure), determining the first similarity between any current frame and the core false alarm picture; step Sd2 (not shown in the figure), determining the second similarity between each associated picture and the derivative picture corresponding to each associated picture; step Sd3 (not shown in the figure), generating a first similarity value corresponding to any current frame according to the first similarity and each second similarity. Specifically, the average of each first similarity and second similarity is calculated as the first similarity value corresponding to any current frame and any picture group.

[0190] In the above-mentioned scenes 1 and 2, when performing similarity comparison, the key area to be identified is the mouth shape of the target object. However, if the mouth shape of the target object changes, and other areas of the target object's body change significantly for the false warning picture, then when the two pictures are compared for similarity, the similarity value may be reduced, and the picture with a false warning may be input into the intelligent algorithm model again, which may cause a false warning again. In order to improve the accuracy of the picture when performing similarity comparison. Therefore, in one embodiment of the present application, step S3022 (not shown in the figure), the inference picture group is compared with any picture group for similarity, and the first similarity value between any current frame and any picture group is generated, including:

[0191] Step Sc1 (not shown in the figure), if there are non-core areas and core areas between any current frame and the core false alarm picture, then determine the non-core similarity value between the inference picture group and any picture group based on the non-core area, and determine the core similarity value between the inference picture group and any picture group based on the core area.

[0192] The core area is the key area where there is a difference between the false warning picture and the normal warning picture. The core area is the area marked by the user on the false warning picture when marking the false warning event. For example, between pictures a3 and b3, the action where the distance between the user's face and the food is greater than the second threshold is the non-core area, and the user's mouth action is the core area marked by the manager.

[0193] If the user does not mark the area, it may be because the false alarm was only caused by factors such as occlusion when the user inferred the event. At this time, the user can classify the event as an occlusion factor again; if the false alarm picture is indeed caused by the reasons in the above scenario 1 or scenario 2, then if the user does not mark it, the management system will pop up a window to prompt the user to mark the core area.

[0194] Step Sc2 (not shown in the figure) generates a first similarity value between any current frame and any picture group based on the core similarity value, the first weight value corresponding to the core similarity value, the non-core similarity value and the second weight value corresponding to the non-core similarity value.

[0195] Specifically, if the picture group includes a core false alarm picture and at least one associated picture, the calculation of each first similarity and each second similarity follows: Similarity = the first weight value corresponding to the picture * the corresponding core similarity value + the corresponding second weight value * the corresponding non-core similarity value; wherein the first weight value and the second weight value are both constants and can be set based on experience; by configuring the first weight value and the second weight value, the proportion of the similarity of the core area is increased, which can increase the focus on the core area when comparing the similarity of the inference picture group with the picture group.

[0196] By calculating the values ​​of each first similarity and second similarity according to the similarity calculation formula, and then calculating the average of all similarities according to the first similarity and the second similarity, the first similarity value between any current frame and any picture group can be determined.

[0197] If the picture group includes only core false positive pictures, the first similarity value between any current frame and any picture group=the first weight value*core similarity value+the second weight value*non-core similarity value.

[0198] Each edge device sends the abnormal identification result to the cloud server. The cloud server is responsible for managing multiple edge devices. There are a large number of warning events that need to be processed, and there may be situations where the management personnel cannot handle the warning events in time. Therefore, the video surveillance system configured in the embodiment of the present application also includes a terminal device that can be connected to the edge device. When the authorized object in the monitoring area carries the terminal device close to the corresponding sensing area where the edge device is located, the communication module configured on the edge device can detect the terminal device. At this time, the terminal device is a terminal to be communicated. When the terminal device is outside the sensing area, the terminal device is a non-communication terminal. If the terminal device is a terminal to be communicated, the abnormal identification result is sent to the terminal to be communicated, and the authorized object in the monitoring area can promptly handle the warning event through the terminal to be communicated.

[0199] In one embodiment of the present application, in step S50: inputting the non-similar frames into the intelligent algorithm model corresponding to the algorithm configuration information to generate a recognition result, and then further comprising:

[0200] Step Se1 (not shown in the figure): if the recognition result indicates an abnormality, the abnormal recognition result is sent to the cloud server and stored. Step Se2 (not shown in the figure): if the number of abnormal recognition results is greater than or equal to one, a terminal to be communicated is searched, and the terminal to be communicated is a terminal that can communicate with the edge device. Step Se3 (not shown in the figure): if the terminal to be communicated is detected within the set sensing range, the abnormal recognition result is sent to the terminal to be communicated.

[0201] Specifically, the communication module can use Bluetooth technology (Bluetooth sensor) to search for mobile terminals. The name of the matched mobile terminal has been recorded on the communication module, and there is no need to identify it again when it is used next time. Specifically, the communication module enters the standby state, and the Bluetooth sensor of the communication module transmits a low-frequency signal to search whether there is a Bluetooth module of a paired mobile terminal within the first preset range; if so, it is determined that the mobile terminal is within the close range of the communication module, and the search range of Bluetooth is less than or equal to 10 meters.

[0202] The communication module can also use NFC (Near Field Communication) to search for paired mobile terminals within a range of 1 meter to make a close-range determination. NFC technology is a short-range, high-frequency radio technology with two reading modes, active and passive. At a frequency of 13.56 MHz, the effective use distance is within 20 cm, and the transmission speed includes 106 Kbit / s, 212 Kbit / s or 424 Kbit / s. Currently, near-field communication has passed ISO / IECIS18092 international standards, ECMA-340 standards and ETSITS102190 standards.

[0203] That is, the terminal devices in the factory area and the communication modules on each edge device have been paired during configuration, and the terminal to be communicated is the terminal that enters the preset search range of the communication module. When searching for mobile terminals, the communication module automatically selects the paired mobile terminals. After the communication module and the mobile terminal are identified, the name of the matched mobile terminal is recorded on the communication module, and there is no need to identify it again when it is used next time.

[0204] After the communication terminal receives the abnormal identification result, the authorized object with processing authority can handle the abnormal event on site. Therefore, the authorized object with processing authority can handle the abnormal event in time as long as it carries the communication terminal, whether it is in the remote monitoring center or in the factory. In addition, in the absence of abnormal events, the communication module is in a closed state to save energy.

[0205] In one embodiment of the present application, step Se3 (not shown in the figure) sends the abnormal identification result to the terminal to be communicated if the terminal to be communicated is detected within the set sensing range, and then includes: step Se4 (not shown in the figure) if a reply instruction based on any abnormal identification result issued by the terminal to be communicated is obtained, any stored abnormal identification result is deleted and a processing identifier is generated; step Se5 (not shown in the figure) sends the processing identifier to the cloud server, and marks any abnormal identification result in the cloud server according to the processing identifier. If the communication module on the edge device receives a reply instruction issued by the authorized object through the terminal to be communicated, the reply instruction is a reply to the corresponding abnormal identification result. By marking the warning event that has been replied on the terminal to be communicated, it can avoid repeated processing of the authorized object at the cloud management platform.

[0206] Reference Figure 5 The above embodiment introduces a video analysis method based on edge-cloud collaboration from the perspective of method flow. The following embodiment introduces a video analysis device 100 based on edge-cloud collaboration from the perspective of a virtual module or a virtual unit. Please see the following embodiment for details.

[0207] A video analysis device 100 based on edge-cloud collaboration, the device comprising:

[0208] The acquisition module 1001 is used to acquire the monitoring video stream collected by the target camera;

[0209] The frame extraction module 1002 is used to extract frames from the monitoring video stream based on the algorithm configuration information corresponding to the acquired target camera to obtain multiple current frames;

[0210] The comparison module 1003 is used to compare the similarity of each current frame with the acquired false alarm picture set, and generate a similarity value corresponding to each current frame;

[0211] A screening module 1004 is used to compare each similarity value with the false alarm threshold, and to take pictures with similarity values ​​greater than or equal to the false alarm threshold as similar frames, and to take pictures with similarity values ​​less than the false alarm threshold as non-similar frames;

[0212] The recognition module 1005 is used to input the non-similar frames into the intelligent algorithm model corresponding to the algorithm configuration information, generate recognition results, and filter the similar frames.

[0213] In a possible implementation, when the comparison module 1003 performs a similarity comparison between any current frame and the acquired false alarm picture set to generate a similarity value corresponding to any current frame, it is specifically used to:

[0214] The false alarm picture set includes at least one picture group, and the picture group includes a marked core false alarm picture;

[0215] Compare the similarity between any current frame and each picture group respectively, generate respective first similarity values, and take the first similarity value with the largest value as the similarity value corresponding to any current frame;

[0216] The step of comparing the similarity between any current frame and any picture group to generate a first similarity value includes:

[0217] If the picture group includes associated pictures related to the core false positive picture, then generating an inference picture group according to any current frame, the inference picture group includes derivative pictures corresponding to each associated picture;

[0218] The inference picture group is compared with any picture group for similarity, and a first similarity value between any current frame and any picture group is generated.

[0219] In a possible implementation, when the comparison module 1003 generates an inference picture group according to any current frame, it is specifically used to:

[0220] Determine a first timestamp of a core false positive image and a second timestamp of any current frame;

[0221] determining a time difference between the first timestamp and the second timestamp;

[0222] Get the timestamp of each associated image;

[0223] Determine the collection time point corresponding to each associated image according to the timestamp and time difference of each associated image;

[0224] Extracting each picture in the monitoring video stream as each derivative picture according to each acquisition time point;

[0225] Generate an inference picture group based on any current frame and each derived picture.

[0226] In a possible implementation, when the comparison module 1003 compares the inference picture group with any picture group to generate a first similarity value between any current frame and any picture group, it is specifically used to:

[0227] Determine a first similarity between any current frame and a core false positive image;

[0228] Determine a second similarity between each associated image and each associated image's corresponding derivative image;

[0229] A first similarity value between any current frame and any picture group is generated according to the first similarity and each second similarity.

[0230] In a possible implementation, when the frame extraction module 1002 extracts frames from the surveillance video stream based on the acquired algorithm configuration information corresponding to the target camera to obtain multiple current frames, it is specifically used to:

[0231] Decapsulate the video transmission protocol and generate surveillance video streams in h264 or h265 format;

[0232] Decode the surveillance video stream to obtain picture frame information in RGB color space or YUV color space;

[0233] Extract frame data regularly according to the acquired frame extraction interval;

[0234] Perform scaling operations on the frame data to obtain a picture with a set resolution;

[0235] The image after the scaling operation is encoded to obtain multiple current frames.

[0236] In a possible implementation, the analysis device further includes a connection module, which is used to generate a recognition result when the non-similar frame is input into the target algorithm model corresponding to the algorithm configuration information, and when the recognition result represents an abnormality, the abnormal recognition result is sent to the cloud server, and the abnormal recognition result is stored;

[0237] If the number of abnormal identification results is greater than or equal to one, searching for a terminal to be communicated with, the terminal to be communicated with being a terminal that can communicate with the edge device;

[0238] If the terminal to be communicated is detected within the set sensing range, the abnormal recognition result is sent to the terminal to be communicated.

[0239] In a possible implementation, after sending the abnormal identification result to the terminal to be communicated, the connection module is specifically used to:

[0240] If a reply instruction based on any abnormal identification result is obtained from the communication terminal, any stored abnormal identification result is deleted and a processing identifier is generated;

[0241] The processing identifier is sent to the cloud server, and any abnormal identification result in the cloud server is marked according to the processing identifier.

[0242] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0243] The present application also introduces an electronic device from the perspective of a physical device, such as Figure 6 As shown, Figure 6 The electronic device 1100 shown includes: a processor 1101 and a memory 1103. The processor 1101 and the memory 1103 are connected, such as through a bus 1102. Optionally, the electronic device 1100 may also include a transceiver 1104. It should be noted that in actual applications, the transceiver 1104 is not limited to one, and the structure of the electronic device 1100 does not constitute a limitation on the embodiments of the present application.

[0244] Processor 1101 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 1101 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like.

[0245] The bus 1102 may include a path to transmit information between the above components. The bus 1102 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. The bus 1102 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.

[0246] The memory 1103 may be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disk storage (including compressed optical disk, laser disk, optical disk, digital versatile disk, Blu-ray disk, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0247] The memory 1103 is used to store the application code for executing the solution of the present application, and the execution is controlled by the processor 1101. The processor 1101 is used to execute the application code stored in the memory 1103 to implement the contents shown in the above method embodiment.

[0248] The electronic devices include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), and fixed terminals such as digital TVs, desktop computers, etc. It can also be a server, etc. Figure 6 The electronic device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0249] It should be understood that, although the steps in the flowchart of the accompanying drawings are displayed in sequence as indicated by the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a part of the sub-steps or stages of other steps.

[0250] The above are only some implementation methods of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A video analysis method based on edge-cloud collaboration, characterized in that: include: Get the surveillance video stream collected by the target camera; Extract frames from the surveillance video stream based on the acquired algorithm configuration information corresponding to the target camera to obtain multiple current frames; Compare each current frame with the acquired false warning picture set for similarity, and generate a similarity value corresponding to each current frame; Compare each similarity value with the false alarm threshold, take pictures with similarity values ​​greater than or equal to the false alarm threshold as similar frames, and take pictures with similarity values ​​less than the false alarm threshold as non-similar frames; Input the non-similar frames into the intelligent algorithm model corresponding to the algorithm configuration information, generate recognition results, and filter the similar frames; If the recognition result indicates an abnormality, the abnormal recognition result is sent to a cloud server and stored; If the number of the abnormal identification results is greater than or equal to one, searching for a terminal to be communicated with, the terminal to be communicated with being a terminal capable of communicating with an edge device; If the terminal to be communicated is detected within the set sensing range, the abnormal identification result is sent to the terminal to be communicated; If a reply instruction based on any of the abnormal identification results is obtained from the communication terminal, any of the stored abnormal identification results is deleted and a processing identifier is generated; The processing identifier is sent to the cloud server, and any of the abnormal identification results in the cloud server is marked according to the processing identifier.

2. The method according to claim 1, characterized in that Comparing the similarity between any current frame and the acquired false alarm picture set, and generating a similarity value corresponding to any current frame, including: The false alarm picture set includes at least one picture group, and the picture group includes marked core false alarm pictures; Compare the similarity between any current frame and each of the picture groups respectively to generate respective first similarity values, and use the first similarity value with the largest value as the similarity value corresponding to any current frame; The step of comparing the similarity between any current frame and any picture group to generate a first similarity value includes: If the picture group includes associated pictures related to the core false alarm picture, generating an inference picture group according to any current frame, wherein the inference picture group includes derivative pictures corresponding to each of the associated pictures; The inference picture group is compared with any one of the picture groups in terms of similarity to generate a first similarity value between any one of the current frames and any one of the picture groups.

3. The method according to claim 2, characterized in that Generating an inference picture group according to any current frame includes: Determine a first timestamp of the core false alarm picture and a second timestamp of any current frame; determining a time difference between the first timestamp and the second timestamp; Obtaining the timestamp of each of the associated pictures; Determine the acquisition time point corresponding to each associated picture according to the timestamp of each associated picture and the time difference; Extracting each picture in the monitoring video stream as each derivative picture according to each acquisition time point; The inference picture group is generated according to any one current frame and each of the derived pictures.

4. The method according to claim 2, characterized in that: The comparing the inference picture group with the any picture group in similarity to generate a first similarity value between the any current frame and the any picture group includes: Determine a first similarity between any current frame and the core false positive picture; Determine a second similarity between each of the associated pictures and the derived pictures corresponding to each of the associated pictures; A first similarity value between any current frame and any picture group is generated according to the first similarity and each of the second similarities.

5. The method according to claim 1, characterized in that The step of extracting frames from the monitoring video stream based on the acquired algorithm configuration information corresponding to the target camera to obtain multiple current frames includes: Decapsulate the video transmission protocol to generate the monitoring video stream in h264 or h265 format; Decoding the surveillance video stream to obtain picture frame information in RGB color space or YUV color space; Extract frame data regularly according to the acquired frame extraction interval; Performing a scaling operation on the frame data to obtain a set resolution image; The image after the scaling operation is encoded to obtain multiple current frames.

6. A video analysis device based on edge-cloud collaboration, characterized in that: The device comprises: An acquisition module is used to acquire the surveillance video stream collected by the target camera; A frame extraction module is used to extract frames from the monitoring video stream based on the acquired algorithm configuration information corresponding to the target camera to obtain multiple current frames; A comparison module is used to compare the similarity of each current frame with the acquired false warning picture set, and generate a similarity value corresponding to each current frame; A screening module, used to compare each similarity value with a false alarm threshold, and to take pictures with similarity values ​​greater than or equal to the false alarm threshold as similar frames, and to take pictures with similarity values ​​less than the false alarm threshold as non-similar frames; A recognition module, used to input the non-similar frames into the intelligent algorithm model corresponding to the algorithm configuration information, generate recognition results, and filter the similar frames; A connection module is used to send the abnormal identification result to the cloud server and store the abnormal identification result when the identification result represents an abnormality; if the number of the abnormal identification results is greater than or equal to one, search for a terminal to be communicated, where the terminal to be communicated is a terminal that can communicate with an edge device; if the terminal to be communicated is detected within a set sensing range, send the abnormal identification result to the terminal to be communicated; if a reply instruction based on any of the abnormal identification results is obtained from the terminal to be communicated, delete any of the stored abnormal identification results and generate a processing identifier; send the processing identifier to the cloud server, and mark any of the abnormal identification results in the cloud server according to the processing identifier.

7. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Transformer substation online intelligent patrol system

    CN113381511A

  • Character secondary verification method and device based on gaits

    CN113469095A