Video analysis method, management method of video analysis, and related devices

By combining the interaction data from other cameras during the execution phase of the video analysis task, the problems of resource consumption and low efficiency in traditional video surveillance systems are solved, achieving efficient target search and tracking while reducing computational load and resource requirements.

CN113378616BActive Publication Date: 2026-07-24HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
Filing Date
2021-03-01
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Traditional video surveillance systems suffer from resource consumption and inefficiency when manually reviewing videos due to the increased number of cameras. This is especially true for target search and tracking tasks, which require a large amount of computation and storage of feature values, leading to an explosive increase in resource demand.

Method used

By combining interactive data from other cameras during the execution phase of video analysis tasks, the search and calculation of a large number of feature values ​​can be avoided. Collaborative analysis using smart cameras or video analysis devices can narrow the scope of analysis and improve efficiency.

Benefits of technology

It reduces resource consumption, lowers computational load, improves resource utilization and analysis efficiency, avoids the need for high-performance computing clusters and large-capacity resources, and achieves efficient target search and tracking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113378616B_ABST
    Figure CN113378616B_ABST
Patent Text Reader

Abstract

The application provides a video analysis method applied to the field of video analysis, and includes the following steps: obtaining a first video captured by a first camera, obtaining interaction data of a second video, and then performing a first video analysis task on the first video based on the interaction data of the second video to obtain an analysis result of the first video. Since the interaction data obtained from other video tasks is combined for analysis in the execution stage of the video analysis task, when some artificial intelligence (AI) related applications trigger tasks such as target search and target tracking, the data stored in the execution stage of the video analysis task can be used for response, without the need to search and calculate a large number of feature values, thereby avoiding explosive resource consumption and improving response efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to Chinese Patent Application No. 202010159116.5, filed on March 9, 2020, entitled “Distributed Visual Analysis Method and System”, the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of video analytics, and more particularly to a video analytics method, a video analytics management method, and related devices, equipment, computer-readable storage media, and computer program products. Background Technology

[0003] Video surveillance is an important means of security and prevention. Traditional video surveillance systems include front-end cameras, transmission cables, and a video monitoring platform. Monitoring personnel can view the video footage of the monitored area captured by the front-end cameras in the monitoring room through the video monitoring platform, thus solving the problem of manpower costs associated with on-site monitoring.

[0004] As the number of cameras increases and the complexity of video surveillance tasks grows, the cost of manually reviewing videos is also rising. To address this, the industry has introduced artificial intelligence (AI) technology to assist in video analysis. Specifically, when applications need to perform tasks such as target search and tracking, they first extract features from the video streams captured by each camera in the background, obtaining feature vectors corresponding to each camera's video stream, and then storing these feature vectors. When the application triggers a task, it searches and calculates based on the stored feature vectors according to the task, thereby achieving target search and tracking.

[0005] However, when an application triggers a task, the search and computation of stored feature vectors based on the task results in a surge in resource consumption, leading to a surge in resource demands. Furthermore, this approach requires a large amount of computation, making it relatively inefficient. Summary of the Invention

[0006] This application provides a video analysis method that analyzes interactive data obtained from other video analysis tasks during the execution phase of the video analysis task. When an application triggers a task, it can directly respond based on the data stored during the execution phase of the video analysis task, without needing to search and calculate a large number of feature values. This avoids explosive resource consumption and demand, reduces computational load, and improves response efficiency. This application also provides a video analysis management method, as well as corresponding apparatus, devices, computer-readable storage media, and computer program products.

[0007] Firstly, this application provides a video analysis method. When the camera includes a smart camera with analysis capabilities, the video analysis method can be executed by the smart camera. When the camera is a non-smart camera without analysis capabilities (also known as a regular camera), the video analysis method can also be executed by a background video surveillance platform (e.g., a video analysis device within the video surveillance platform).

[0008] The camera or video analysis device can acquire video captured by the camera (e.g., a first video) and interactive data of video captured by other cameras (e.g., a second video). The interactive data is obtained by performing a video analysis task on the video. Then, based on the interactive data of the second video, the camera or video analysis device performs a first video analysis task on the first video to obtain the analysis result of the first video.

[0009] This method analyzes data by combining interactive data from other video analysis tasks during the execution phase of the video analysis task. Thus, when the application triggers tasks such as target search and tracking, it can directly analyze the data stored during the execution phase of the video analysis task (e.g., data related to the interactive data of the second video obtained from the first video analysis task), without having to search and calculate a large number of feature values. This avoids explosive resource consumption and demand, eliminating the need for high-performance computing clusters and large-capacity resources, improving resource utilization and reducing costs. Furthermore, analyzing the first video based on the interactive data of the second video enables collaborative analysis, narrowing the analysis scope, reducing computational load, and improving analysis efficiency.

[0010] In some possible implementations, the camera or video analysis device may also acquire action logic. This action logic includes one or more of the following parameters: a task type parameter for executing the first video analysis task, a time parameter for executing the first video analysis task, and a condition parameter for executing the first video analysis task.

[0011] The task type parameter describes the type of video analysis task. Task types can include face detection (detecting a specified face), body detection (detecting a specified body), vehicle license plate detection (detecting a specified license plate), vehicle feature detection (detecting a specified vehicle using feature values), crowd counting, vehicle counting, or specific behavior detection. Specific behaviors can be set according to requirements; for example, specific behaviors can be one or more of the following: not wearing a mask, fighting, occupying public space illegally, or using a mobile phone while driving.

[0012] The time parameter describes the execution time of the video analysis task. This execution time can include the start time, and further, it can include the duration or end time. The start time can be determined based on the distance between the cameras. Furthermore, the start time can also be determined based on the target's speed.

[0013] The condition parameters may include a similarity threshold. This similarity threshold can be used to determine whether the target in the first video and the target to be tracked in the second video are the same target. The similarity threshold can be set based on empirical values, for example, it can be set to 0.93.

[0014] Action logic can also be used to instruct adjustment logic for the camera. This adjustment logic includes adjusting the camera's orientation and / or focal length. For example, in a specific behavior detection and alarm scenario, action logic may include adjusting the camera's orientation and focal length to focus on the target for behavior detection.

[0015] Based on the interaction data between the action logic and the second video, the camera or video analysis device performs a first video analysis task on the first video. On the one hand, it can realize collaborative analysis of videos based on related regions. On the other hand, by executing video analysis tasks through action logic, it can achieve targeted analysis, such as analysis of faces or analysis of the human body, thus improving analysis efficiency and response efficiency.

[0016] In some possible implementations, the camera or video analytics device can receive the motion logic sent by the management device. This motion logic instructs the analysis of the video. This motion logic can be written by an administrator using a programming interface provided by the management device. The administrator can be a professional capable of writing motion logic.

[0017] In some possible implementations, the analysis results of the first video include data associated with the interaction data of the second video. In this way, the camera or video analysis device can perform analysis based on the data associated with the interaction data of the second video, enabling rapid response to tasks such as target detection and tracking without having to search and calculate a large number of feature values. This avoids explosive resource consumption and demand, and improves response efficiency.

[0018] In some possible implementations, the interactive data of the second video includes information about the target to be tracked appearing in the second video. The information about the target to be tracked may include any one or more of the target's attributes, feature values, or target images. Considering the transmission overhead of the interactive data, the target image can also be stored, and then its storage address can be used instead of the target image in the interactive data for transmission. This can significantly reduce the amount of data transmitted and lower transmission overhead.

[0019] Accordingly, the camera or video analysis device can identify the target in the first video, obtain the target's information, and then, based on the target's information and the information of the target to be tracked in the second video, obtain the analysis result of the first video. The analysis result includes information about targets whose information similarity to the target to be tracked in the second video meets preset conditions.

[0020] This allows for target tracking during the video analysis phase, specifically by using video from related regions. This avoids the need for extensive feature comparisons and searches during target detection and tracking tasks, preventing explosive resource consumption and demand. Furthermore, it enables direct responses to target detection and tracking tasks based on data from the video analysis phase, thus improving response efficiency.

[0021] In some possible implementations, the camera or video analysis device can identify targets in the first video and obtain information about the targets based on a pre-trained artificial intelligence model, such as a human detection model, a face detection model, a crowd counting model, or a vehicle counting model.

[0022] By using artificial intelligence models to identify targets in the first video, automated video surveillance can be achieved. This can reduce labor costs and avoid errors caused by manual monitoring, thereby improving the comprehensiveness and accuracy of video surveillance.

[0023] In some possible implementations, a subscription relationship exists between the first camera and the second camera. This subscription relationship corresponds to the subscription relationship for interactive data of the videos captured by the cameras. For example, if the first camera subscribes to the second camera, then a first video analysis task that analyzes the first video captured by the first camera subscribes to the interactive data of the second video captured by the second camera.

[0024] The subscription relationship is pre-built by the management device. Specifically, the management device can establish subscription relationships between cameras through different modes. In this way, when performing video analysis, the management device can send interactive data of videos captured by other cameras subscribed to by the camera to the camera or video analysis device based on the subscription relationship, so that the camera or video analysis device can achieve collaborative analysis based on the interactive data and improve analysis efficiency.

[0025] In some possible implementations, the video analysis method described above can be performed by the first camera. The first camera can be a smart camera. Similarly, the interactive data of the second video can be obtained by the second camera performing a second video analysis task on the second video. The second camera can also be a smart camera.

[0026] In some possible implementations, the first camera acquires the interaction data of the second video from the second camera. For example, when the cameras are connected in a point-to-point manner, they can directly send and receive interaction data based on their subscription relationship without needing a management device as an intermediary. This improves the efficiency of sending and receiving interaction data and avoids the resource consumption caused by relaying through a management device.

[0027] In some possible implementations, the first camera can also obtain the interaction data of the second video from the management device. Specifically, the management device can be used to forward the interaction data based on a subscription relationship. The camera or video analysis device performs a second analysis task on the second video to obtain the interaction data of the second video, and can report this interaction data to the management device. In this way, the management device can send the interaction data of the second video to the first camera according to the subscription relationship. This simplifies the processing logic of the camera and lowers the threshold for camera operation.

[0028] In some possible implementations, the camera or video analysis device can generate interactive data for the first video based on the analysis results. For example, some or all information from the analysis results of the first video can be used as interactive data, and then the interactive data can be sent to a management device or a third camera. The third camera and the first camera have a subscription relationship. This enables continuous analysis, such as continuous tracking of a target.

[0029] In some possible implementations, the method is performed by a video analysis device that has established a communication connection with the first camera and the second camera.

[0030] In some possible implementations, the video analysis method of this application is applied to security monitoring. The first video analysis task includes one or more of the following tasks: person detection and tracking, vehicle detection and tracking, crowd counting and statistics, vehicle counting and statistics, and specific behavior detection and alarm.

[0031] Secondly, this application provides a video analysis management method. This method can be executed by a management device. The management device is communicatively connected to multiple cameras. The management device can be a software device, which can be deployed on a general-purpose device, such as a server. The management device can also be a hardware device with management functions. For ease of description, this application uses the management device as an example of a software device.

[0032] Specifically, the management device receives interactive data from a second video, which is captured by the second camera. The management device then sends the interactive data to the first camera or a video analysis device, enabling the first camera or the video analysis device to perform a first video analysis task on the first video based on the interactive data. The first camera and the second camera have a subscription relationship, and the first video is captured by the first camera.

[0033] This method analyzes interactive data from other video analysis tasks during the execution phase of the video analysis task. Thus, when the application triggers tasks such as target search and tracking, it can directly analyze the data stored during the execution phase of the video analysis task, without having to search and calculate a large number of feature values. This avoids explosive resource consumption and demand, narrows the analysis scope, reduces computational load, and improves analysis efficiency.

[0034] In some possible implementations, before the management device receives the interactive data of the second video, it may also establish a subscription relationship between the first camera and the second camera among the at least one camera. The management device can perform the subscription relationship establishment step once, and when responding to subsequent tasks, it can directly respond based on the established subscription relationship, thereby avoiding unnecessary resource waste and improving response efficiency.

[0035] In some possible implementations, the management device can establish a subscription relationship between the first camera and the second camera in the at least one camera configuration mode. Specifically, the management device receives a subscription instruction sent by the first camera, wherein the subscription instruction includes subscription parameters, which may be, for example, a distance parameter. The management device queries the second camera that satisfies the subscription parameters based on the subscription parameters and establishes a subscription relationship between the first camera and the second camera.

[0036] In some possible implementations, the management device can establish a subscription relationship between the first camera and the second camera in the at least one camera configuration through a passive execution mode. Specifically, the management device receives a subscription instruction sent by a manager, which may include the identifier of the camera to be subscribed to, and then the management device establishes the subscription relationship between the first camera and the second camera according to the subscription instruction.

[0037] In this way, the management device can establish subscription relationships at different stages through the pre-configuration mode or passive execution mode, such as when starting a video analysis task or during the execution of a video analysis task, and realize collaborative analysis based on the subscription relationship to meet personalized analysis needs.

[0038] In some possible implementations, the management device may also send action logic to the first camera or the video analysis device, so that the first camera or the video analysis device performs the first video analysis task on the first video based on the interaction data of the second video and the action logic. This enables targeted video analysis based on action logic, improving analysis efficiency.

[0039] In some possible implementations, before sending the action logic to the first camera or the video analytics device, the management device may also obtain action logic written by an administrator through its programmable interface. In some embodiments, the management device may also obtain action logic configured by a user through a video analytics application.

[0040] Thirdly, this application provides a video analysis device. The device includes:

[0041] A communication unit is used to acquire a first video captured by a first camera and interactive data for acquiring a second video, wherein the second video is acquired by a second camera;

[0042] An analysis unit is configured to perform a first video analysis task on the first video based on the interaction data of the second video, and obtain the analysis result of the first video, wherein the interaction data of the second video is obtained by performing a second video analysis task on the second video.

[0043] In some possible implementations, the communication unit is further configured to:

[0044] Obtain action logic, wherein the action logic includes one or more of the following parameters: task type parameter for executing the first video analysis task, time parameter for executing the first video analysis task, and condition parameter for executing the first video analysis task;

[0045] The analysis unit is specifically used for:

[0046] Based on the action logic and the interaction data of the second video, a first video analysis task is performed on the first video.

[0047] In some possible implementations, the communication unit is specifically used for:

[0048] The system receives the action logic sent by the management device, wherein the action logic is written by the administrator according to the programming interface provided by the management device.

[0049] In some possible implementations, the analysis results of the first video include data associated with the interaction data of the second video.

[0050] In some possible implementations, the interactive data of the second video includes information about the target to be tracked that appears in the second video;

[0051] The analysis unit is specifically used for:

[0052] Identify the target in the first video and obtain information about the target;

[0053] Based on the information of the target and the information of the target to be tracked in the second video, the analysis result of the first video is obtained. The analysis result includes information of the target whose information similarity with the target to be tracked in the second video meets a preset condition.

[0054] In some possible implementations, the analysis unit is specifically used for:

[0055] Based on a pre-trained artificial intelligence model, targets in the first video are identified, and information about the targets is obtained.

[0056] In some possible implementations, a subscription relationship exists between the first camera and the second camera, and this subscription relationship is pre-built by a management device.

[0057] In some possible implementations, the analysis unit is further used for:

[0058] The interactive data of the first video is generated based on the analysis results of the first video;

[0059] The communication unit is also used for:

[0060] The interactive data of the first video is sent to the management device or the third camera, wherein the third camera and the first camera have a subscription relationship.

[0061] Fourthly, this application provides a management device. The management device is communicatively connected to multiple cameras, and the management device includes:

[0062] A communication unit is configured to receive interactive data of a second video, wherein the second video is captured by the second camera; and to send the interactive data of the second video to the first camera or a video analysis device, so that the first camera or the video analysis device performs a first video analysis task on the first video based on the interactive data of the second video, wherein there is a subscription relationship between the first camera and the second camera, and the first video is captured by the first camera.

[0063] In some possible implementations, the device further includes:

[0064] A subscription unit is used to establish a subscription relationship between the first camera and the second camera among the at least one cameras.

[0065] In some possible implementations, the subscription unit is specifically used for:

[0066] Receive a subscription instruction sent by the first camera, wherein the subscription instruction includes subscription parameters;

[0067] The system queries the second camera that satisfies the subscription parameters and establishes a subscription relationship between the first camera and the second camera.

[0068] In some possible implementations, the subscription unit is specifically used for:

[0069] Receive subscription instructions sent by administrators;

[0070] The subscription relationship between the first camera and the second camera is established according to the subscription instruction.

[0071] In some possible implementations, the communication unit is further used for:

[0072] The action logic is sent to the first camera or the video analysis device, so that the first camera or the video analysis device performs the first video analysis task on the first video based on the interaction data of the second video and the action logic.

[0073] In some possible implementations, the communication unit is further used for:

[0074] The system can acquire action logic written by administrators through the programmable interface of the management device, or acquire action logic configured by users through a video analytics application.

[0075] Fifthly, this application provides a device including a processor and a memory. The processor and the memory communicate with each other. The memory stores executable program code, and the processor reads the executable program code stored in the memory to implement the functions of the video analysis device according to any implementation of the third aspect of this application, or to implement the functions of the management device according to any implementation of the fourth aspect of this application.

[0076] Sixthly, this application provides a camera. The camera includes a processor, a memory, and an image sensor. The image sensor is used to acquire a first video, the memory stores executable program code, and the processor reads the executable program code to implement the functions of the video analysis device described in the third aspect of this application.

[0077] In a seventh aspect, this application provides a computer-readable storage medium storing instructions that run in a device that implements the functions of a video analysis apparatus according to any implementation of the third aspect of this application, or the functions of a management apparatus according to any implementation of the fourth aspect of this application.

[0078] Eighthly, this application provides a computer program product containing instructions, which, when executed by a device, enables the device to perform the functions of a video analysis device as described in any implementation of the third aspect of this application, or the functions of a management device as described in any implementation of the fourth aspect of this application.

[0079] Based on the implementation methods provided in the above aspects, this application can be further combined to provide more implementation methods. Attached Figure Description

[0080] To more clearly illustrate the technical methods of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly described below.

[0081] Figure 1 An application scenario diagram of a video analysis method provided in this application embodiment;

[0082] Figure 2A A schematic diagram of an interface for triggering target detection or trajectory search in an application provided in this embodiment of the application;

[0083] Figure 2B A schematic diagram of a trajectory search process provided in an embodiment of this application;

[0084] Figure 3 A trajectory diagram of a target provided in an embodiment of this application;

[0085] Figure 4 An application scenario diagram of a video analysis method provided in this application embodiment;

[0086] Figure 5 An application scenario diagram of a video analysis method provided in this application embodiment;

[0087] Figure 6 A flowchart illustrating a video analysis method provided in this application embodiment;

[0088] Figure 7 A positional relationship diagram of a camera provided in an embodiment of this application;

[0089] Figure 8 A flowchart illustrating a video analytics management method provided in this application embodiment;

[0090] Figure 9 An interactive flowchart of a video analysis method provided in an embodiment of this application;

[0091] Figure 10 An interactive flowchart of a video analysis method provided in an embodiment of this application;

[0092] Figure 11 This is a schematic diagram of the structure of a video analysis device provided in an embodiment of this application;

[0093] Figure 12 This is a schematic diagram of the structure of a management device provided in an embodiment of this application;

[0094] Figure 13 This is a schematic diagram of the structure of a device provided in an embodiment of this application;

[0095] Figure 14 This is a schematic diagram of the structure of a device provided in an embodiment of this application;

[0096] Figure 15This is a schematic diagram of the structure of a camera provided in an embodiment of this application. Detailed Implementation

[0097] The terms "first" and "second" used in the embodiments of this application are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined with "first" and "second" may explicitly or implicitly include one or more of that feature.

[0098] First, some technical terms involved in the embodiments of this application will be introduced.

[0099] Video surveillance refers to the method of monitoring a geographical area using video surveillance equipment (also known as cameras, which can be fixed in a location, or cameras in drones, patrol vehicles, etc.) that record the on-site conditions of that area. Video surveillance technically solves the problem of high labor costs associated with manual monitoring (personnel conducting on-site inspections), and therefore has been widely used in many situations. For example, video surveillance can be applied to community management, enterprise management, and road traffic management to ensure the personal and property safety of community residents, the security of enterprise assets, and smooth traffic flow.

[0100] While traditional video surveillance has solved the problem of manual monitoring, it still requires manual review of the videos for effective monitoring. As the number of video surveillance cameras deployed increases, so does the workload of manually reviewing the videos, leading to increased labor costs. Therefore, the industry has introduced artificial intelligence (AI) technology to assist in video analysis, reducing the workload of manual video review and lowering labor costs.

[0101] Artificial intelligence (AI) simulates human thought processes and / or intelligent behaviors (such as learning, reasoning, thinking, and planning) through computer programs running on computers. As a branch of AI, deep learning (DL) has made significant progress in tasks such as image recognition. Therefore, the industry primarily uses deep learning to assist in video analysis for video surveillance.

[0102] Deep learning extracts features from data to discover its deeper characteristic representations, enabling it to perform classification, regression, and prediction. Deep learning is primarily applied in perception and decision-making scenarios within the field of artificial intelligence, such as image recognition, speech recognition, natural language translation, and computer gaming.

[0103] In AI-based video surveillance scenarios, the video surveillance platform first decodes the videos captured by various cameras (e.g., cameras installed in different geographical areas) to obtain images corresponding to each video. Then, it uses a deep learning model to extract feature vectors from the images and stores these feature vectors. These feature vectors can be represented by binary data, also known as feature values. When an application triggers a task, such as target search and tracking, the video surveillance platform can search and calculate the stored feature values ​​according to the task. For example, it can search the stored feature values ​​based on search criteria, and then calculate the similarity between the searched feature values ​​and preset feature values ​​to obtain the target's trajectory, thus achieving target search and tracking.

[0104] However, when an application triggers a task, the video surveillance platform experiences a surge in resource consumption due to the search and computation of stored feature values. This leads to a sudden spike in resource demands. For example, the video surveillance platform may be configured with a high-performance computing cluster to support feature value searching and a large memory capacity to support feature value comparison. These resources (such as computing and memory resources) may be idle during periods other than feature value searching and comparison, resulting in low resource utilization. Furthermore, when an application triggers a task, the video surveillance platform needs to perform a significant amount of computation to obtain the results required by the application (e.g., the trajectory of the target), leading to relatively low application response efficiency.

[0105] In view of this, embodiments of this application provide a video analysis method. When the camera includes a smart camera with analysis capabilities, the video analysis method can be executed by the smart camera. When the camera is a non-smart camera without analysis capabilities (also known as a regular camera), the video analysis method can also be executed by a background video surveillance platform (e.g., a video analysis device within the video surveillance platform).

[0106] Specifically, the camera or video analysis device can acquire the video captured by the camera (referred to as the first video in this application embodiment for ease of description), and acquire the interactive data of the video captured by other cameras (referred to as the second video in this application embodiment for ease of description). The interactive data is obtained by performing a video analysis task on the video. Then, based on the interactive data of the second video, the camera or video analysis device performs a first video analysis task on the first video to obtain the analysis result of the first video.

[0107] This method analyzes data by combining interactive data from other video analysis tasks during the execution phase of the video analysis task. Thus, when the application triggers tasks such as target search and tracking, it can directly analyze the data stored during the execution phase of the video analysis task (e.g., data related to the interactive data of the second video obtained from the first video analysis task), without having to search and calculate a large number of feature values. This avoids explosive resource consumption and demand, eliminating the need for high-performance computing clusters and large-capacity resources, improving resource utilization and reducing costs. Furthermore, analyzing the first video based on the interactive data of the second video enables collaborative analysis, narrowing the analysis scope, reducing computational load, and improving analysis efficiency.

[0108] For example, in a target search and tracking task, when analyzing the first video captured by a camera, interactive data from a second video captured by other cameras is obtained. This interactive data can include the type, color, and feature values ​​of the target to be tracked. When applying this interactive data to the analysis of the first video, all targets captured in the first video can be detected and their features extracted. These targets are then compared with the type, color, and feature values ​​of the tracked target in the interactive data to determine the information of targets in the first video whose similarity to the tracked target meets preset conditions, thus obtaining the analysis results. The analysis results of the first video can then be reported to a management device, or interactive data from the first video can be generated based on the analysis results for use in the analysis of other videos.

[0109] It should be noted that this method is not a video analysis method developed for a specific business, but can be widely applied to various video surveillance scenarios. For example, it can be used in criminal investigation scenarios to track criminal suspects, in public security management scenarios to monitor events that violate public security regulations (such as fights), in traffic management scenarios to monitor vehicles involved in accidents, or in shopping mall management scenarios to identify very important customers (VIPs) and to disperse dense crowds, thereby optimizing pedestrian flow.

[0110] To make the technical solution of this application clearer and easier to understand, the application scenarios of this application are described below with reference to the accompanying drawings.

[0111] See Figure 1 The diagram shows an application scenario of the video analysis method. The scenario includes multiple cameras 100, a video surveillance platform 200, and an application 300. The multiple cameras 100 are connected to the video surveillance platform 200, and the video surveillance platform 200 is connected to the application 300.

[0112] Application 300 can be a video analytics application, which can provide one or more functions such as trajectory search, target detection, target tracking, target counting, and specific behavior detection and alarm. The target can be an entity such as a person, vehicle, or animal, and the specific behavior refers to pre-defined detectable behaviors, such as fighting or wearing a mask. Based on the video surveillance platform 200, the video analytics application can provide functions such as person detection and tracking, vehicle detection and tracking, crowd counting, and vehicle and specific behavior detection and alarm. In this embodiment, application 300 can be a dedicated client for implementing trajectory search or target detection functions, or it can be a browser for providing trajectory search and target detection services.

[0113] Application 300 can provide an interactive interface. This interface can be a graphical user interface (GUI) or a command user interface (CUI). For ease of understanding, the following example uses a GUI as the interactive interface.

[0114] like Figure 2A As shown, application 300 presents videos captured by multiple cameras 100 to the user through a GUI, such as videos 1 to N captured by cameras 1 to N. The GUI also includes target detection controls and trajectory search controls. Users (specifically, personnel using application 300, such as security personnel or investigators) can trigger target detection and trajectory search operations through these controls. For example, a user can trigger the trajectory search control, input an image of the target to be searched, and thus trigger a trajectory search operation for that target. In response to the trajectory search operation, application 300 can send a trajectory search request to the video surveillance platform 200.

[0115] The video surveillance platform 200 includes a management device 202 and a storage and retrieval device 204. When the multiple cameras 100 connected to the video surveillance platform 200 include ordinary cameras (i.e., cameras that cannot perform video analysis and processing), the video surveillance platform 200 may also include a video analysis device 206. The video analysis device 206 is used to perform video analysis tasks on the video captured by the video cameras 100. When the multiple cameras 100 connected to the video surveillance platform 200 are all intelligent cameras (i.e., cameras capable of video analysis and processing), the functions of the video analysis device 206 can also be performed by the intelligent cameras.

[0116] The management device 202 is used to establish subscription relationships between the various cameras in the multiple cameras 100. For example, if the first camera and the second camera are cameras in different geographical locations, the management device 202 can receive a subscription instruction sent by the first camera. This subscription instruction includes subscription parameters. The management device 202 can then query the second camera that meets the subscription parameters and establish a subscription relationship between the first camera and the second camera.

[0117] A first camera captures a first video, and a second camera captures a second video. The second camera or video analysis device 206 can perform a second video analysis task on the second video to obtain interactive data of the second video. This interactive data is used for interaction between video analysis tasks or between cameras. Specifically, the management device 202 is also used to receive the interactive data of the second video and, based on the subscription relationship between the first and second cameras, send the interactive data of the second video to the first camera or video analysis device 206. Thus, the first camera or video analysis device 206 can perform a first video analysis task on the first video based on the interactive data of the second video to obtain the analysis result of the first video. The analysis result of the first video includes data associated with the interactive data of the second video.

[0118] In some possible implementations, the management device 202 further includes a programming interface unit for acquiring action logic written by the administrator through the programmable interface. The management device 202 is used to send the action logic to each camera or video analysis device 206. The action logic is used to instruct the analysis logic for the video. The action logic may include any one or more of the following: task type parameters for performing the video analysis task, time parameters for performing the video analysis task, and condition parameters for performing the video analysis task. For example, the first camera or video analysis device 206 can perform a first video analysis task on the first video based on the action logic for the first video analysis task and the interaction data of the second video, thereby obtaining data associated with the interaction data of the second video.

[0119] The storage and search device 204 is used to store the analysis results of each video, such as the analysis results of the second video and the analysis results of the first video. The analysis results of different videos are related. For example, if the analysis result of the first video is related to the interaction data of the second video, then the analysis results of the first video and the analysis results of the second video stored in the search device 204 are related. This relationship can be reflected by sharing a common association identifier or by a relationship table; this application does not limit this. Furthermore, the analysis results of the first video can also be used to generate interaction data for the first video. This interaction data can be sent to the management device 202 or other cameras (such as a third camera) that have a subscription relationship with the first camera for analysis of the third video captured by the third camera. Accordingly, the analysis results of the third video captured by the third camera are also related to the interaction data of the first video; that is, the analysis results of the third video and the analysis results of the first video are also related.

[0120] The storage and search device 204 can also provide a search interface through which search criteria can be received, and data that meets the search criteria can be retrieved. For example, the storage and search device 204 supports searching by time, geographical location (e.g., camera location), camera identifier, attributes, feature values, images, etc. It should be noted that when the storage and search device 204 performs a search based on an image, the image can be converted into feature values ​​first, and then the search can be performed.

[0121] The storage and search device 204 is also configured to generate a trajectory of the target based on the interaction data of the second video and data associated with the interaction data of the second video in response to a trajectory search request. For example, see Figure 2BWhen a user needs to obtain the trajectory of a tracked target, the first, second, and third cameras, having subscribed to each other, will send interactive data of the second video to the first camera when the second camera detects the tracked target. The first camera, based on this interactive data, will analyze the first video and also detect the tracked target, then send the same interactive data to the third camera. The third camera, based on this interactive data, will analyze the third video and detect the tracked target. The analysis results of the second, first, and third videos can be stored in the storage and search device 204 using a relational table, and these analysis results are interconnected. Based on these interconnected analysis results, when searching for the trajectory of a tracked target, it is easier to obtain the trajectory without performing extensive feature value searches and comparisons, avoiding explosive resource consumption and demands, improving resource utilization, and increasing analysis efficiency by eliminating the need for extensive computation.

[0122] Furthermore, the storage and search device 204 can also return the target's trajectory to the application 300, so that the application 300 can present the target's trajectory to the user through a user interface such as a GUI. See details... Figure 3 The interface diagram showing the target's trajectory illustrates that application 300 can acquire the geographical locations of multiple cameras 100, generate a distribution diagram of the cameras 100 based on their geographical locations, and then display the target's trajectory on the distribution diagram of the cameras 100. For example... Figure 3 As shown, cameras 100 are represented by circles. When a camera detects a target, its image is displayed at the corresponding position in the distribution diagram, representing the target's presence at that location. Application 300 can connect the corresponding targets in the distribution diagram sequentially with directed arrows according to the time sequence of target detection, presenting the target's trajectory based on the path formed by multiple directed arrows.

[0123] In some other possible implementations, the functions of the management device 202 in the video surveillance platform 200 can also be implemented by the application 300. See also Figure 4 The diagram illustrates an application scenario for the video analytics method. This scenario includes multiple cameras 100, a video surveillance platform 200, and an application 300. The video surveillance platform 200 includes a storage and search management device 204, and further includes a video analytics device 206. The functions of the storage and search management device 204 and the video analytics device 206 are described below. Figure 1 The illustrated embodiment is described below. Application 300 is also used to implement... Figure 1 The function of the management device 202 in the illustrated embodiment.

[0124] Specifically, application 300 includes management device 302. Management device 302 can be a functional module, plugin, mini-program, etc., of application 300. Management device 302 is used to receive subscription instructions sent by administrators and establish a subscription relationship between the first camera and the second camera based on the subscription instructions. In this way, management device 302 can send interactive data of the second video to the first camera or video analysis device 206 according to the subscription relationship. When the management device is implemented by application 300, it can be a user of application 300 acting as an administrator, or it can be a professional administrator configuring application 300 before the user uses its functions. When the management device is independent of application 300, the administrator is usually not a user of application 300, but a professional who establishes subscription relationships and configures action logic.

[0125] In some possible implementations, the management device 302 is further configured to acquire the action logic configured by the administrator or user through application 300 and send it to the first camera or video analysis device 206. Accordingly, the first camera or video analysis device 206 can perform a first video analysis task on the first video based on the action logic and the interaction data of the second video, thereby obtaining the analysis results of the first video. Figure 1 In the scenario shown, camera 100 can be a bullet camera, a PTZ camera, or a camera on a turnstile or access control system. Furthermore, camera 100 can be a camera installed in a fixed location, or a camera attached to a mobile device, such as a camera on a drone or a camera on a vehicle.

[0126] Application 300 can be deployed on terminals. Terminals include, but are not limited to, desktop computers, laptops, tablets, smartphones, or smart wearable devices. Among them, smart wearable devices can include smartwatches, smart bracelets, smart glasses, etc.

[0127] The video surveillance platform 200 can be deployed in a cloud environment or a local data center. A cloud environment refers to a cloud computing cluster owned by a cloud service provider, used to provide computing, storage, and communication resources. The cloud computing cluster can include a central cloud and an edge cloud. A central cloud includes at least one central computing device (e.g., a central server), and an edge cloud includes at least one edge computing device (e.g., an edge server). A local data center refers to the data center to which the user belongs. A local data center includes at least one local computing device, such as at least one local server.

[0128] The various components of the video surveillance platform 200 can be deployed centrally in a cloud environment or a local data center, or they can be deployed in a distributed manner in either a cloud environment or a local data center. This embodiment of the application illustrates the example of the video surveillance platform 200 being deployed in a distributed manner in a cloud environment.

[0129] See Figure 5 The system architecture diagram of the video analysis method shown illustrates that the video analysis device 206 of the video surveillance platform 200 is deployed in the edge cloud, while the management device 202 and storage and search device 204 of the video surveillance platform 200 are deployed in the central cloud. The video surveillance platform 200 in the cloud environment can perform collaborative analysis of videos captured by the camera 100. Application 300 is deployed on the terminal. The terminal can run application 300 and interact with the video surveillance platform 200 in the cloud environment to achieve video surveillance.

[0130] When the video surveillance platform 200 is deployed in a cloud environment, video analytics methods can be provided to users as cloud services. Specifically, in the cloud environment, an instance of the video surveillance platform 200 runs, and application 300 can respond to user-triggered application operations, such as trajectory search operations, and interact with the instance of the video surveillance platform 200 to realize the application's functions.

[0131] When the video surveillance platform 200 is deployed in a local data center, specifically, the user can install the software package of the video surveillance platform 200 on a local server and then run the installed video surveillance platform 200 to implement the video analysis method of this application embodiment. In some embodiments, the above-mentioned software package can also be an installation-free software package, and the user can directly run the above-mentioned installation-free software package on the server to implement the video analysis method of this application embodiment.

[0132] Next, taking a video surveillance scenario where the cameras associated with the video surveillance platform 200 are ordinary cameras, and therefore the video surveillance platform 206 performs video analysis tasks on the videos captured by each camera as an example, the video analysis method provided in this application embodiment will be described in detail.

[0133] See Figure 6 The flowchart shown illustrates a video analysis method, which includes:

[0134] S602: The video analysis device 206 acquires the first video captured by the first camera.

[0135] The first camera specifically refers to any one or more cameras among a plurality of cameras deployed in the monitored area. The monitored area is the geographical area covered by the camera. In some embodiments, the monitored area may include any one or more geographical areas such as shopping malls, residential communities, companies, and roads. The video captured by the first camera is called the first video.

[0136] The video analysis device 206 can acquire the first video captured by the first camera in real time, receive the first video actively reported by the first camera, or periodically acquire the first video from the first camera. Furthermore, considering transmission efficiency, the first camera can compress the first video to obtain a first compressed video. The video analysis device 206 can acquire the first compressed video and then decompress it to obtain the first video.

[0137] S604: The video analysis device 206 acquires the interactive data of the second video.

[0138] The second video was captured by a second camera. The second camera is one or more cameras among multiple cameras deployed in the monitored area. The second camera and the first camera are located in different geographical locations. Interactive data is information used to describe the video content, specifically information about targets appearing in the video. Here, a target refers to an objectively existing object in the physical world; for example, a target can be a person, animal, or other movable living organism, or a vehicle or other movable non-living object.

[0139] Interactive data for the second video can be obtained by performing a second video analysis task on the second video. The type of interactive data can vary depending on the video analysis task. Examples of interactive data are provided below for different video analysis tasks.

[0140] For example, when the video analysis task is target detection or target tracking, the interactive data can be any one or more of the target's attributes, feature values, or target images. Considering the transmission overhead of the interactive data, the video analysis device 206 can also store the target image and then use the storage address of the target image instead of the target image in the interactive data for transmission. This can significantly reduce the amount of data transmitted and lower transmission overhead. The storage address can be represented by binary, encoded values ​​(such as base64 encoded values), or a Uniform Resource Locator (URL).

[0141] The attributes of a target are abstract descriptions of that target. For example, when the target is a person, the attributes could include any one or more of the following: the target type is a person, the person's gender, the person's height, and the person's clothing (e.g., the type and color of clothes, shoes, hats, and / or glasses). As another example, when the target is a vehicle, the attributes could include the target type being a vehicle, license plate number, vehicle type, and vehicle color.

[0142] A feature value is a value that expresses the characteristics of a target. A feature value can be a feature vector obtained by feature extraction from an image, a tensor, or a binary value based on the aforementioned feature vector. To improve accuracy and reduce computation, the video analysis device 206 can extract features from specific parts of the target in the image to obtain feature values. For example, if the target is a person, the video analysis device 206 can extract features from a face to obtain feature values.

[0143] A target image refers to an image frame in a video stream that includes a target. In some possible implementations, the target image can also be an image of a specific part of the target within the target's image frame. For example, if the target is a person, the target image can include a face image, a body image, and so on.

[0144] For example, when the video analytics task is target counting (such as crowd counting or vehicle counting), the interactive data can also include any one or more of the following: target quantity, density, etc. Density measures the degree of target concentration and is typically represented by the number of targets per unit area. When the video analytics task is specific behavior detection and alarm, the interactive data can also include target behavior, etc.

[0145] In geographical areas such as train stations, subway stations, and squares, the video analysis device 206 can also detect the number of people in these areas using an AI model. When the number of people exceeds a set first threshold, a risk warning can be issued, such as a reminder of a high risk of stampede. Similarly, in geographical areas such as roads and parking lots, the video analysis device 206 can also detect the number of vehicles in these areas using an AI model. When the number of vehicles exceeds a set second threshold, the user can be advised to take a detour or park in another geographical area. The first and second thresholds can be set based on empirical values, and this embodiment does not limit their specific settings.

[0146] For ease of understanding, this application also provides an example of interactive data, as shown below:

[0147]

[0148]

[0149] In the example above, the interaction data includes the target's attributes (e.g., type: human, gender, clothing color), facial feature values, and encoded values ​​of the facial image. The interaction data also includes a data identifier, which uniquely identifies the interaction data.

[0150] In a target tracking scenario, the interactive data of the second video includes information about the target to be tracked appearing in the second video. In some embodiments, the video analysis device 206 can determine that the target appearing in the second video is a target to be tracked when it matches a preset monitoring target. Specifically, the video analysis device 206 can extract features from the target appearing in the second video and from the image of the monitoring target, then calculate the distance between the feature values, and determine whether the target appearing in the second video matches the monitoring target based on this distance. In other embodiments, the video analysis device 206 can determine the target to be tracked in response to a user's specified operation on a target appearing in the second video. For example, application 300 can present a second video to a user, who can specify a target appearing in the second video. Application 300 can report the user-triggered specified operation on the target to the video analysis device 206. The video analysis device 206 can then determine the user-specified target as a target to be tracked in response to the user-triggered specified operation.

[0151] In some possible implementations, a subscription relationship exists between the first camera and the second camera. Correspondingly, a subscription relationship exists between the interactive data of the first video captured by the first camera and the interactive data of the second video captured by the second camera. The video analysis device 206 can receive the interactive data of the second video sent by the management device 202 based on the aforementioned subscription relationship.

[0152] The subscription relationship can be pre-built by the management device 202. The management device 202 can establish the subscription relationship through various modes. The different modes are described in detail below.

[0153] One mode is a pre-configuration mode. Specifically, the management device 202 receives a subscription instruction sent by the first camera, wherein the subscription instruction includes subscription parameters. Then, the management device 202 queries the second camera that satisfies the subscription parameters and establishes a subscription relationship between the first camera and the second camera.

[0154] The subscription parameters may include distance parameters. When the video analysis device 206 starts the first video analysis task, the first camera can send a subscription instruction to the management device 202. This subscription instruction includes the following subscription parameter: the surrounding distance is less than or equal to n kilometers (km). This subscription parameter is a distance parameter, where n is greater than 0. n can be set according to empirical values, for example, it can be 2. Correspondingly, this subscription parameter can be represented as listener.add(nearby<2km,face_data). In this way, the management device 202 can determine the second camera that meets the subscription parameters based on the positional relationship between the cameras 100 and the above distance parameter, thereby establishing a subscription relationship between the first camera and the second camera.

[0155] Furthermore, the management device 202 can more accurately determine the second cameras that have a subscription relationship with the first camera based on distance, combined with historical average vehicle speed and / or walking time. For example, in a hit-and-run vehicle tracking scenario, a subscription relationship can be established between the first camera and cameras that may detect the vehicle within the next hour, based on distance and historical average vehicle speed. Similarly, in a person tracking scenario, a subscription relationship can be established between the first camera and cameras that may detect the pedestrian within the next hour, based on distance and walking time.

[0156] Among them, the distance between the 100 cameras, the historical average vehicle speed, and the walking time can be stored in a graph format, specifically in a graph database. Figure 7 An example of a positional relationship diagram of camera 100 is provided, such as... Figure 7 As shown, each camera 100 is a node in the graph, and the distance between cameras 100 and the historical average vehicle speed between two cameras 100 can be stored as edges of the graph.

[0157] Another mode is the passive execution mode. Specifically, the management device 202 receives a subscription instruction sent by the administrator, and establishes a subscription relationship between the first camera and the second camera based on the subscription instruction. In some scenarios, the administrator can generate a subscription instruction through a programmable interface, which may include the identifier of the camera 100 to which the subscription relationship is to be established, and then send the subscription instruction to the management device 202. In this way, the management device 202 can directly establish a subscription relationship between the first camera and the second camera based on the aforementioned subscription instruction. When performing video analysis, the video analysis device 206 can passively receive interactive data from other video analysis tasks, such as interactive data from the second video. Furthermore, the video analysis device 206 can also receive action logic and perform a first video analysis task on the first video according to the action logic. For example, the video analysis device 206 can perform face detection, human body detection, etc. on the first video according to the action logic.

[0158] In some possible implementations, the subscription relationship can also be established by the cameras 100 through a communication mode. Specifically, the cameras 100 can communicate with each other, and thus, they can establish subscription relationships based on information such as distance. For example, the first camera can broadcast a subscription request message to other cameras, requesting to establish a subscription relationship with cameras within 2km of the first camera. Cameras receiving this subscription request message can determine whether they are within 2km; if so, they return a subscription response message. In this way, the first camera establishes a subscription relationship with the second camera. The first camera can then report the established subscription relationship to the management device 202.

[0159] For security reasons, the management device 202 can also authenticate subscription permissions before establishing a subscription relationship. When subscription permission authentication is successful, the management device 202 allows the establishment of a subscription relationship; when subscription permission authentication fails, the management device 202 refuses to establish a subscription relationship. When the subscription relationship is established between cameras 100 through a communication mode, the camera 100 can perform subscription permission authentication. This authentication can be implemented using a signature verification mechanism based on a public-private key pair.

[0160] It should be noted that S602 and S604 can be executed in parallel or in a set order. For example, the video analysis device 206 can execute S602 first and then S604, or execute S604 first and then S602. This application embodiment is only used as an example of executing S602 first and then S604.

[0161] S606: The video analysis device 206 performs a first video analysis task on the first video based on the interactive data of the second video, and obtains the analysis result of the first video.

[0162] The interactive data of the second video includes relevant data about the targets appearing in the second video, such as target attributes, target feature values, target images, target quantity, or target density. The video analysis device 206 can perform a first video analysis task on the first video based on the interactive data using an AI model, obtaining the analysis results of the first video. The analysis results of the first video include data related to the interactive data of the second video.

[0163] Specifically, the video analysis device 206 can identify targets in the first video. For example, it can use a pre-trained AI model to identify targets in the first video and obtain target information. The pre-trained AI model can include a face detection model, a human body detection model, a crowd counting model, a vehicle counting model, etc. Then, the video analysis device 206 obtains the analysis results of the first video based on the target information and the information of the target to be tracked in the second video. The analysis results include information of targets whose information similarity to the target to be tracked in the second video meets preset conditions. Meeting the preset conditions for similarity can be that the similarity reaches a preset similarity threshold or that the similarity reaches the maximum, etc. The following is an example of a person detection and tracking scenario. Specifically, the video analysis device 206 can identify people in the first video and obtain information about at least one person in the first video. The person information can include at least one of gender, clothing color, facial feature values, and facial image. The video analysis device 206 compares the information of at least one person in the first video with the information of the person to be tracked in the second video. For example, it can calculate the similarity of facial feature values. When the similarity reaches a preset similarity threshold (e.g., 0.95), the person in the first video is determined to be the person to be tracked. The video analysis device 206 can obtain the analysis results of the first video based on the information of the person.

[0164] In some possible implementations, the video analysis device 206 can also acquire action logic, such as receiving action logic issued by the management device 202. This action logic can be written by the administrator according to the programming interface provided by the management device 202. Then, based on the action logic and the interaction data of the second video, the video analysis device 206 performs a first video analysis task on the first video to obtain the analysis result of the first video.

[0165] The action logic is used to instruct the analysis logic for the video. The action logic can include one or more of the following parameters: task type parameter for performing the video analysis task, time parameter for performing the video analysis task, and condition parameter for performing the video analysis task.

[0166] The task type parameter describes the type of video analysis task. Task types can include face detection (detecting a specified face), body detection (detecting a specified body), vehicle license plate detection (detecting a specified license plate), vehicle feature detection (detecting a specified vehicle using feature values), crowd counting, vehicle counting, or specific behavior detection. Specific behaviors can be set according to requirements; for example, specific behaviors can be one or more of the following: not wearing a mask, fighting, occupying public space illegally, or using a mobile phone while driving.

[0167] The time parameter describes the execution time of the video analysis task. This execution time can include the start time, and further, it can include the execution duration or the execution end time. The start time can be determined based on the distance between the cameras 100. For example, if the distance between the cameras is 2km, the video analysis task can be scheduled to begin after 100 seconds based on the distance and speed.

[0168] The conditional parameters may include a similarity threshold. This similarity threshold can be used to determine whether a target in the first video and a target to be tracked in the second video are the same target. For example, if the similarity between the target and the target to be tracked is greater than the similarity threshold, they are determined to be the same target; otherwise, they are different targets. The similarity threshold can be set based on empirical values, for example, it can be set to 0.93.

[0169] In other possible implementations, action logic can also be used to instruct adjustment logic for camera 100. This adjustment logic includes adjusting the orientation and / or focal length of camera 100. For example, in a specific behavior detection and alarm scenario, the action logic may include adjusting the orientation and focal length of camera 100 to enable focusing on the target for behavior detection.

[0170] The following example illustrates how the video analysis device 206 performs a first video analysis task on the first video based on action logic and interaction data of the second video in a target tracking scenario.

[0171] Specifically, the action logic received by the video analysis device 206 from the management device 202 is as follows:

[0172]

[0173]

[0174] The video analysis device 206 can detect the target in the first video after 100 seconds according to the above-described action logic, and obtain information such as the target's attributes, feature values, and image. The video analysis device 206 compares the information of the target in the first video (attributes, feature values, or target image) with the information of the target to be tracked in the second video to determine the similarity between the target in the first video and the target to be tracked in the second video. This similarity is then compared with a similarity threshold to determine whether they are the same target. When the target appearing in the second video is the same target as the target appearing in the first video, the video analysis device 206 can generate the target's trajectory based on the video source location (specifically, the location of the camera 100), thereby achieving target detection and tracking.

[0175] In some possible implementations, a subscription relationship exists between the third camera and the first camera. The process of establishing this subscription relationship can be found in the relevant description above. The video analysis device 206 can also generate interactive data for the first video based on the analysis results of the first video. The interactive data for the first video may include information about targets appearing in the first video. This interactive data may include any one or more of the target's attributes, feature values, or target image. The video analysis device 206 can then send the interactive data of the first video to the management device 202. Thus, when the management device 202 analyzes the third video captured by the third camera, it can combine the aforementioned interactive data of the first video to perform a third video analysis task.

[0176] Based on the above description, the video analysis method provided in this application analyzes the interaction data obtained from other video analysis tasks during the execution phase of the video analysis task to obtain data associated with the aforementioned interaction data. Thus, when the application triggers tasks such as target search and target tracking, video analysis can be performed directly based on the associated data stored during the execution phase of the video analysis task, without needing to search and calculate a large number of feature values. This avoids explosive resource consumption and demand, eliminating the need for high-performance computing clusters and large-capacity resources, thereby improving resource utilization and reducing costs. Furthermore, analyzing the first video based on the interaction data of the second video enables collaborative analysis, narrowing the analysis scope, reducing computational load, and improving analysis efficiency.

[0177] Figure 6 The illustrated embodiment describes the video analysis method from the perspective of the video analysis device 206. When the camera is a smart camera, the above-described video analysis method can also be executed by each camera itself (e.g., the first camera). Specifically, the first camera acquires a first video, then acquires interaction data from a second video, and then performs a first video analysis task on the first video based on the interaction data from the second video. In some embodiments, the video analysis method can also be executed collaboratively by the camera and the video analysis device 206.

[0178] The first camera can acquire interactive data of the second video from the management device 202, for example, by receiving interactive data of the second video sent by the management device 202 based on the subscription relationship between the first and second cameras. In some embodiments, the second camera can be a smart camera, which can perform a second video analysis task on the second video to obtain the interactive data of the second video. Correspondingly, the first camera can also acquire interactive data of the second video from the second camera.

[0179] In some possible implementations, the first camera may also acquire motion logic, such as receiving motion logic sent by the management device 202, and then performing a first video analysis task on the first video based on the interaction data of the motion logic and the second video to obtain the analysis result of the first video.

[0180] When the first camera receives the action logic sent by the management device 202, it can also detect whether it possesses the function corresponding to the action logic. If it does, the first camera can perform a first video analysis task on the first video based on the interaction data between the action logic and the second video; if it does not, the first camera can forward the action logic to the video analysis device 206, which will then perform the first video analysis task on the first video based on the interaction data between the action logic and the second video.

[0181] Furthermore, a subscription relationship exists between the third camera and the first camera. The first camera can also generate interactive data for the first video based on the analysis results, and then send this interactive data to the third camera. Thus, when analyzing the third video captured by the third camera, the third camera can combine it with the interactive data from the first video for collaborative analysis. This enables continuous tracking of the target.

[0182] It should also be noted that, Figure 6 In the illustrated embodiment, the establishment of the subscription relationship and the distribution of action logic are demonstrated using the management device 202. When the cameras 100 are connected in a peer-to-peer (P2P) manner, they can directly subscribe to interactive data and distribute action logic.

[0183] In some possible implementations, the establishment of subscription relationships and the distribution of action logic can also be handled by video analytics applications (e.g., Figure 1 The video analytics application (300) is implemented as follows. The video analytics application can provide an action logic configuration interface, which can be a GUI or a CUI, allowing users to configure action logic. In this way, the video analytics application can obtain the action logic and then send it to the corresponding camera or to the video analytics device 206. Similarly, the video analytics application can establish subscription relationships between cameras 100 using a pre-configuration mode or a passive execution mode.

[0184] The video analysis method has been described in detail above. Next, the management method of video analysis will be introduced from the perspective of the management device (e.g., management device 202).

[0185] See Figure 8 The flowchart shown illustrates a video analytics management method, which includes:

[0186] S802: The management device 202 establishes a subscription relationship between the first camera and the second camera among multiple cameras.

[0187] The management device 202 can establish a subscription relationship between the first camera and the second camera in multiple modes, such as a pre-configuration mode and / or a passive execution mode. The pre-configuration mode and the passive execution mode are described in detail below.

[0188] In pre-configuration mode, the management device 202 receives a subscription instruction from the first camera, wherein the subscription instruction includes subscription parameters. Then, the management device 202 queries the second camera that satisfies the subscription parameters and establishes a subscription relationship between the first camera and the second camera. The subscription parameters may include distance parameters.

[0189] Specifically, when the first camera or video analysis device 206 initiates the first video analysis task, the first camera can send a subscription instruction to the management device 202. This subscription instruction includes subscription parameters, such as a distance parameter indicating that the distance to the first camera is less than a set distance. Thus, the management device 202 can determine the second camera that meets the subscription parameters based on the positional relationship between the cameras 100 and the aforementioned distance parameter, thereby establishing a subscription relationship between the first camera and the second camera.

[0190] Furthermore, the management device 202 can more accurately identify the second camera that has a subscription relationship with the first camera based on distance, combined with historical average vehicle speed and / or walking time.

[0191] In passive execution mode, the management device 202 receives a subscription instruction sent by the administrator and establishes a subscription relationship between the first camera and the second camera based on the subscription instruction. In some scenarios, the administrator can generate a subscription instruction through a programmable interface. This subscription instruction may include the identifier of the camera 100 to which the subscription relationship is to be established, and then send the subscription instruction to the management device 202. In this way, the management device 202 can directly establish a subscription relationship between the first camera and the second camera based on the aforementioned subscription instruction.

[0192] It should be noted that the video analysis and management method of this application embodiment may also omit the above-described S802. For example, after establishing a subscription relationship between the first camera and the second camera, the management device 202 can perform subsequent tasks based on this subscription relationship, without having to establish a subscription relationship for each task.

[0193] S804: The management device 202 receives the interactive data of the second video.

[0194] The second video is the video captured by the second camera. The interactive data of the second video can be obtained by the second camera or by the video analysis device 206 performing a second video analysis task on the second video. The management device 202 can obtain the interactive data of the second video from the second camera or the video analysis device 206.

[0195] The interactive data of the second video may include information about the target to be tracked that appears in the second video. The target to be tracked may be a specified target, or a target determined by the second camera or video analysis device 206 through matching targets appearing in the second video with monitored targets.

[0196] The types of interactive data can vary depending on the video analytics task. The following examples illustrate interactive data for different video analytics tasks.

[0197] For example, when the second video analysis task is target detection and tracking (person detection and tracking, vehicle detection and tracking, etc.), the interactive data of the second video can be any one or more of the following information: the attributes, feature values, and images of the target to be tracked. As another example, when the second video analysis task is target counting and statistics (crowd counting, vehicle counting, etc.), the interactive data of the second video can be one or more of the following information: the number of targets or the density of targets.

[0198] The above-mentioned S802 and S804 can be executed in parallel or in a set order. For example, the video analysis device 206 can execute S802 first and then S804, or execute S804 first and then S802. This application embodiment is only used as an example to illustrate the execution of S802 first and then S804.

[0199] S806: The management device 202 sends the interaction data of the second video to the first camera or the video analysis device 206 according to the subscription relationship.

[0200] When the first camera is a standard camera (without video analysis capabilities), the management device 202 sends the interaction data of the second video to the video analysis device 206 based on the subscription relationship between the first and second cameras. In this way, the video analysis device 206 can perform a first video analysis task on the first video based on the interaction data of the second video.

[0201] When the first camera is a smart camera (with video analytics capabilities), the management device 202 sends the interaction data of the second video to the first camera based on the subscription relationship between the first and second cameras. In some embodiments, the management device 202 may also send the interaction data of the second video to the video analytics device 206. Thus, either the first camera or the video analytics device 206 can perform a first video analytics task on the first video based on the interaction data of the second video.

[0202] In some possible implementations, the management device 202 may also send action logic to the first camera or the video analysis device, so that the first camera or the video analysis device performs a first video analysis task on the first video based on the interaction data of the second video and the action logic.

[0203] Before the management device 202 sends the action logic to the first camera or video analysis device, the management device 202 can obtain the action logic written by the administrator through a programmable interface. In some embodiments, the action logic can also be configured by a video analysis application.

[0204] To make the technical solution of this application clearer and easier to understand, the video analysis method of this application embodiment will be introduced from the perspective of interaction in combination with specific application scenarios below.

[0205] See Figure 9 The flowchart shown illustrates a video analysis method applied to a face-based trajectory search scenario, comprising the following steps:

[0206] S902: Camera A sends a subscription command to management device 202.

[0207] The subscription instruction includes subscription parameters. These parameters can be distance parameters. For example, if the distance parameter is within 2km, the subscription instruction instructs the management device 202 to establish a subscription relationship between camera A and cameras within 2km of camera A.

[0208] S904: Management device 202 queries for cameras that meet the subscription parameters.

[0209] The management device 202 can determine which cameras meet the subscription parameters based on their geographical location relationships. Specifically, if the management device 202 has the geographical location relationships of the cameras stored locally, it can search locally for cameras that meet the subscription parameters. If the management device 202 does not have the geographical location relationships of the cameras stored locally, it can send a query request to a location device. The location device can respond to the query request by returning a query response to the management device 202. This query response includes a list of cameras that meet the subscription parameters.

[0210] S906: Management device 202 completes subscription relationship record.

[0211] In this embodiment, it is assumed that camera B is a camera that meets the subscription parameters, for example, camera B is a camera within 2km of camera A. Correspondingly, the management device 202 can record the correspondence between camera A and camera B to establish a subscription relationship between camera A and camera B.

[0212] It should be noted that S902 to S906 are merely one implementation method for establishing a subscription relationship, specifically through a pre-configuration mode. In other possible implementations of this application embodiment, the management device 202 may also establish a subscription relationship through other methods, such as a passive execution mode.

[0213] S908: Camera B analyzes the video captured by Camera B and obtains the video analysis results.

[0214] Camera B is an intelligent camera that can perform video analysis tasks on the video captured by camera B, thereby analyzing the video and obtaining the analysis results of the video captured by camera B.

[0215] S909: Camera B reports the analysis results of the video captured by camera B to management device 202.

[0216] The analysis results include information about the face F. For example, the analysis results may include the feature values ​​of the face F and an image of the face F. In some possible implementations, the analysis results may also include information about the human body F, such as gender, clothing color, and other clothing information.

[0217] S910: The management device 202 sends the analysis results of the video captured by the camera B to the storage and search device 204 to store the analysis results.

[0218] S911: Management device 202 sends action logic to camera A.

[0219] Specifically, the management device 202 can receive action logic written by the administrator through the programmable interface provided by the management device 202. When the management device 202 detects that the similarity between face F and the monitored face reaches a preset threshold, it can send the action logic to a camera that has a subscription relationship with camera B, such as camera A. In some embodiments, the management device 202 can also send action logic to camera A based on the instruction from camera B.

[0220] It should be noted that the management device 202 can issue corresponding action logic according to actual needs. For example, when camera A has good face capture capability, the management device 202 can issue action logic to instruct face monitoring; when camera A has good human body capture capability, the management device 202 can issue action logic to instruct human body monitoring.

[0221] S912: The management device 202 sends interactive data, including information about the face F, to the camera A.

[0222] Specifically, based on the subscription relationship between camera A and camera B, management device 202 sends interactive data of the video captured by camera B to camera A. Management device 202 can obtain the interactive data of the video captured by camera B based on some or all of the information in the analysis results of the video captured by camera B. In some embodiments, this interactive data includes information about a face F.

[0223] S911 and S912 can be executed in parallel or sequentially according to a set order. For example, the management device 202 can first issue the action logic and then issue the interaction data; or the management device 202 can first issue the interaction data and then issue the action logic.

[0224] S914: Camera A caches action logic and interactive data including face information F.

[0225] Camera A caches the aforementioned action logic and interaction data for later comparison. Camera A can use the cached action logic and interaction data based on certain strategies, such as time-sensitivity strategies. In some possible implementations, camera A can use the cached action logic and interaction data after a set start time and delete it after a set deletion time. For example, if the start time is 10 minutes and the deletion time is 10 hours, camera A can use the aforementioned action logic and interaction data 10 minutes after the cache time expires and delete it after the cache time expires 10 hours later.

[0226] S916: Camera A analyzes the video captured by camera A according to the action logic to obtain information about the person in camera A.

[0227] Camera A is an intelligent camera that can perform video analysis tasks on the video captured by camera A based on the action logic issued by management device 202, thereby analyzing the video and obtaining information about the people in camera A. This information may include facial information. For example, the information about the people in camera A may include the feature values ​​of face G and the image of face G. In some possible implementations, the information about the people in camera A may also include information about the human body G, such as gender, clothing color, and other clothing details.

[0228] Specifically, the action logic includes at least one of the following: task type parameter, time parameter, and conditional parameter for executing the video analysis task. In this example, the task type is assumed to be face monitoring, the analysis starts after 100 seconds, and the similarity threshold is 0.95 (conditional parameter). Based on the task type and the analysis start time, camera A calls a face detection model to identify faces in the video captured by camera A, obtaining face information. This face information includes information about face G, such as the feature values ​​of face G and the image of face G.

[0229] S918: Camera A determines that face F and face G correspond to the same person based on the similarity between face F and face G.

[0230] Specifically, camera A can calculate the distance between the feature values ​​of face F and face G to determine the similarity between face F and face G. The similarity is then compared with a similarity threshold. If the similarity is greater than or equal to the similarity threshold, it is determined that face F and face G belong to the same person.

[0231] S920: Camera A reports the analysis results of the video captured by camera A to management device 202.

[0232] The analysis results of the video captured by camera A include information about face G. For example, the analysis results of the video captured by camera A include the feature values ​​of face G and the image of face G. The analysis results also include the correlation between face F and face G.

[0233] S922: The management device 202 sends the analysis results of the video captured by the camera A to the storage and search device 204 so that the storage and search device 204 stores the analysis results.

[0234] The analysis results contain the association between face F and face G. In this embodiment, it is assumed that face F and face G correspond to the same person.

[0235] S924: Application 300 sends a trajectory search request to storage and search device 204.

[0236] Specifically, application 300 can generate a trajectory search request in response to a user-triggered trajectory search operation. The trajectory search request can carry a facial image of the monitored target, such as a suspect. Then, application 300 sends the trajectory search request to storage and search device 204.

[0237] S926: Storage and search device 204 generates a trajectory based on the analysis results of the video captured by camera A and the video captured by camera B.

[0238] Specifically, the storage and search device 204 can perform a search based on the face image carried in the trajectory search request. When the face image matches the image of face F, the storage and search device 204 can connect the geographical locations of camera A and camera B according to the association between face F and face G to form a trajectory.

[0239] S928: The storage and search device 204 sends the trajectory to the application 300.

[0240] S930: Application 300 presents the trajectory to the user.

[0241] In some possible implementations, the camera can have built-in motion logic. For example, camera A can have built-in motion logic, so the video analysis method of this application embodiment may not need to execute S911 as described above. Camera A can perform video analysis tasks on the video captured by camera A based on the built-in motion logic.

[0242] In some possible implementations, cameras can also send action logic to other cameras to assist in completing certain tasks while performing their own logical processing. For example, after camera B detects a face, it acquires the face and body image of that person and notifies nearby cameras, such as camera A, to coordinate tracking of that person. Specifically, camera B sends interactive data and action logic to camera A. When there are relatively few cameras that can capture faces (or under certain conditions, it is impossible to capture faces that meet the requirements), but there are more cameras that can monitor the human body, collaborative analysis using cameras that can monitor the human body can better track the person.

[0243] Based on the above description, it can be seen that the embodiments of this application establish a subscription relationship to realize customized associated regions, and construct networked trajectory tracking by collaboratively analyzing the videos in the associated regions. On the one hand, this avoids massive feature value search and feature value comparison operations, avoiding explosive resource consumption and thus avoiding explosive resource demand; on the other hand, it improves analysis efficiency by reducing the amount of computation.

[0244] Next, see Figure 10The flowchart shown illustrates a video analysis method applied to specific behavior detection and alarm scenarios. The specific behavior could be the act of not wearing a mask. The method specifically includes the following steps:

[0245] S1002: Management device 202 receives action logic written by the administrator through the programming interface.

[0246] The management device 202 provides a programming interface. Administrators can use this interface to write action logic using a programming language. Specifically, the action logic can be logic for detecting and alarming specific behaviors. The specific behavior could be the act of not wearing a mask. The management device 202 can receive the action logic written by administrators through the programming interface. In some embodiments, the management device 202 can also receive action logic configured by administrators.

[0247] S1004: The management device 202 determines the camera corresponding to camera A based on the camera location information.

[0248] Specifically, the management device 202 can determine the cameras used for crowd counting and the cameras used for detecting mask-wearing behavior based on the camera location information, so as to establish a subscription relationship between the cameras. Among them, the camera used for crowd counting can be a high-point camera, which has a wider observation area, but the details in the image may not be clear. The camera used for detecting mask-wearing behavior, on the other hand, can clearly display the details in the image.

[0249] In this embodiment, it is assumed that the cameras within the monitored area include camera A and camera B. Camera A is a high-point camera within the area, used for crowd counting, and camera B is used to detect mask-wearing behavior.

[0250] S1006: Management device 202 completes subscription relationship record.

[0251] Specifically, the management device 202 can record the correspondence between camera A and camera B to establish a subscription relationship between them. In this way, the management device 202 can forward interactive data based on the subscription relationship, thereby enabling data interaction between camera A and camera B.

[0252] S1008: Camera A analyzes the video captured by Camera A and obtains the analysis results.

[0253] Camera A is an intelligent camera that can perform video analysis tasks on the videos it captures, thereby analyzing the video and obtaining analysis results. These results include the number of people in the video.

[0254] S1010: Camera A reports the analysis results of the video captured by camera A to management device 202.

[0255] S1011: Management device 202 sends action logic to camera B.

[0256] Specifically, the management device 202 can receive action logic written by the administrator through the programmable interface provided by the management device 202. The action logic includes at least one of the following: task type parameters, time parameters, and condition parameters for performing the video analysis task. In this example, it is assumed that the task type is specific behavior detection, such as mask-wearing behavior detection. The condition parameters include the task's triggering conditions. In this example, it is assumed that the triggering condition is that the number of people in the video exceeds a first threshold.

[0257] S1012: The management device 202 sends interactive data of the video captured by camera A to camera B.

[0258] The interactive data in the video captured by camera A comes from the analysis results of the video captured by camera A. This interactive data may be part or all of the information in the analysis results of the video captured by camera A. In some embodiments, the interactive data in the video captured by camera A may include the number of people in the video.

[0259] S1014: Camera B, based on action logic, determines whether the number of people in the video exceeds the first threshold. If so, proceed to S1016.

[0260] S1016: Camera B analyzes the video captured by Camera B based on action logic to detect people who are not wearing masks.

[0261] Camera B is an intelligent camera that can perform video analysis tasks on the videos it captures, thereby analyzing the videos and obtaining analysis results. These results include information on individuals not wearing masks. This information can include any one or more of the following: attributes, feature values, and facial images of the individuals not wearing masks.

[0262] S1018: Camera B reports the analysis results, including information on people not wearing masks, to management device 202.

[0263] S1020: The management device 202 sends the analysis results, including information on people not wearing masks, to the application 300.

[0264] S1022: Application 300 will issue an alarm based on information about people not wearing masks.

[0265] Specifically, application 300 can generate alarm messages based on information about people not wearing masks, typically presenting alarm information as a notification. In some embodiments, application 300 can also broadcast information about people not wearing masks to achieve alarm notification.

[0266] This embodiment uses the management device 202 to illustrate the forwarding of interactive data and the distribution of action logic. In other possible implementations of this application embodiment, camera A can also directly send the action logic and interactive data, including the number of people in the video, to camera B. Based on the action logic, camera B performs a step to determine whether the number of people in the video is greater than a first threshold, thereby determining whether to trigger mask-wearing behavior detection.

[0267] In some possible implementations, after camera A reports the video analysis results to management device 202, management device 202 can first execute a step to determine whether the number of people in the video exceeds a first threshold. If so, management device 202 sends action logic to camera B. Based on this action logic, camera B performs mask-wearing behavior detection on the video captured by camera B.

[0268] It should be noted that if the number of people in the video falls below a first threshold for a certain period of time, camera B can stop video analysis, for example, by stopping the detection of mask-wearing behavior, in order to free up resources. Specifically, camera A or management device 202 can instruct camera B to stop detecting mask-wearing behavior by sending action logic.

[0269] It should be noted that, Figure 9 , Figure 10 The embodiments shown all use cameras A and B as smart cameras for illustration. In some possible implementations, cameras A and B can also be ordinary cameras. In this case, the video analysis device 206 can also perform the analysis of the video captured by camera A and the video captured by other cameras. The video analysis device 206 can perform video analysis tasks on the corresponding videos based on action logic, thereby realizing video analysis.

[0270] The above text combined Figures 1 to 10 The video analysis method and video analysis management method provided in the embodiments of this application have been described in detail. The apparatus and equipment provided in the embodiments of this application will be described below with reference to the accompanying drawings.

[0271] See Figure 11 The schematic diagram of the video analysis device 206 shown indicates that the device 206 includes:

[0272] Communication unit 2062 is used to acquire the first video captured by the first camera;

[0273] The communication unit 2062 is also used to acquire interactive data of the second video, wherein the second video is captured by the second camera;

[0274] Analysis unit 2064 is used to perform a first video analysis task on the first video based on the interaction data of the second video to obtain the analysis result of the first video, wherein the interaction data of the second video is obtained by performing a second video analysis task on the second video.

[0275] In some possible implementations, the communication unit 2062 is further configured to:

[0276] Obtain action logic, wherein the action logic includes one or more of the following parameters: task type parameter for executing the first video analysis task, time parameter for executing the first video analysis task, and condition parameter for executing the first video analysis task;

[0277] The analysis unit 2064 is specifically used for:

[0278] Based on the action logic and the interaction data of the second video, a first video analysis task is performed on the first video.

[0279] In some possible implementations, the communication unit 2062 is specifically used for:

[0280] The system receives the action logic sent by the management device 202, wherein the action logic is written by the administrator according to the programming interface provided by the management device 202.

[0281] In some possible implementations, the analysis results of the first video include data associated with the interaction data of the second video.

[0282] In some possible implementations, the interactive data of the second video includes information about the target to be tracked that appears in the second video;

[0283] The analysis unit 2064 is specifically used for:

[0284] Identify the target in the first video and obtain information about the target;

[0285] Based on the information of the target and the information of the target to be tracked in the second video, the analysis result of the first video is obtained. The analysis result includes information of the target whose information similarity with the target to be tracked in the second video meets a preset condition.

[0286] In some possible implementations, the analysis unit 2064 is specifically used for:

[0287] Based on a pre-trained AI model, targets in the first video are identified, and information about the targets is obtained.

[0288] In some possible implementations, a subscription relationship exists between the first camera and the second camera, which is pre-built by the management device 202.

[0289] In some possible implementations, the analysis unit 2064 is further configured to:

[0290] The interactive data of the first video is generated based on the analysis results of the first video;

[0291] The communication unit 2062 is further configured to:

[0292] The interactive data of the first video is sent to the management device 202 or the third camera, wherein the third camera and the first camera have a subscription relationship.

[0293] In some possible implementations, the video analysis device 206 establishes a communication connection with the first camera and the second camera.

[0294] In some possible implementations, the video analysis device 206 is applied to security monitoring, and the first video analysis task includes one or more of the following tasks: person detection and tracking, vehicle detection and tracking, crowd counting and statistics, vehicle counting and statistics, and specific behavior detection and alarm.

[0295] The video analysis device 206 according to the embodiments of this application can correspondingly execute the methods described in the embodiments of this application, and the above and other operations and / or functions of each module / unit of the video analysis device 206 are respectively for implementing Figure 6 For the sake of brevity, the corresponding processes of each method in the illustrated embodiments will not be described in detail here.

[0296] See next Figure 12 The schematic diagram of the management device 202 shown is provided. The management device 202 is communicatively connected to multiple cameras. The device 202 includes:

[0297] The communication unit 2024 is used to receive interactive data of the second video, wherein the second video is captured by the second camera;

[0298] The communication unit 2024 is further configured to send the interactive data of the second video to the first camera or the video analysis device, so that the first camera or the video analysis device performs a first video analysis task on the first video based on the interactive data of the second video, wherein there is a subscription relationship between the first camera and the second camera, and the first video is captured by the first camera.

[0299] In some possible implementations, the device 202 further includes:

[0300] The subscription unit 2022 is used to establish a subscription relationship between the first camera and the second camera among the at least one cameras.

[0301] In some possible implementations, the subscription unit 2022 is specifically used for:

[0302] Receive a subscription instruction sent by the first camera, wherein the subscription instruction includes subscription parameters;

[0303] The system queries the second camera that satisfies the subscription parameters and establishes a subscription relationship between the first camera and the second camera.

[0304] In some possible implementations, the subscription unit 2022 is specifically used for:

[0305] Receive subscription instructions sent by administrators;

[0306] The subscription relationship between the first camera and the second camera is established according to the subscription instruction.

[0307] In some possible implementations, the communication unit 2024 is further used for:

[0308] The action logic is sent to the first camera or the video analysis device 206 so that the first camera or the video analysis device 206 performs a first video analysis task on the first video based on the interaction data of the second video and the action logic.

[0309] In some possible implementations, the communication unit 2024 is further used for:

[0310] Before sending the action logic to the first camera or video analysis device, obtain the action logic written by the administrator through the programmable interface, or obtain the action logic configured by the user through the video analysis application.

[0311] The management device 202 according to the embodiments of this application can correspondingly execute the methods described in the embodiments of this application, and the above and other operations and / or functions of each module / unit of the management device 202 are respectively for implementing Figure 8 For the sake of brevity, the corresponding processes of each method in the illustrated embodiments will not be described in detail here.

[0312] This application also provides a device 1300. The device 1300 can be a server or server cluster in a cloud environment, or a server or server cluster in a local data center. The device 1300 is specifically used to implement, for example... Figure 11 The video analysis device 206 in the illustrated embodiment has the following functions.

[0313] Figure 13 A structural schematic diagram of device 1300 is provided, as follows: Figure 13 As shown, device 1300 includes a bus 1301, a processor 1302, a communication interface 1303, and a memory 1304. The processor 1302, the memory 1304, and the communication interface 1303 communicate with each other via the bus 1301.

[0314] Bus 1301 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be divided into address buses, data buses, control buses, etc. For ease of representation, Figure 13 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0315] The processor 1302 can be any one or more of the following processors: central processing unit (CPU), graphics processing unit (GPU), microprocessor (MP), or digital signal processor (DSP).

[0316] The communication interface 1303 is used for communication with external devices. For example, the communication interface 1303 is used to acquire the first video captured by the first camera and the interactive data of the second video, or to receive the action logic sent by the management device 202 and send the interactive data of the first video to the management device 202, etc.

[0317] Memory 1304 may include volatile memory, such as random access memory (RAM). Memory 1304 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).

[0318] The memory 1304 stores executable code, and the processor 1302 executes the executable code to perform the aforementioned video analysis method.

[0319] Specifically, in achieving Figure 11 In the case of the illustrated embodiment, and Figure 11 When each unit of the video analysis device 206 described in the embodiment is implemented by software, it performs... Figure 11 The software or program code required for the functions of the communication unit 2062 and the analysis unit 2064 can be partially or entirely stored in the memory 1304. The processor 1302 executes the program code corresponding to each unit stored in the memory 1304 to perform the video analysis method.

[0320] This application also provides a device 1400. The device 1400 can be a server or server cluster in a cloud environment, or a server or server cluster in a local data center. The device 1400 is specifically used to implement, for example... Figure 12 The function of the management device 202 in the illustrated embodiment.

[0321] Figure 14 A structural schematic diagram of a device 1400 is provided, as follows: Figure 14 As shown, device 1400 includes a bus 1401, a processor 1402, a communication interface 1403, and a memory 1404. The processor 1402, the memory 1404, and the communication interface 1403 communicate with each other via the bus 1401.

[0322] The specific implementations of bus 1401, processor 1402, communication interface 1403, and memory 1404 can be found in [reference needed]. Figure 13 The illustrated embodiment is described below. Specifically, in the implementation... Figure 12 In the case of the illustrated embodiment, and Figure 12 When the modules or units of the management device 202 described in the embodiment are implemented by software, the following steps are performed: Figure 12The software or program code required for the functions of the subscription unit 2022 and communication unit 2024 can be partially or entirely stored in the memory 1404. The processor 1402 executes the program code corresponding to each unit stored in the memory 1404 to execute the video analysis management method.

[0323] This application also provides a camera 100, which can be a smart camera. Specifically, the camera 100 is used to implement the video analysis method provided in this application.

[0324] Figure 15 A structural schematic diagram of a camera 100 is provided, as follows: Figure 15 As shown, the camera 100 includes a bus 1501, a processor 1502, a communication interface 1503, a memory 1504, and an image sensor 1505. The processor 1502, memory 1504, communication interface 1503, and image sensor 1505 communicate with each other via the bus 1501.

[0325] The specific implementations of bus 1501, processor 1502, communication interface 1503, and memory 1504 can be found in [reference needed]. Figure 13 The illustrated embodiment is described below. The image sensor 1505 is a photosensitive element used to convert a light image on a photosensitive surface into an electrical signal proportional to the light image, thereby enabling video acquisition. The image sensor 1505 may include different types of sensors such as charge-coupled devices (CCDs) and complementary metal-oxide semiconductors (CMOS).

[0326] Specifically, the image sensor 1505 is used to acquire a first video. The processor 1502 can acquire the first video acquired by the image sensor 1505 through the bus 1501. The communication interface 1503 is used to acquire interactive data of a second video. The second video is captured by a second camera. The communication interface 1503 transmits the interactive data of the second video to the processor 1502 through the bus 1501. The processor 1502 executes the program code in the memory 1504, executes the first video analysis task based on the interactive data of the second video, and obtains the analysis result of the first video.

[0327] This application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store, or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct device 1300 to execute the video analysis method applied to video analysis device 206, or device 1400 to execute the management method applied to video analysis device 202, or camera 100 to execute the video analysis method.

[0328] This application also provides a computer program product comprising one or more computer instructions. When the computer instructions are loaded and executed on a computing device, all or part of the processes or functions described in this application are generated.

[0329] The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, or data center to another website, computer, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means.

[0330] When the computer program product is executed by a computer, the computer performs any of the aforementioned video analysis method or video analysis management method. The computer program product can be a software installation package; when it is necessary to use any of the aforementioned video analysis method or video analysis management method, the computer program product can be downloaded and executed on the computer.

[0331] The descriptions of the processes or structures corresponding to the above figures each have their own emphasis. For parts of a process or structure that are not described in detail, please refer to the relevant descriptions of other processes or structures.

Claims

1. A video analysis method, characterized in that, The method includes: Acquire the first video captured by the first camera; Acquire interactive data of the second video, wherein the second video is captured by the second camera, and there is a subscription relationship between the first camera and the second camera; During the execution phase of the video analysis task, based on the interaction data of the second video, the first video analysis task is performed on the first video to obtain the analysis result of the first video. The analysis result of the first video includes data associated with the interaction data of the second video, and the interaction data of the second video is obtained by performing the second video analysis task on the second video. In response to a trajectory search request, a trajectory of the target is generated based on the interaction data of the second video and the data associated with the interaction data of the second video.

2. The method according to claim 1, characterized in that, The method further includes: Obtain action logic, wherein the action logic includes one or more of the following parameters: task type parameter for executing the first video analysis task, time parameter for executing the first video analysis task, and condition parameter for executing the first video analysis task; The first video analysis task, performed on the first video based on the interaction data of the second video, includes: Based on the action logic and the interaction data of the second video, a first video analysis task is performed on the first video.

3. The method according to claim 2, characterized in that, The acquisition of action logic includes: receiving the action logic sent by the management device, wherein the action logic is written by the administrator according to the programming interface provided by the management device.

4. The method according to any one of claims 1-3, characterized in that, The interactive data of the second video includes information about the target to be tracked that appears in the second video; The step of performing a first video analysis task on the first video based on the interaction data of the second video to obtain the analysis results of the first video includes: Identify the target in the first video and obtain information about the target; Based on the information of the target and the information of the target to be tracked in the second video, the analysis result of the first video is obtained. The analysis result includes information of the target whose information similarity with the target to be tracked in the second video meets a preset condition.

5. The method according to claim 4, characterized in that, The step of identifying the target in the first video and obtaining information about the target includes: Based on a pre-trained artificial intelligence (AI) model, targets in the first video are identified, and information about the targets is obtained.

6. The method according to any one of claims 1-3, characterized in that, The subscription relationship is pre-built by the management device.

7. The method according to any one of claims 1-3, characterized in that, The method is executed by the first camera; the interactive data of the second video is obtained by the second camera performing a second video analysis task on the second video.

8. The method according to claim 7, characterized in that, The step of obtaining the interactive data for the second video includes: The first camera acquires interactive data from the second camera in the second video, or... The first camera obtains the interaction data of the second video from the management device.

9. The method according to any one of claims 1-3, characterized in that, The method further includes: The interactive data of the first video is generated based on the analysis results of the first video; The interactive data of the first video is sent to the management device or the third camera, wherein the third camera and the first camera have a subscription relationship.

10. The method according to any one of claims 1-3, characterized in that, The method is executed by a video analysis device, which establishes a communication connection with the first camera and the second camera.

11. The method according to any one of claims 1-3, characterized in that, The method is applied to security monitoring, and the first video analysis task includes one or more of the following tasks: person detection and tracking, vehicle detection and tracking, crowd counting and statistics, vehicle counting and statistics, and specific behavior detection and alarm.

12. A video analytics management method, characterized in that, The method is applied to a management device that is communicatively connected to multiple cameras, and the method includes: The management device receives interactive data of the second video, wherein the second video is captured by the second camera; During the execution phase of the video analysis task, the management device sends the interaction data of the second video to the first camera or the video analysis device, so that the first camera or the video analysis device performs a first video analysis task on the first video based on the interaction data of the second video to obtain the analysis result of the first video. The first camera and the second camera have a subscription relationship, and the first video is captured by the first camera. The analysis result of the first video includes data associated with the interaction data of the second video, which is stored in a storage and search device for generating a target trajectory based on the interaction data of the second video and the data associated with the interaction data of the second video in response to a trajectory search request.

13. The method according to claim 12, characterized in that, Before the management device receives the interactive data of the second video, the method further includes: The management device establishes a subscription relationship between the first camera and the second camera among the plurality of cameras.

14. The method according to claim 13, characterized in that, The management device establishes a subscription relationship between the first camera and the second camera among the plurality of cameras, including: The management device receives a subscription instruction sent by the first camera, wherein the subscription instruction includes subscription parameters; The management device queries the second camera that meets the subscription parameters according to the subscription parameters, and establishes a subscription relationship between the first camera and the second camera.

15. The method according to claim 13, characterized in that, The management device establishes a subscription relationship between the first camera and the second camera among the plurality of cameras, including: The management device receives subscription instructions sent by the administrator; The management device establishes a subscription relationship between the first camera and the second camera based on the subscription instruction.

16. The method according to any one of claims 12-15, characterized in that, The method further includes: The management device sends action logic to the first camera or the video analysis device, so that the first camera or the video analysis device performs the first video analysis task on the first video based on the interaction data of the second video and the action logic.

17. The method according to claim 16, characterized in that, Before the management device sends the action logic to the first camera or the video analysis device, the method further includes: The management device acquires action logic written by the administrator through the programmable interface of the management device, or the management device acquires action logic configured by the user through a video analytics application.

18. A video analysis device, characterized in that, include: The communication unit is used to acquire a first video captured by a first camera and to acquire interactive data of a second video during the execution phase of a video analysis task, wherein the second video is acquired by a second camera and there is a subscription relationship between the first camera and the second camera. An analysis unit is configured to perform a first video analysis task on the first video based on the interaction data of the second video, and obtain an analysis result of the first video. The analysis result of the first video includes data associated with the interaction data of the second video, which is obtained by performing a second video analysis task on the second video. The analysis result of the first video is stored in a storage and search device. The unit is configured to generate a trajectory of a target based on the interaction data of the second video and the data associated with the interaction data of the second video in response to a trajectory search request.

19. The apparatus according to claim 18, characterized in that, The communication unit is further used for: Obtain action logic, wherein the action logic includes one or more of the following parameters: task type parameter for executing the first video analysis task, time parameter for executing the first video analysis task, and condition parameter for executing the first video analysis task; The analysis unit is specifically used for: Based on the action logic and the interaction data of the second video, a first video analysis task is performed on the first video.

20. The apparatus according to claim 19, characterized in that, The communication unit is specifically used for: The system receives the action logic sent by the management device, wherein the action logic is written by the administrator according to the programming interface provided by the management device.

21. The apparatus according to any one of claims 18-20, characterized in that, The interactive data of the second video includes information about the target to be tracked that appears in the second video; The analysis unit is specifically used for: Identify the target in the first video and obtain information about the target; Based on the information of the target and the information of the target to be tracked in the second video, the analysis result of the first video is obtained. The analysis result includes information of the target whose information similarity with the target to be tracked in the second video meets a preset condition.

22. The apparatus according to claim 21, characterized in that, The analysis unit is specifically used for: Based on a pre-trained artificial intelligence (AI) model, targets in the first video are identified, and information about the targets is obtained.

23. The apparatus according to any one of claims 18-20, characterized in that, The subscription relationship is pre-built by the management device.

24. The apparatus according to any one of claims 18-20, characterized in that, The analysis unit is also used for: The interactive data of the first video is generated based on the analysis results of the first video; The communication unit is also used for: The interactive data of the first video is sent to the management device or the third camera, wherein the third camera and the first camera have a subscription relationship.

25. A management device, characterized in that, The management device is communicatively connected to multiple cameras, and the management device includes: A communication unit is configured to receive interactive data from a second video, wherein the second video is captured by a second camera; and to send the interactive data from the second video to a first camera or a video analysis device, so that the first camera or the video analysis device performs a first video analysis task on the first video based on the interactive data from the second video during the execution phase of the video analysis task. The analysis result of the first video includes data associated with the interactive data from the second video. There is a subscription relationship between the first camera and the second camera, and the first video is captured by the first camera. The analysis result of the first video is stored in a storage and search device, which is configured to generate a trajectory of a target based on the interactive data from the second video and the data associated with the interactive data from the second video in response to a trajectory search request.

26. The apparatus according to claim 25, characterized in that, The device further includes: The subscription unit is used to establish a subscription relationship between the first camera and the second camera among the plurality of cameras.

27. The management device according to claim 26, characterized in that, The subscription unit is specifically used for: Receive a subscription instruction sent by the first camera, wherein the subscription instruction includes subscription parameters; The system queries the second camera that satisfies the subscription parameters and establishes a subscription relationship between the first camera and the second camera.

28. The management device according to claim 26, characterized in that, The subscription unit is specifically used for: Receive subscription instructions sent by administrators; The subscription relationship between the first camera and the second camera is established according to the subscription instruction.

29. The management device according to any one of claims 25-28, characterized in that, The communication unit is also used for: The action logic is sent to the first camera or the video analysis device, so that the first camera or the video analysis device performs the first video analysis task on the first video based on the interaction data of the second video and the action logic.

30. The management device according to any one of claims 25-28, characterized in that, The communication unit is also used for: The system can acquire action logic written by administrators through the programmable interface of the management device, or acquire action logic configured by users through a video analytics application.

31. A device, characterized in that, The device includes a processor and a memory, the memory storing executable program code, the processor reading the executable program code stored in the memory to implement the function of the video analysis device according to any one of claims 18-24, or to implement the function of the management device according to any one of claims 25-30.

32. A camera, characterized in that, The camera includes a processor, a memory, and an image sensor. The image sensor is used to acquire a first video. The memory stores executable program code. The processor reads the executable program code to implement the function of the video analysis device according to any one of claims 18-24.

33. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed in the device, enable the device to perform the functions of the video analysis apparatus according to any one of claims 18-24, or the management apparatus according to any one of claims 25-30.

34. A computer program product comprising instructions, wherein when the instructions in the computer program product are executed by a device, the device performs the function of the video analysis device according to any one of claims 18-24, or performs the function of the management device according to any one of claims 25-30.