Video intelligent analysis method and system based on flexible analysis framework

By introducing a flexible analysis framework into the video intelligent analysis framework, users can customize the analysis process and adjust the algorithm, solving the shortcomings of analysis process and algorithm adjustment in the existing technology, and improving the practicality and scalability of video intelligent analysis.

CN117011767BActive Publication Date: 2025-06-06CHONGQING NORMAL UNIVERSITY +2
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202310978060.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-04
Publication Date
2025-06-06
Estimated Expiration
2043-08-04

AI Technical Summary

Technical Problem

The existing video intelligent analysis framework is poor in flexibility, and users are unable to customize the analysis process or adjust the algorithm, resulting in poor practicality and scalability of the analysis.

Method used

Using a video intelligent analysis method based on the flexible analysis framework, we can define user events and task lists to realize the customization of the video intelligent analysis process and user events, and assign tasks to the corresponding algorithm executors through the scheduler, and call corresponding algorithms and data to execute tasks.

Benefits of technology

It improves the practicality and scalability of intelligent video analysis, allows users to dynamically combine different algorithms to meet various scenario needs, realizes algorithm adjustment and modification, and enhances the ability of intelligent video surveillance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117011767B_ABST
    Figure CN117011767B_ABST
Patent Text Reader

Abstract

The present invention specifically relates to a video intelligent analysis method and system based on a flexible analysis framework. The method includes: pre-defining user events according to scene requirements, and determining the tasks and task lists that need to be executed when each user event is processed; when a certain video frame triggers a user event: parsing the triggered user event into a number of tasks to be executed; assigning the tasks to be executed to the corresponding algorithm executor; calling the corresponding algorithm and data according to the task list of the corresponding task to execute the task, and obtaining the corresponding algorithm execution result; when all tasks of the corresponding video frame are executed, outputting the algorithm execution results of all tasks of the corresponding video frame as its user event processing result; and using the user event processing results of all video frames in the video frame sequence as the video intelligent analysis result. The present invention can realize the customization of the analysis process and user events as well as the adjustment and modification of the algorithm, that is, it provides a flexible video intelligent analysis framework.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of big data and video image processing, and in particular to a video intelligent analysis method and system based on a flexible analysis framework. Background Art

[0002] The implementation of intelligent video surveillance mainly adopts artificial intelligence technologies such as pattern recognition, image processing, machine learning, and deep learning, combined with software engineering design ideas, and with the help of the powerful data processing capabilities of computers, it extracts key information from the video stream, analyzes and extracts the key information, so as to realize fully automatic and all-weather automatic judgment of abnormal situations in the monitoring screen, identify targets, quickly and accurately locate target positions, capture abnormal situations and other functions, thereby achieving intelligent effects such as pre-, mid-, and post-event early warning, processing, and evidence collection.

[0003] With the development of computer vision and artificial intelligence, more and more video analysis algorithms are embedded in surveillance cameras, enabling them to independently implement operations such as face detection, crowd counting, intelligent tracking, zooming, focusing, etc., or directly connect to video cloud servers to analyze and process video streams. On the other hand, the existing technology has also proposed and constructed intelligent video analysis systems, which mainly have two implementation architectures, namely edge computing and cloud computing architectures.

[0004] However, the existing technologies are mostly designed based on a single algorithm or a group of algorithms corresponding to relevant scenarios. In the actual video intelligent analysis process, users cannot define the analysis process, nor can they adjust or modify the algorithm. That is, the flexibility of the existing video intelligent analysis framework is poor, which leads to poor practicality and scalability of the analysis. Therefore, how to improve the practicality and scalability of video intelligent analysis is a technical problem that needs to be solved urgently. Summary of the invention

[0005] In view of the above-mentioned deficiencies in the prior art, the technical problem to be solved by the present invention is: how to provide a video intelligent analysis method and system based on a flexible analysis framework, which can realize the customization of analysis processes and user events and the adjustment and modification of algorithms, that is, to provide a flexible video intelligent analysis framework, thereby improving the practicality and scalability of video intelligent analysis, and providing a new idea for intelligent video surveillance.

[0006] In order to solve the above technical problems, the present invention adopts the following technical solutions:

[0007] The video intelligent analysis method based on the flexible analysis framework includes:

[0008] S1: Define user events in advance according to scenario requirements, and determine the tasks and task list to be performed when processing each user event; the task list includes the algorithm information and data required by the task;

[0009] S2: Pull the video stream to be processed and extract the video frame sequence from the video stream;

[0010] S3: Perform user event detection on the video frame sequence. When a video frame triggers a user event, the following user event processing flow is executed:

[0011] S301: Parsing the triggered user event into a number of tasks to be executed;

[0012] S302: Allocating the tasks to be executed to the corresponding algorithm executors through the scheduler;

[0013] S303: The algorithm executor calls the corresponding algorithm and data according to the task list of the corresponding task to execute the task, obtains the corresponding algorithm execution result and stores it;

[0014] If the algorithm execution result of a task corresponding to the video frame triggers a new user event, jump to step S3 and add a parallel user event processing flow for the corresponding video frame;

[0015] S304: Determine through the scheduler whether all tasks corresponding to the video frame have been executed: if so, output the algorithm execution results of all tasks corresponding to the video frame as its user event processing result; otherwise, return to step S302;

[0016] S4: The user event processing results of all video frames in the video frame sequence are used as video intelligent analysis results.

[0017] Preferably, if there are several tasks to be processed in sequence, the scheduler assigns each task in turn;

[0018] If there are several tasks to be processed in parallel, the scheduler will assign each task at the same time.

[0019] Preferably, the called data includes video stream information, a video frame sequence corresponding to the video stream, an algorithm execution result generated by a previous algorithm, and a video frame sequence that is finally processed.

[0020] Preferably, the task list also includes algorithm input requirements and algorithm output results;

[0021] Before the scheduler assigns a task, it determines whether the task can be executed based on the task list of the task.

[0022] Preferably, each video stream pulled is stored in a source stream data queue, and a unique identifier is assigned to each video stream;

[0023] Create a result set for the video frame that triggers the user event through the intermediate data queue, and store the algorithm execution results of all tasks of the video frame in the corresponding result set;

[0024] The final data queue stores all the video frames for which tasks have been completed and their user event processing results, and each video frame is arranged in the order of the original video frame sequence.

[0025] The present invention also discloses a video intelligent analysis system based on a flexible analysis framework, which is implemented based on the video intelligent analysis method based on the flexible analysis framework in the present invention, including:

[0026] Algorithm Manager, which is used to store and manage all algorithms set up;

[0027] An event detection module is used to extract a video frame sequence from a video stream and perform user event detection on the video frame sequence;

[0028] The task definer is used to parse the triggered user event into several tasks to be executed when a certain video frame is detected to trigger a user event;

[0029] The task definer predefines user events according to scenario requirements and determines the tasks and task lists that need to be performed when processing each user event; the task list contains the algorithm information and data required by the task;

[0030] The scheduler is used to obtain the tasks to be executed from the task definer and assign them to the corresponding algorithm executors;

[0031] An algorithm executor is used to call the corresponding algorithm and data according to the task list of the corresponding task to execute the task, obtain the corresponding algorithm execution result and store it;

[0032] The data management module is used to store the pulled video stream and the algorithm execution results generated after the task is executed;

[0033] Finally, it is determined whether all tasks corresponding to the video frame have been completed: if so, the algorithm execution results of all tasks corresponding to the video frame are output as the user event processing results; otherwise, the scheduler continues to obtain and assign tasks.

[0034] Preferably, the algorithm manager performs unified interface management on each algorithm according to the source, type, function, operation efficiency and conditions of the algorithm; each algorithm has a unique algorithm signature.

[0035] Preferably, the task list also includes algorithm input requirements and algorithm output results;

[0036] Before the scheduler assigns a task, it determines whether the task can be executed based on the task list of the task.

[0037] Preferably, the data management module divides the queues according to the unique identifier of the video stream, and allocates a queue for each video stream at the source stream data, intermediate data and final data stages;

[0038] Each video stream pulled is stored in the source stream data queue, and a unique identifier is assigned to each video stream;

[0039] Create a result set for the video frame that triggers the user event through the intermediate data queue, and store the algorithm execution results of all tasks of the video frame in the corresponding result set;

[0040] The final data queue stores all the video frames for which tasks have been completed and their user event processing results, and each video frame is arranged in the order of the original video frame sequence.

[0041] Preferably, the scheduler places the task processing results output by the algorithm executor into the intermediate data queue of the data management module: if the current task is the only task or the last task of the current video frame, the task processing results of the current video frame and all its tasks are placed in the final data queue; otherwise, check whether there is a corresponding video frame in the intermediate data queue. If not, place the corresponding video frame in the intermediate data queue and place the corresponding task processing results in the result set of the video frame; it is also used to place the video frame and its user event processing results into the final data queue when it is detected that all tasks of the corresponding video frame have been executed or the set task processing threshold has been reached.

[0042] Compared with the prior art, the video intelligent analysis method and system based on the flexible analysis framework in the present invention have the following beneficial effects:

[0043] The present invention realizes the customization of video intelligent analysis process and user events by defining user events and determining the tasks that need to be performed when processing user events. The task list includes the algorithm information and data required for the task, so that the corresponding algorithm and data can be called when the task is subsequently executed. Since the user event is definable, the various algorithms involved in the entire user event analysis process can also be defined and combined, rather than being a processing process defined by the program during programming. The algorithm can be adjusted and modified (for example, after a face detection event, a task of counting the number of people present can be flexibly added, that is, an algorithm for counting the number of people entering and leaving can be added), and different algorithms can be dynamically combined to meet the needs of various scenarios, that is, a flexible video intelligent analysis framework is provided, thereby improving the practicality and extensibility of video intelligent analysis, and providing a new idea for intelligent video surveillance.

[0044] The present invention assigns the tasks to be executed to the algorithm executor through the scheduler, and then the algorithm executor calls the corresponding algorithm and data to execute the tasks. The algorithm executor can be distributed, so that the algorithm execution tasks can be assigned to different distributed computers for execution, so as to achieve the effect of system scalability and ensure system throughput. In addition, the flexible video intelligent analysis framework of the present invention can control the number of video frames and algorithms detected, and can realize the processing of multiple different video streams in one system, thereby further improving the practicality and scalability of video intelligent analysis.

[0045] The present invention can also trigger new user events during the task execution process (for example, when executing the face recognition task of the face detection event, if a stranger is detected, a stranger event is triggered, which in turn generates a new execution task), that is, user events are decomposed into tasks, tasks trigger new events, and new events are decomposed into new tasks, thereby triggering new processing flows. This further reflects the flexibility of the video intelligent analysis framework of the present invention.

[0046] The present invention divides the video intelligent analysis system into functional modules such as algorithm management, data management, event detection, task definition, task scheduling and algorithm execution, which meets the single responsibility of the functional modules and avoids excessive coupling. It can abstractly design the system architecture from a higher level, so that the system can flexibly guarantee the execution of each task, and then analyze the full-scene content of the video in more detail and more actively. At the same time, it can realize flexible streaming processing through user event definition and task scheduling of the task definer, rather than a rigid algorithm pulling a stream for rough calculation, thereby improving the quality of video intelligent analysis and better realizing intelligent video monitoring. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to make the purpose, technical solution and advantages of the invention more clear, the present invention will be further described in detail below with reference to the accompanying drawings, in which:

[0048] Figure 1 It is a logic block diagram of a video intelligent analysis method based on a flexible analysis framework;

[0049] Figure 2 This is the architecture diagram of the video intelligent analysis system based on the flexible analysis framework;

[0050] Figure 3 It is a processing flow chart of the video flexible processing system;

[0051] Figure 4 Flowchart for the scheduler to perform flexible scheduling;

[0052] Figure 5 Generates a flowchart of the task processing flow for the task definer. DETAILED DESCRIPTION

[0053] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. The components of the embodiments of the present invention generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but only represents selected embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work belong to the scope of protection of the present invention.

[0054] It should be noted that similar numbers and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in the subsequent drawings. In the description of the present invention, it should be noted that the orientation or position relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inside", "outside", etc. is based on the orientation or position relationship shown in the drawings, or the orientation or position relationship in which the invention product is usually placed when used, which is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation of the present invention. In addition, the terms "first", "second", "third", etc. are only used to distinguish the description, and cannot be understood as indicating or implying relative importance. In addition, the terms "horizontal", "vertical", etc. do not mean that the components are required to be absolutely horizontal or suspended, but can be slightly tilted. For example, "horizontal" only means that its direction is more horizontal than "vertical", and does not mean that the structure must be completely horizontal, but can be slightly tilted. In the description of the present invention, it is also necessary to explain that, unless otherwise clearly specified and limited, the terms "set", "install", "connect", and "connect" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two elements. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0055] The following is a further detailed description through specific implementation methods:

[0056] Embodiment 1:

[0057] This embodiment discloses a video intelligent analysis method based on a flexible analysis framework.

[0058] like Figure 1 As shown, the video intelligent analysis method based on the flexible analysis framework includes:

[0059] S1: Define user events in advance according to scenario requirements, and determine the tasks and task list to be performed when processing each user event; the task list includes the algorithm information and data required by the task;

[0060] S2: Pull the video stream to be processed and extract the video frame sequence from the video stream;

[0061] S3: Perform user event detection on the video frame sequence. When a certain video frame triggers a user event, the following user event processing flow is entered:

[0062] In this embodiment, the triggering of user events can be divided into two types. One is triggered by an external sensor or generated by a single frame image. For example, when the infrared sensor corresponding to the camera senses a moving object, a face detection event is triggered. It can also detect faces for each frame of the image. As long as there is a frame input, the detection event will be triggered. The other is generated by the algorithm execution result of other events. For example, after the face detection event detects a face, the stranger recognition task will be executed. When a stranger is recognized, the stranger event is triggered, which will generate a new execution task.

[0063] S301: Parsing the triggered user event into a number of tasks to be executed;

[0064] S302: Allocating the tasks to be executed to the corresponding algorithm executors through the scheduler;

[0065] In this embodiment, if there are several tasks to be processed in sequence, the scheduler allocates each task in sequence; if there are several tasks to be processed in parallel, the scheduler allocates each task at the same time.

[0066] The algorithm executor is responsible for calling a specific algorithm to execute a task instance. The specific algorithm executor that executes a task instance is determined by the scheduler. The algorithm executor obtains the algorithm type and version from the task and the algorithm program from the algorithm manager. For example, the scheduler assigns task instance A to algorithm executor B for execution, but executor B does not have the algorithm of task instance A, so it needs to request the algorithm program from the algorithm manager. Usually, the algorithm program is cached after it is obtained, and it does not need to be obtained next time unless an algorithm version mismatch is detected.

[0067] S303: The algorithm executor calls the corresponding algorithm and data according to the task list of the corresponding task to execute the task, obtains the corresponding algorithm execution result and stores it;

[0068] In this embodiment, if the algorithm execution result of a certain task triggers a new user event, the process jumps to step S3 and adds a parallel user event processing flow for the corresponding video frame.

[0069] The algorithm executor is a program that runs on different machines to execute the algorithm. The algorithm executor obtains the input of the algorithm, starts the algorithm execution, obtains the result and processes the result. Usually, the algorithm executor runs on multiple machines, and the scheduler manages these algorithm executors. The scheduler determines which algorithm executor will execute according to the executor's capabilities (for example: network bandwidth, CPU and GPU performance, memory size, etc.) and load (CPU, GPU, memory usage).

[0070] S304: Determine through the scheduler whether all tasks corresponding to the video frame have been executed: if so, output the algorithm execution results of all tasks corresponding to the video frame as its user event processing result; otherwise, return to step S302;

[0071] In this embodiment, the algorithm is finally executed on the algorithm executor, which is responsible for starting the execution of the algorithm and monitoring the execution status of the algorithm: in progress, normal termination, abnormal termination, etc. The scheduler communicates with the algorithm executor to monitor the execution status of the task. If the algorithm execution is completed, the scheduler can start the execution of the next task.

[0072] S4: The user event processing results of all video frames in the video frame sequence are used as video intelligent analysis results.

[0073] It should be noted that the algorithm refers to a program that extracts information from video frames or transforms video frames. The program completes inseparable functions, such as an algorithm that converts the image of a video frame into a grayscale image, and an algorithm that detects faces from a video frame.

[0074] The user event refers to the event that the user wants to trigger in the video stream, such as a face detection event. After a face is detected in the video, a series of tasks need to be processed, such as a helmet wearing detection task, a stranger recognition task, etc. When executing a task, a new user event may be triggered. For example, after a stranger is detected, a stranger event is triggered, which in turn generates a new execution task.

[0075] The task is what needs to be done in response to user events, such as the helmet wearing detection task mentioned above. The core of this task is the helmet detection algorithm. The task also defines how to obtain input, such as the result of face detection as the input of the helmet detection algorithm. The task also defines how to process the output of the algorithm, such as the output of the anchor frame coordinates of the face not wearing a helmet when the helmet detection algorithm determines that it is to be worn in the future. Specifically, all the minimum processing units are abstracted into tasks, and each task has a task list, which contains three contents: data references that need to be processed by the algorithm, the current algorithm, the input requirements of the algorithm, and the processing of the algorithm output.

[0076] The present invention realizes the customization of video intelligent analysis process and user events by defining user events and determining the tasks that need to be performed when processing user events. The task list includes the algorithm information and data required for the task, so that the corresponding algorithm and data can be called when the task is subsequently executed. Since the user event is definable, the various algorithms involved in the entire user event analysis process can also be defined and combined, rather than being a processing process defined by the program during programming. The algorithm can be adjusted and modified (for example, after a face detection event, a task of counting the number of people present can be flexibly added, that is, an algorithm for counting the number of people entering and leaving can be added), and different algorithms can be dynamically combined to meet the needs of various scenarios, that is, a flexible video intelligent analysis framework is provided, thereby improving the practicality and extensibility of video intelligent analysis, and providing a new idea for intelligent video surveillance.

[0077] The present invention assigns the tasks to be executed to the algorithm executor through the scheduler, and then the algorithm executor calls the corresponding algorithm and data to execute the tasks. The algorithm executor can be distributed, so that the algorithm execution tasks can be assigned to different distributed computers for execution, so as to achieve the effect of system scalability and ensure system throughput. In addition, the flexible video intelligent analysis framework of the present invention can control the number of video frames and algorithms detected, and can realize the processing of multiple different video streams in one system, thereby further improving the practicality and scalability of video intelligent analysis.

[0078] The present invention can also trigger new user events during the task execution process (for example, when executing the face recognition task of the face detection event, if a stranger is detected, a stranger event is triggered, which in turn generates a new execution task), that is, user events are decomposed into tasks, tasks trigger new events, and new events are decomposed into new tasks, thereby triggering new processing flows. This further reflects the flexibility of the video intelligent analysis framework of the present invention.

[0079] During the specific implementation process, the task list also includes algorithm input requirements and algorithm output results;

[0080] Before the scheduler assigns a task, it determines whether the task can be executed based on the task list of the task.

[0081] In this embodiment, the task includes the definition of the algorithm input, such as the color space, width, height, bit depth, etc. of the input image. If the input does not meet the input requirements of the algorithm, the task is not executable. Usually, the input of a task is the output of the previous task. The algorithm of the previous task does not guarantee that its output meets the input requirements of the next algorithm, so the algorithm is not executable.

[0082] In the present invention, before assigning a task, the scheduler determines whether the task is executable according to the task list, and then can screen out executable tasks for assignment, thereby improving the efficiency of user event processing and task execution.

[0083] During the specific implementation process, a queue is allocated to each video stream at the source stream data, intermediate data and final data stages.

[0084] 1) Source stream data queue

[0085] Each video stream pulled is stored through the source stream data queue, and a unique identifier is assigned to each video stream.

[0086] In this embodiment, all cameras connected to the system are managed in a unified manner. In this system (FRTA architecture), each camera will only pull one stream. Although current high-definition cameras can support the simultaneous pulling of dozens of streams, pulling too many video streams at the same time will pose a greater challenge to network bandwidth and computer memory. At the same time, repeated processing operations will be generated when performing algorithm analysis, which will also cause a waste of computing resources. Therefore, this system adopts the solution of pulling only one video stream from one camera.

[0087] The source stream data queue manages all video streams in a unified manner, and maintains a piece of meta-information (description information) for each stream, which is similar to the file control block in file management. The description information includes the source of the video stream (camera-related information and the unit to which the camera belongs, etc.), the video encapsulation format, encoder, decoder, frame resolution, bit rate, frame rate, stream pull address, stream push address, stream puller and stream pusher, etc. At the same time, in order to ensure the communication pressure between the various modules of the system, a unique identifier is assigned to each stream.

[0088] Each video stream caches a frame sequence in the system, which is the video frame sequence pulled from the camera. Each frame of the video will carry a lot of information when it is captured from the camera, but only some information related to the picture is commonly used in frame sequence or picture processing, such as picture width and height, picture pixel format, picture step, picture depth, etc. When copying or transmitting pictures, this system only transmits this necessary information to reduce bandwidth costs, and at the same time assigns a unique identifier in the stream to each frame of the picture as the position of the video stream to which the picture belongs.

[0089] It is impossible to cache all the video frames of a stream in the memory. On the one hand, the previous frames have been processed and there is no need to cache them. On the other hand, the stream puller is still grabbing video frames from the camera. If these video frames are not cached in time, it may block the subsequent video frame capture and affect the processing flow of the video stream, which will cause the real-time video playback to appear stuck, frame loss, and even cause memory overflow errors. Therefore, this system will save the processed video frames in time.

[0090] 2) Intermediate data queue

[0091] A result set is created for the video frame that triggers the user event through the intermediate data queue, and the algorithm execution results of all tasks of the video frame are stored in the corresponding result set.

[0092] In this embodiment, the intermediate data is the result generated by the algorithm in addition to the initial video stream and the final data. The reason for separating the results generated by the algorithm is also based on the principle of single responsibility, placing data from different periods in different places for easy management and maintenance.

[0093] When there is a video frame in the source stream queue, the system dequeues the video frame and puts it into the intermediate data queue to wait until all tasks on the frame are completed. If the scheduler detects that the task on the frame is not completed, it will continue to wait for the task to be executed or hand over the unexecuted task to the algorithm executor. The scheduler will write the results of the completed task into the collection associated with the frame; when all tasks on the frame are completed or timed out, the frame will be dequeued from the intermediate data queue and placed in the final data queue.

[0094] 3) Final data queue

[0095] The final data queue stores all the video frames for which tasks have been completed and their user event processing results, and each video frame is arranged in the order of the original video frame sequence.

[0096] In this embodiment, the final data is composed of the original video frame sequence and the corresponding text data, which indicates that the video frame has been processed and all the data needs to be arranged in the time sequence of the original video frame.

[0097] Generally, the frame rate of the streaming video is 25 frames per second, and the interval between one frame is 40 milliseconds. Taking into account the time consumption of unprocessed streams, an initial waiting time of 30 milliseconds is set in the architecture of this article, that is, the maximum waiting time for the stream pusher or data storage is 30 milliseconds. If the minimum video frame in the current queue is not the next frame of the pushed video frame, this frame is pushed directly without waiting.

[0098] The architecture provides a configurable waiting time. Users can manually or the system can automatically adjust the appropriate waiting time. The adjustment strategy is to monitor the task completion time, task generation and scheduling time on the video frame. Generally, when pushing the real-time stream to users for viewing, adjustment is required to ensure the smoothness of the video. Therefore, the total processing time of a frame of image should not exceed 40 milliseconds or the frame interval time. There are also some strategies in scheduling to ensure the smoothness of the video.

[0099] The present invention stores video streams and video frame sequences before, during and after processing respectively through source stream data queues, intermediate data queues and final data queues, so that all data in the video intelligent analysis process can be centrally managed, and the intermediate data generated by the algorithm can be effectively reused, thereby improving data management effects and making full use of system computing resources to reduce equipment overhead.

[0100] Embodiment 2:

[0101] This embodiment discloses a video intelligent analysis system based on a flexible analysis framework, which is implemented based on the video intelligent analysis method based on a flexible analysis framework in the first embodiment.

[0102] like Figure 2 and Figure 3 As shown, the video intelligent analysis system based on the flexible analysis framework includes:

[0103] Algorithm Manager, which is used to store and manage all algorithms set up;

[0104] The data management module is used to store the pulled video stream and the algorithm execution results generated after the task is executed;

[0105] An event detection module is used to extract a video frame sequence from a video stream and perform user event detection on the video frame sequence;

[0106] The task definer is used to parse the triggered user event into several tasks to be executed when a certain video frame is detected to trigger a user event; the task definer pre-defines the user event according to the scene requirements, and determines the tasks and task list to be executed when each user event is processed; the task list contains the algorithm information and data required by the algorithm;

[0107] In this embodiment, the triggering of user events can be divided into two types. One is triggered by an external sensor or generated by a single frame image. For example, when the infrared sensor corresponding to the camera senses a moving object, a face detection event is triggered. It can also detect faces for each frame of the image. As long as there is a frame input, the detection event will be triggered. The other is generated by the algorithm execution result of other events. For example, after the face detection event detects a face, the stranger recognition task will be executed. When a stranger is recognized, the stranger event is triggered, which will generate a new execution task.

[0108] The scheduler is used to obtain the tasks to be executed from the task definer and assign them to the corresponding algorithm executors;

[0109] An algorithm executor is used to call the corresponding algorithm and data according to the task list of the corresponding task to execute the task, obtain the corresponding algorithm execution result and store it;

[0110] In this embodiment, if there are several tasks to be processed in sequence, the scheduler allocates each task in sequence; if there are several tasks to be processed in parallel, the scheduler allocates each task at the same time.

[0111] The algorithm executor is responsible for calling a specific algorithm to execute a task instance. The specific algorithm executor that executes a task instance is determined by the scheduler. The algorithm executor obtains the algorithm type and version from the task and the algorithm program from the algorithm manager. For example, the scheduler assigns task instance A to algorithm executor B for execution, but executor B does not have the algorithm of task instance A, so it needs to request the algorithm program from the algorithm manager. Usually, the algorithm program is cached after it is obtained, and it does not need to be obtained next time unless an algorithm version mismatch is detected.

[0112] The algorithm executor is a program that runs on different machines to execute the algorithm. The algorithm executor obtains the input of the algorithm, starts the algorithm execution, obtains the result and processes the result. Usually, the algorithm executor runs on multiple machines, and the scheduler manages these algorithm executors. The scheduler determines which algorithm executor will execute according to the executor's capabilities (for example: network bandwidth, CPU and GPU performance, memory size, etc.) and load (CPU, GPU, memory usage).

[0113] Finally, it is determined whether all tasks corresponding to the video frame have been completed: if so, the algorithm execution results of all tasks corresponding to the video frame are output as the user event processing results; otherwise, the scheduler continues to obtain and assign tasks.

[0114] In this embodiment, the algorithm is finally executed on the algorithm executor, which is responsible for starting the execution of the algorithm and monitoring the execution status of the algorithm: in progress, normal termination, abnormal termination, etc. The scheduler communicates with the algorithm executor to monitor the execution status of the task. If the algorithm execution is completed, the scheduler can start the execution of the next task.

[0115] It should be noted that the algorithm refers to a program that extracts information from video frames or transforms video frames. The program completes inseparable functions, such as an algorithm that converts the image of a video frame into a grayscale image, and an algorithm that detects faces from a video frame.

[0116] The user event refers to the event that the user wants to trigger in the video stream, such as a face detection event. After a face is detected in the video, a series of tasks need to be processed, such as a helmet wearing detection task, a stranger recognition task, etc. When executing a task, a new user event may be triggered. For example, after a stranger is detected, a stranger event is triggered, which in turn generates a new execution task.

[0117] The task is what needs to be done in response to user events, such as the helmet wearing detection task mentioned above. The core of this task is the helmet detection algorithm. The task also defines how to obtain input, such as the result of face detection as the input of the helmet detection algorithm. The task also defines how to process the output of the algorithm, such as the output of the anchor frame coordinates of the face not wearing a helmet when the helmet detection algorithm determines that it is to be worn in the future. Specifically, all the minimum processing units are abstracted into tasks, and each task has a task list, which contains three contents: data references that need to be processed by the algorithm, the current algorithm, the input requirements of the algorithm, and the processing of the algorithm output.

[0118] The present invention divides the video intelligent analysis system into functional modules such as algorithm management, data management, event detection, task definition, task scheduling and algorithm execution, which meets the single responsibility of the functional modules and avoids excessive coupling. It can abstractly design the system architecture from a higher level, so that the system can flexibly guarantee the execution of each task, and then analyze the full-scene content of the video in more detail and more actively. At the same time, it can realize flexible streaming processing through user event definition and task scheduling of the task definer, rather than a rigid algorithm pulling a stream for rough calculation, thereby improving the quality of video intelligent analysis and better realizing intelligent video monitoring.

[0119] The present invention realizes the customization of video intelligent analysis process and user events by defining user events and determining the tasks that need to be performed when processing user events. The task list includes the algorithm information and data required for the task, so that the corresponding algorithm and data can be called when the task is subsequently executed. Since the user event is definable, the various algorithms involved in the entire user event analysis process can also be defined and combined, rather than being a processing process defined by the program during programming. The algorithm can be adjusted and modified (for example, after a face detection event, a task of counting the number of people present can be flexibly added, that is, an algorithm for counting the number of people entering and leaving can be added), and different algorithms can be dynamically combined to meet the needs of various scenarios, that is, a flexible video intelligent analysis framework is provided, thereby improving the practicality and extensibility of video intelligent analysis, and providing a new idea for intelligent video surveillance.

[0120] The present invention assigns the tasks to be executed to the algorithm executor through the scheduler, and then the algorithm executor calls the corresponding algorithm and data to execute the tasks. The algorithm executor can be distributed, so that the algorithm execution tasks can be assigned to different distributed computers for execution, so as to achieve the effect of system scalability and ensure system throughput. In addition, the flexible video intelligent analysis framework of the present invention can control the number of video frames and algorithms detected, and can realize the processing of multiple different video streams in one system, thereby further improving the practicality and scalability of video intelligent analysis.

[0121] The present invention can also trigger new user events during the task execution process (for example, when executing the face recognition task of the face detection event, if a stranger is detected, a stranger event is triggered, which in turn generates a new execution task), that is, user events are decomposed into tasks, tasks trigger new events, and new events are decomposed into new tasks, thereby triggering new processing flows. This further reflects the flexibility of the video intelligent analysis framework of the present invention.

[0122] 1. Algorithm Manager

[0123] In the specific implementation process, the algorithm manager performs unified interface management for each algorithm based on the source, type, function, operating efficiency and conditions of the algorithm; each algorithm has a unique algorithm signature, which consists of the algorithm name and version number. Because the algorithm executor is responsible for the execution of all algorithms, this requires that the algorithms must have the same interface form, such as naming specifications, parameter transfer methods, etc. The unified interface can ensure that the algorithm executor can start the algorithm execution in a unified way.

[0124] Due to the different sources, types and functions of algorithms, there are some differences between algorithms besides logic. Some algorithms come from already developed frameworks, such as OpenCV; some algorithms come from models trained by users themselves; or users choose to use models provided by certain cloud platforms. At the same time, there are also some types of distinctions in algorithms. Algorithms may be conventional logic processing algorithms, or they may be machine learning or deep learning algorithms. Conventional algorithms may not have high requirements for GPUs, but the performance differences of the latter on CPUs and GPUs cannot be ignored. In addition, the effects of algorithms are also different. An algorithm with particularly high accuracy may take longer, while an algorithm with poor accuracy can also meet general needs.

[0125] Therefore, the algorithm manager manages the algorithms in a unified interface according to the source, type, function, operating efficiency and conditions of the algorithms, and can adapt to multi-threading and sequential execution. It also automatically adapts to the devices that the user can provide and selects the appropriate algorithm to provide differentiated services. The algorithm manager manages the algorithms according to the algorithm signature. The algorithm signature consists of the algorithm name and version number and is unique. Users will have a uniqueness check when registering the algorithm. Algorithms with the same function use the same name, but their version numbers cannot be the same to distinguish them from each other.

[0126] In addition, the algorithm needs to explain its input and output requirements. Because this system is aimed at video stream processing, algorithms whose input is non-image or video frame sequence are not considered. In terms of algorithm input, there are differences in image color space, width, height, step size, bit depth, etc. In terms of algorithm output, due to different algorithm sources, different processing is required. For algorithms embedded during system implementation, text information can be directly output. For third-party applications that output images or videos, they can be marked or the original video frames can be replaced. The algorithm manager in the system is more about maintaining the existence of the algorithm. Other modules need to make corresponding adjustments to the input and output of the algorithm.

[0127] In order to ensure that the execution of the algorithm is not restricted by the machine, all algorithms in the system are adjusted to be stateless. Although some algorithms have a logical order, the requirements for images are not consistent. In addition, in order to reuse images or videos as much as possible, the image itself cannot be modified unless the image is only used by the current algorithm. All algorithm data is stored in the blackboard, and all algorithms get data from the blackboard.

[0128] 2. Task definition

[0129] User events refer to events that users set on a video stream in the front end. Events here are events that users want to trigger in the video stream. All events will save information such as users, cameras, video streams, and algorithms in the user event manager. After being parsed by the task parser, user events will become several tasks, which will be bound to the video stream and written into the task definer.

[0130] The task definer mainly manages the relationship between video streams and algorithms. There are one-to-one, one-to-many, many-to-one or many-to-many relationships between video streams and algorithms. The task definer maintains an optimal correspondence based on the settings between video streams and algorithms in order to reduce the number of repeated calculations and data memory usage. The task definer uses a directed acyclic graph to associate the relationship between tasks. There are peer and sequential relationships between tasks, such as Figure 4 shown.

[0131] 2.1 User Events

[0132] The user sets the event bound to a video stream on the front end. The event here is the event that the user wants to trigger in the video stream. All events will save information such as users, cameras, video streams, and algorithms in the user event manager. After being parsed by the task parser, the user event will become several individual tasks, which will be bound to the video stream and written into the task definer.

[0133] 2.2 Task Definer

[0134] The task definer uses a directed acyclic graph to associate tasks, and there are peer and sequential relationships between tasks. In this system, all the smallest processing units are abstracted as tasks, and each task has a task list, which contains three contents: data references that need to be processed by the algorithm, the current algorithm, the input requirements of the algorithm, and the output processing of the algorithm. Video processing algorithms or image algorithms almost all require that the input data is images, so non-images or non-video frame sequences are not considered.

[0135] Most of the algorithms bound to a video stream can be calculated at the same time. This algorithm is relatively simple to process. As long as the computing resources allow, it can be executed concurrently or in parallel, and the order of results will not be affected. However, some algorithm processing requires some pre-algorithms for some reasons. There are many reasons, such as the order of events set by the user; in addition, some algorithms are time-consuming, or the video changes are relatively small. In these cases, processing every frame is undoubtedly an infeasible or unnecessary calculation. Therefore, some focusing algorithms are used to perceive the content changes of the original video frame, and then the video frame is scaled or segmented, and finally the image address that meets the algorithm input requirements is passed to the algorithm. This processing will save a lot of computing resources and can further ensure the real-time performance of the video.

[0136] The task definer is also used to deduce the algorithm's pre-algorithm and the processing method of the output result based on the algorithm's input and output. For example, if the algorithm requires a grayscale image as input, the task definer will detect whether the current image is a grayscale image. If not, it will automatically call the image's pixel conversion algorithm to convert the color image into a grayscale image.

[0137] 2.3 Task Factory

[0138] The task factory is used to generate the corresponding algorithm object or load the corresponding algorithm model according to the algorithm signature and algorithm path of the algorithm in the task.

[0139] In this embodiment, the task factory mainly produces the corresponding algorithm object or loads the model of the corresponding algorithm according to the algorithm in the task. Through the factory mode, not only can the algorithm object required for the creation task of the creation logic not be exposed, but also the algorithm object can be managed. Loading the algorithm model takes a certain amount of time. If a new model is loaded and a new object is created every time the algorithm is used, a lot of unnecessary time will inevitably be wasted. Therefore, it is worthwhile to spend a little memory to cache some algorithm models that will be used.

[0140] 3. Data management module (blackboard)

[0141] In the specific implementation process, the data management module divides the queues according to the unique identifier of the video stream, and allocates a queue to each video stream in the source stream data, intermediate data and final data stages.

[0142] In order to increase the dimension of real-time video analysis and the linkage of multiple videos, the FRTA architecture manages all data in a centralized manner. The most important data is the video stream, in addition to the text data of the algorithm's video processing results. The architecture provides a unified base class for all data to facilitate data transmission and expansion.

[0143] After processing a picture or a video frame sequence, the algorithm needs to cache the processing results for use by subsequent algorithms. Multiple algorithms process a piece of data at the same time, and their completion time must be different. Therefore, it is also necessary to cache the data generated by the first completed task in order to release resources for processing other tasks. Therefore, the blackboard mode is used to cache data in the framework of the present invention.

[0144] The blackboard model has four basic roles: data, data producer, data consumer, and controller. The data is the video stream information in this architecture, the frame sequence corresponding to the video stream, the results generated by the algorithm, and the final processed video frame sequence; the roles of data producers and data consumers are interchangeable, and a producer may also use the data generated by another producer, thus becoming a consumer. The same consumer generates data and puts it into the blackboard to become a producer. Therefore, all producers and consumers are unified into tasks in the architecture; the controller controls the interaction between data and algorithms, and this part of the function is merged into the scheduler in the architecture of this article.

[0145] In addition, the data cache form of the blackboard can be selected by the architecture implementer, who can choose to implement the data cache by himself, or use a database or message queue framework instead. In order to facilitate management and avoid overly complex architecture, the blackboard architecture divides queues according to the unique identifier of the flow, and allocates a queue for each flow in the source flow, intermediate and final data stages.

[0146] 3.1 Source Stream Data Queue

[0147] Each video stream pulled is stored through the source stream data queue, and a unique identifier is assigned to each video stream.

[0148] In this embodiment, all cameras connected to the system are managed in a unified manner. In this system, each camera will only pull one stream. Although current high-definition cameras can support the simultaneous pulling of dozens of streams, pulling too many video streams at the same time will pose a greater challenge to network bandwidth and computer memory. At the same time, repeated processing operations will be generated when performing algorithm analysis, which will also cause a waste of computing resources. Therefore, this system adopts the solution of pulling only one video stream from one camera.

[0149] The source stream data queue manages all video streams in a unified manner, and maintains a piece of meta-information (description information) for each stream, which is similar to the file control block in file management. The description information includes the source of the video stream (camera-related information and the unit to which the camera belongs, etc.), the video encapsulation format, encoder, decoder, frame resolution, bit rate, frame rate, stream pull address, stream push address, stream puller and stream pusher, etc. At the same time, in order to ensure the communication pressure between the various modules of the system, a unique identifier is assigned to each stream.

[0150] Each video stream caches a frame sequence in the system, which is the video frame sequence pulled from the camera. Each frame of the video will carry a lot of information when it is captured from the camera, but only some information related to the picture is commonly used in frame sequence or picture processing, such as picture width and height, picture pixel format, picture step, picture depth, etc. When copying or transmitting pictures, this system only transmits this necessary information to reduce bandwidth costs, and at the same time assigns a unique identifier in the stream to each frame of the picture as the position of the video stream to which the picture belongs.

[0151] It is impossible to cache all the video frames of a stream in the memory. On the one hand, the previous frames have been processed and there is no need to cache them. On the other hand, the stream puller is still grabbing video frames from the camera. If these video frames are not cached in time, it may block the subsequent video frame capture and affect the processing flow of the video stream, which will cause the real-time video playback to appear stuck, frame loss, and even cause memory overflow errors. Therefore, this system will save the processed video frames in time.

[0152] 3.2 Intermediate Data Queue

[0153] A result set is created for the video frame that triggers the user event through the intermediate data queue, and the algorithm execution results of all tasks of the video frame are stored in the corresponding result set.

[0154] In this embodiment, the intermediate data is the result generated by the algorithm in addition to the initial video stream and the final data. The reason for separating the results generated by the algorithm is also based on the principle of single responsibility, placing data from different periods in different places for easy management and maintenance.

[0155] When there is a video frame in the source stream queue, the system dequeues the video frame and puts it into the intermediate data queue to wait until all tasks on the frame are completed. If the scheduler detects that the task on the frame is not completed, it will continue to wait for the task to be executed or hand over the unexecuted task to the algorithm executor. The scheduler will write the results of the completed task into the collection associated with the frame; when all tasks on the frame are completed or timed out, the frame will be dequeued from the intermediate data queue and placed in the final data queue.

[0156] 3.3 Final Data Queue

[0157] The final data queue stores all the video frames for which tasks have been completed and their user event processing results, and each video frame is arranged in the order of the original video frame sequence.

[0158] In this embodiment, the final data is composed of the original video frame sequence and the corresponding text data, which indicates that the video frame has been processed and all the data needs to be arranged in the time sequence of the original video frame.

[0159] Generally, the frame rate of the streaming video is 25 frames per second, and the interval between one frame is 40 milliseconds. Taking into account the time consumption of unprocessed streams, an initial waiting time of 30 milliseconds is set in the architecture of this article, that is, the maximum waiting time for the stream pusher or data storage is 30 milliseconds. If the minimum video frame in the current queue is not the next frame of the pushed video frame, this frame is pushed directly without waiting.

[0160] The architecture provides a configurable waiting time. Users can manually or the system can automatically adjust the appropriate waiting time. The adjustment strategy is to monitor the task completion time, task generation and scheduling time on the video frame. Generally, when pushing the real-time stream to users for viewing, adjustment is required to ensure the smoothness of the video. Therefore, the total processing time of a frame of image should not exceed 40 milliseconds or the frame interval time. There are also some strategies in scheduling to ensure the smoothness of the video.

[0161] The present invention stores video streams and video frame sequences before, during and after processing respectively through source stream data queues, intermediate data queues and final data queues, so that all data in the video intelligent analysis process can be centrally managed, and the intermediate data generated by the algorithm can be effectively reused, thereby improving data management effects and making full use of system computing resources to reduce equipment overhead.

[0162] 4. Scheduler

[0163] During the specific implementation process, the scheduler puts the task processing results output by the algorithm executor into the intermediate data queue of the data management module: if the current task is the only task or the last task of the current video frame, the task processing results of the current video frame and all its tasks are placed in the final data queue; otherwise, check whether there is a corresponding video frame in the intermediate data queue. If not, put the corresponding video frame into the intermediate data queue and put the corresponding task processing results into the result set of the video frame; at the same time, when it is detected that all tasks of the corresponding video frame are executed or the set task processing threshold is reached, the video frame and its user event processing results are placed in the final data queue.

[0164] The scheduler is the intermediate scheduler for all task execution and data flow. According to the principle of single function, the executable tasks and data flow are extracted separately to control whether each task in the entire architecture is executed, the execution time, the execution duration, and the execution exception handling. The scheduler obtains the executable task from the task definer, determines the executability of the task, and if it cannot be executed, determines the reason for the non-executability, and performs different treatments according to the reasons; if it is executable, the algorithm executor is called. If the algorithm executor responds that it is executable, it is handed over to the algorithm executor to execute the task, otherwise it is further processed. The scheduler will put the results generated by the algorithm into the intermediate data queue of the blackboard. If the current algorithm is the only task or the last task of the current video frame, the algorithm result and the current video frame will be put into the final data queue; if not, it will detect whether there is a corresponding video frame in the blackboard. If not, the video frame will be put into the queue, and then the output result of the algorithm will be put into the result set of the corresponding video frame. At the same time, it will detect whether all the tasks of the video frame have been executed or if there is a task timeout. If so, the frame will be put into the final data queue, such as Figure 5 shown.

[0165] Another job of the scheduler is to manage the entire system. When a stream stops analyzing, or there is a problem with pulling the stream, or there is an exception in pushing the stream, or there is an exception in saving the video, the scheduler will send corresponding instructions to each module, close some unnecessary algorithm objects, clear the corresponding data memory, and notify the system of processing results and other operations.

[0166] 5. Algorithm Executor

[0167] In this embodiment, the algorithm executor is extracted separately in order to manage and allocate all computing resources in a unified manner. This design not only supports single-machine resource management, but also supports distributed or cloud computing. In other words, the architecture does not need to care who is executing the task. The scheduler only needs to hand over the task to the algorithm executor and wait for an execution result.

[0168] All tasks from the scheduler are regarded as a basic task execution unit. The algorithm executor obtains the algorithm and data in the task according to the task list. The execution system is responsible for calling and executing the algorithm and returning the result to the algorithm executor. The algorithm executor processes the result according to the result processing method in the task list and returns the algorithm result to the scheduler.

[0169] During the specific implementation process, the task list also includes algorithm input requirements and algorithm output results;

[0170] Before the scheduler assigns a task, it determines whether the task can be executed based on the task list of the task.

[0171] In this embodiment, the task includes the definition of the algorithm input, such as the color space, width, height, bit depth, etc. of the input image. If the input does not meet the input requirements of the algorithm, the task is not executable. Usually, the input of a task is the output of the previous task. The algorithm of the previous task does not guarantee that its output meets the input requirements of the next algorithm, so the algorithm is not executable.

[0172] In the present invention, before assigning a task, the scheduler determines whether the task is executable according to the task list, and then can screen out executable tasks for assignment, thereby improving the efficiency of user event processing and task execution.

[0173] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit the technical solution. Those skilled in the art should understand that those modifications or equivalent substitutions of the technical solution of the present invention that do not depart from the purpose and scope of the technical solution should be included in the scope of the claims of the present invention.

Claims

1. Video intelligent analysis method based on flexible analysis framework, It is characterized in that include: S1: Define user events in advance according to scenario requirements, and determine the tasks and task lists that need to be performed when processing each user event; The task list includes the algorithm information and data required by the task; the task list also includes the algorithm input requirements and algorithm output results; S2: Pull the video stream to be processed and extract the video frame sequence from the video stream; S3: Perform user event detection on the video frame sequence. When a video frame triggers a user event, the following user event processing flow is executed: S301: Parsing the triggered user event into a number of tasks to be executed; S302: Allocate the task to be executed to the corresponding algorithm executor through the scheduler; before the scheduler assigns the task, it determines whether the task can be executed according to the task list of the task; S303: The algorithm executor calls the corresponding algorithm and data according to the task list of the corresponding task to execute the task, obtains the corresponding algorithm execution result and stores it; If the algorithm execution result of a task corresponding to the video frame triggers a new user event, jump to step S3 and add a parallel user event processing flow for the corresponding video frame; S304: Determine through the scheduler whether all tasks corresponding to the video frame have been executed: if so, output the algorithm execution results of all tasks corresponding to the video frame as its user event processing result; Otherwise, return to step S302; Each video stream pulled is stored in the source stream data queue, and a unique identifier is assigned to each video stream; Create a result set for the video frame that triggers the user event through the intermediate data queue, and store the algorithm execution results of all tasks of the video frame in the corresponding result set; The final data queue stores all the video frames of the completed tasks and their user event processing results, and each video frame is arranged in the order of the original video frame sequence; S4: The user event processing results of all video frames in the video frame sequence are used as video intelligent analysis results.

2. The video intelligent analysis method based on the flexible analysis framework as claimed in claim 1, Features: In step S302, if there are several tasks to be processed in sequence, the scheduler assigns each task in sequence; If there are several tasks to be processed in parallel, the scheduler will assign each task at the same time.

3. The video intelligent analysis method based on the flexible analysis framework as claimed in claim 1, Features: In step S303, the called data includes video stream information, a video frame sequence corresponding to the video stream, an algorithm execution result generated by a previous algorithm, and a video frame sequence that is finally processed.

4. Video intelligent analysis system based on flexible analysis framework, It is characterized in that The video intelligent analysis method based on the flexible analysis framework described in claim 1 is implemented, comprising: Algorithm Manager, which is used to store and manage all algorithms set up; An event detection module is used to extract a video frame sequence from a video stream and perform user event detection on the video frame sequence; The task definer is used to parse the triggered user event into several tasks to be executed when a certain video frame is detected to trigger a user event; The task definer predefines user events according to scenario requirements and determines the tasks and task lists that need to be performed when processing each user event; the task list contains the algorithm information and data required by the task; The scheduler is used to obtain the tasks to be executed from the task definer and assign them to the corresponding algorithm executors; An algorithm executor is used to call the corresponding algorithm and data according to the task list of the corresponding task to execute the task, obtain the corresponding algorithm execution result and store it; The data management module is used to store the pulled video stream and the algorithm execution results generated after the task is executed; Finally, it is determined whether all tasks corresponding to the video frame have been completed: if so, the algorithm execution results of all tasks corresponding to the video frame are output as the user event processing results; otherwise, the scheduler continues to obtain and assign tasks.

5. The video intelligent analysis system based on the flexible analysis framework as claimed in claim 4, Features: The algorithm manager performs unified interface management on each algorithm based on the algorithm's source, type, function, operating efficiency and conditions; each algorithm has a unique algorithm signature.

6. The video intelligent analysis system based on the flexible analysis framework as claimed in claim 4, Features: The task list also includes algorithm input requirements and algorithm output results; Before the scheduler assigns a task, it determines whether the task can be executed based on the task list of the task.

7. The video intelligent analysis system based on the flexible analysis framework as claimed in claim 4, Features: The data management module divides the queues according to the unique identifier of the video stream, and allocates a queue for each video stream in the source stream data, intermediate data and final data stages; Each video stream pulled is stored in the source stream data queue, and a unique identifier is assigned to each video stream; Create a result set for the video frame that triggers the user event through the intermediate data queue, and store the algorithm execution results of all tasks of the video frame in the corresponding result set; The final data queue stores all the video frames for which tasks have been completed and their user event processing results, and each video frame is arranged in the order of the original video frame sequence.

8. The video intelligent analysis system based on the flexible analysis framework as claimed in claim 7, Features: The scheduler places the task processing results output by the algorithm executor into the intermediate data queue of the data management module: if the current task is the only task or the last task of the current video frame, the task processing results of the current video frame and all its tasks are placed in the final data queue; otherwise, check whether there is a corresponding video frame in the intermediate data queue. If not, place the corresponding video frame in the intermediate data queue and place the corresponding task processing results in the result set of the video frame; it is also used to place the video frame and its user event processing results into the final data queue when it is detected that all tasks of the corresponding video frame have been executed or the set task processing threshold has been reached.

Citation Information

Cited By

  • Cloud-side collaborative video content intelligent analysis and understanding system

    CN121459250A