A configurable video analytics method based on modular algorithm units
By using a modular algorithm unit library and a graphical interface configuration, flexible configuration and resource optimization management of video analysis tasks are achieved, solving the problem of traditional systems being unable to quickly expand new functions and improving the system's flexibility and resource utilization efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGDONG AVCIT TECH HLDG CO LTD
- Filing Date
- 2026-03-23
- Publication Date
- 2026-05-29
Smart Images

Figure CN122120416A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence and video surveillance system technology, and in particular to a configurable video analysis method based on modular algorithm units. Background Technology
[0002] Currently, with the rapid development of intelligent video surveillance, the Internet of Things (IoT), and artificial intelligence (AI) technologies, deep learning-based video analytics is widely used in many fields such as security, intelligent transportation, smart cities, access control, and retail analytics. Traditional intelligent video processing systems typically tightly couple cameras with pre-defined algorithm tasks: each camera is bound to one or a few dedicated algorithm programs or processes, and these tasks are permanently embedded in the device or server upon deployment. This traditional approach requires modifying or redeploying the program on the device or server to add or modify algorithm tasks, hindering the rapid expansion of new functionalities and resulting in long deployment cycles and high maintenance costs. Summary of the Invention
[0003] To address at least one of the aforementioned technical problems, this disclosure proposes a configurable video analysis method based on modular algorithm units in its first aspect. The method includes: a pre-built modular algorithm unit library storing multiple independent algorithm module definitions, each including a unique module identifier, input / output data specifications, and computational resource requirements; responding to a user's new application command on a graphical interface, displaying an algorithm selection list, receiving at least one target algorithm module selected by the user from the list, and receiving a runtime strategy configuration for the target algorithm module, including a deployment time interval setting; displaying a camera resource list, receiving a user's batch selection operation for multiple target cameras, generating an association mapping table between cameras and target algorithm modules, and instantiating the target algorithm module into an independent application runtime process on a computing node according to computational resource requirements; configuring data routing based on the association mapping table to input the video streams from the selected target cameras to the independent application runtime process in real time, and triggering the process to execute a video analysis task; and generating and displaying a visual application card for each independent application runtime process in the application management main interface, each visual application card serving as an independent interactive entry point for the application runtime process, and displaying the latest detected event images of the application runtime process in real time.
[0004] Preferably, the operation strategy configuration also includes: selecting an application cover image, receiving an icon selected by the user from a preset icon library or a custom image uploaded, and setting it as the display cover of the visual application card; setting an application name, receiving a text string entered by the user, and setting it as the title of the visual application card; and setting an alert threshold, receiving a sensitivity parameter or confidence threshold set by the user, used to control the alert triggering conditions of independent application runtime processes.
[0005] Preferably, the deployment time interval setting specifically includes: providing a time axis configuration component based on weeks, the component containing seven time dimension entries from Monday to Sunday; for each day, providing an editable time period slider or input box to set one or more discontinuous deployment time intervals; and configuring an independent application runtime process to perform analysis processing and warnings on the input video stream only when the current system time falls within the deployment time interval, and to pause analysis or ignore warning results when the current system time does not fall within the deployment time interval.
[0006] Preferably, the algorithm selection list is displayed in a hierarchical classification structure, including: the first level is the business scenario classification, including: personnel management, vehicle management, environmental monitoring, industry customization, and visual large model algorithms; the second level is the specific algorithm function, belonging to the corresponding business scenario classification, including: face recognition, uniform recognition, loitering detection, area intrusion detection, crowd gathering detection, fall detection, open flame detection, mobile phone use detection, motor vehicle intrusion detection, fire lane blockage detection, non-motorized vehicle capture, license plate recognition, off-duty detection, grass trampling detection, electric bicycle helmetless monitoring, elderly scavenger monitoring, and sleeping on duty detection.
[0007] Preferably, the system displays a list of camera resources and accepts batch selection operations from users for multiple target cameras. This includes: providing a list interface with multiple dimensions, including: channel number, video IP address, video name, and video protocol type; video names include any one or more of the following: server room, A-block elevator, basement parking garage exit, 1st floor lobby reception, freight elevator, takeout locker, exhibition hall entrance, parking garage exit, building entrance, and basement express delivery storage area; and providing a checkbox mechanism that, in response to a user selecting N different cameras at once, maps the video streams of these N cameras to the same independent application runtime process, where N is not less than 1.
[0008] Preferably, the visual application card performs interactive event display, including: displaying a numerical badge of the total number of detected events or the number of unprocessed events in real time overlaid in a preset area of the visual application card; displaying statistical results of detected events in an adjacent area of the card or in a pop-up layer in response to a user's hover operation on the cover area or name area of the visual application card; displaying thumbnails of detected events in response to a user's click operation on the visual application card; and displaying a details page of the detected event in response to a user's click operation on any thumbnail, the details page displaying the time of occurrence, location of occurrence, event type, and complete warning video clip.
[0009] Preferably, the display mode of the visual application card is dynamically switched according to the deployment status of the application runtime process, including: when the application runtime process is in an undeployed state, the visual application card displays a preset image and an undeployed indicator; when the application runtime process is in a deployed state, the step of displaying the latest detected event image in real time is executed.
[0010] Preferably, the method further includes setting a global search component in the main application management interface. The global search component includes a text input area and an image upload trigger control. In response to natural language keywords or phrases entered by the user in the text input area, or in response to a target image uploaded by the user through the image upload trigger control, a corresponding search request is generated. Based on the search request, a multimodal matching search is performed in the historical event database, and a list of historical detection event images matching the search request and their associated attribute information are displayed.
[0011] Preferably, it also includes updating the runtime strategy configuration of the independent application runtime process or the associated camera in response to the user's editing command for the visual application card.
[0012] In a second aspect, this disclosure proposes a configurable video analysis system based on modular algorithm units, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the method as described above.
[0013] Some technical advantages of this disclosure are as follows: This invention provides a configurable video analysis method based on modular algorithm units, including: a pre-built modular algorithm unit library; responding to a user's new application command on a graphical interface, receiving at least one target algorithm module and its running strategy configuration selected by the user from a list; displaying a list of camera resources, receiving the user's batch selection operation for multiple target cameras, generating an association mapping table between cameras and target algorithm modules, and instantiating the target algorithm modules into independent application runtime processes on computing nodes according to computing resource requirements; configuring data routing according to the association mapping table to input the video streams of the selected target cameras to the independent application runtime processes in real time, and triggering the processes to execute video analysis tasks; and generating and displaying a visual application card for each independent application runtime process in the application management main interface, displaying the latest detected event images of the process in real time.
[0014] This method breaks down video analysis tasks into independent, configurable, and reusable small algorithm modules. These modules can be combined into complete tasks or applications through visual or declarative configuration. The generated applications can be batch-bound or dynamically bound to multiple cameras. Furthermore, it can perform runtime orchestration, resource scheduling, version management, and monitoring, thereby enabling simple and rapid processing of new tasks. Attached Figure Description
[0015] To better understand the technical solutions of this disclosure, the following accompanying drawings, which are used to assist in the illustration of the prior art or embodiments, can be referred to. These drawings selectively illustrate the products or methods involved in the prior art or some embodiments of this disclosure. The basic information of these drawings is as follows: Figure 1 This is a flowchart of an embodiment of a configurable video analysis method based on modular algorithm units according to this application. Detailed Implementation
[0016] The following will further describe the technical means or effects involved in this disclosure. Obviously, the provided embodiments (or implementation methods) are only some of the implementation methods covered by this disclosure, and not all of them. Based on the embodiments in this disclosure and the explicit or implicit descriptions in the figures and text, all other embodiments that can be obtained by those skilled in the art without creative effort will be within the scope of protection claimed in this disclosure.
[0017] With the popularization and development of intelligent security systems, current security monitoring systems deploy multiple cameras in many areas. These cameras perform video analysis tasks, such as intrusion detection, object tracking, and facial recognition, using specialized algorithms. In traditional security systems, each camera is typically associated with a fixed task algorithm, making it difficult to quickly expand or modify the security system to meet different needs.
[0018] This disclosure provides a configurable video analytics method based on modular algorithm units. By decomposing tasks into multiple independently configurable algorithm modules, users can intuitively configure tasks through a graphical interface without writing complex code. This improves the system's flexibility, scalability, and resource optimization management.
[0019] like Figure 1 This disclosure provides a configurable video analysis method based on modular algorithm units, comprising: Step S1: Pre-set modular algorithm unit library, which stores multiple independent algorithm module definitions. Each definition includes a unique module identifier, input and output data specifications, and computing resource requirements. Step S2: In response to the user's command to create a new application on the graphical interface, display the algorithm selection list, receive at least one target algorithm module selected by the user from the list, and receive the running strategy configuration for the target algorithm module, including the deployment time interval setting. Step S3: Display the camera resource list, receive the user's batch selection operation for multiple target cameras, generate the association mapping table between cameras and target algorithm modules, and instantiate the target algorithm module as an independent application runtime process on the computing node according to the computing resource requirements. Step S4: Based on the association mapping table, configure data routing to input the video stream from the selected target camera to an independent application runtime process in real time, and trigger the process to execute video analysis tasks; Step S5: In the application management main interface, generate and display a visual application card for each independent application runtime process. Each visual application card serves as an independent interactive entry point for the application runtime process and displays the latest detected event images of the application runtime process in real time.
[0020] Step S1: Pre-set modular algorithm unit library, which stores multiple independent algorithm module definitions. Each definition includes a unique module identifier, input and output data specifications, and computing resource requirements. A pre-built modular algorithm unit library stores multiple algorithms and their corresponding independent algorithm module definitions. Each algorithm, such as face recognition, vehicle detection, and smoke detection, is independently encapsulated and has a standard interface, such as inputting video streams and outputting structured data. The algorithms do not interfere with each other and can be loaded, unloaded, and upgraded independently without affecting other algorithms or the main system program. Each algorithm unit communicates with the main system through a standard interface, enabling hot-swapping. Even if a new algorithm, such as 'helmet detection,' needs to be added, it can simply be placed in the library without requiring system reconstruction.
[0021] The system pre-builds and maintains a modular algorithm unit library. This library stores multiple independent algorithm module definitions. Each algorithm module definition is a standardized metadata description file, meaning that each algorithm module definition includes: a unique module identifier, used to globally and uniquely identify the algorithm type; input and output data specifications, explicitly defining the video stream format required by the algorithm, such as resolution, encoding format, frame rate, and output data structure, such as coordinate boxes, category labels, and confidence scores; computational resource requirements, which quantitatively describe the CPU cores, memory size, GPU memory, and other resource quotas required to run the algorithm, providing a basis for subsequent resource scheduling; and a default configuration template, containing the recommended initial threshold, sensitivity, and other parameters for the algorithm.
[0022] In one embodiment, the pre-built modular algorithm unit library is standardized. Each algorithm module definition in the library is stored in a structured data format such as JSON or YAML. Each algorithm module definition includes a globally unique identifier, such as ALG_FIRE_001 representing the open flame detection algorithm and ALG_HELMET_002 representing the safety helmet detection algorithm. The input specification defined by each algorithm module defines the video stream parameters that the algorithm can accept, such as codec: H.264 / H.265, resolution: 1920x1080, fps: 25. If the input video stream parameters do not conform to the above specifications, the front-end gateway will automatically transcode or refuse access. The output specification defined by each algorithm module defines the format of the algorithm inference result, such as {"type": "fire", "bbox": [x1, y1, x2, y2], "confidence":0.95, "timestamp": 1678888888}. The resource quota for computational resource requirements specifies the minimum resources needed to run the algorithm, such as CPU: 2 cores, memory: 4GB, GPU memory: 2GB. This data will be directly used for resource scheduling decisions in step S3. The default parameter set contains the algorithm's recommended configuration at the factory, such as a confidence threshold of 0.6 and a detection interval of 500ms. Business category tags are used to build a hierarchical list in step S2, such as "Environmental Monitoring" and "Personnel Management".
[0023] The aforementioned modular algorithm unit library can be dynamically expanded to accommodate new user-submitted requirements. When a new algorithm model, such as the latest "electric vehicle entry detection" model, is developed, developers only need to write a new module definition file according to the above standards and upload it to the modular algorithm unit library. The system can then recognize and display the new algorithm the next time a user creates a new application without needing to restart. This mechanism enables "hot-swappable" algorithm capabilities.
[0024] Step S2: In response to the user's command to create a new application on the graphical interface, display the algorithm selection list, receive at least one target algorithm module selected by the user from the list, and receive the running strategy configuration for the target algorithm module, including the deployment time interval setting. In response to a user's "Create Application" command on the graphical interface, the system displays a list of algorithm selections. When a user clicks the "Create Application" button on the graphical interface (which can be the system's main application management interface), the algorithm selection list is displayed on a screen. This list uses a hierarchical classification structure, such as categorized by business scenario: personnel, vehicles, environment, etc., facilitating quick user retrieval. Users can select one or more target algorithm modules from the list, and then configure detailed execution strategies for each selected module. Configuration includes: setting deployment time intervals; for example, the graphical interface provides a weekly timeline component, allowing users to set one or more non-contiguous deployment time periods for each day of the week, such as enabling detection only from 09:00-12:00 and 14:00-18:00 from Monday to Friday. Independent application runtime processes are configured to perform analysis and alerts on the input video stream only when the current system time falls within the deployment time interval; otherwise, analysis is paused or alert results are ignored.
[0025] Users can also select the application's cover and name on the graphical user interface. Users can choose a preset icon or upload a custom image as the application's cover and enter the application's name, making it easy to intuitively identify each created application in the management interface, i.e., the main interface. Multiple applications can be created.
[0026] Users can also fine-tune the algorithm's sensitivity or confidence threshold based on actual conditions such as ambient light and angle, facilitating a balance between false alarm and false alarm rates. In other words, the alert threshold can be set or adjusted. After setting these parameters, the user clicks "Next" on the graphical interface, which then redirects to another graphical interface including selectable cameras. The display device then shows a list of camera resources, containing detailed information such as channel number, IP address, and video name (e.g., "Lobby 1st Floor," "Basement Parking Garage").
[0027] This step involves user interaction with the system through a visual graphical interface, allowing ordinary users to define applications through visual graphics. Each application also has a unique identifier ID. The algorithm selection list above is displayed using a hierarchical classification structure, including: the first level is the business scenario classification, including: personnel management, vehicle management, environmental monitoring, industry customization, and visual large model algorithms; the second level is the specific algorithm function, belonging to the corresponding business scenario classification, including: face recognition, uniform recognition, loitering detection, area intrusion detection, crowd gathering detection, fall detection, open flame detection, mobile phone use detection, motor vehicle intrusion, fire lane obstruction detection, non-motorized vehicle capture, license plate recognition, off-duty detection, grass trampling detection, electric bicycle helmetless detection, elderly scavenger monitoring, and sleeping on duty detection.
[0028] In one embodiment, the user clicks the "Create Application" button on the system's main interface, triggering a new application creation command. In response to this command, the system displays an application wizard window, with an algorithm selection list on the left side. This list is not arranged in a flat format, but rather uses a hierarchical classification structure. The first level is business scenarios, including categories such as "Personnel Management," "Vehicle Management," "Environmental Monitoring," "Industry Customization," and "Visual Large Model Algorithms." When a user clicks on "Environmental Monitoring," the list expands.
[0029] The second level is the specific functions: for example, under "Environmental Monitoring", specific algorithms such as "Open Flame Detection", "Smoke Detection", and "Fire Exit Blockage" are displayed.
[0030] Users can select one of the target algorithms, such as "open flame detection," to create an application. Users can also select two or more target algorithms, such as "open flame detection" and "smoke detection," to create a composite application. Each application generated above, whether composite or not, has a unique identifier ID.
[0031] After selecting one or more algorithms, the deep configuration of the running strategy is completed. The configuration panel is displayed on the right side of the interface. Users need to complete at least the following key configurations: First, the identity customization of the newly created application, including selecting the application's cover image. The system provides an icon library, such as a flame icon, a safety helmet icon, etc. Users can also upload local images, such as photos of the project site, as the cover image of the newly created application. This image will be used as the cover image of the visual application card in step S5. Second, setting warning thresholds, receiving user-defined sensitivity parameters or confidence thresholds, is used to control the warning triggering conditions of independent application runtime processes.
[0032] The in-depth configuration of the above-mentioned operating strategy also includes name setting, where the user enters "Open Flame Monitoring in Area A of Chemical Plant" as the title of the newly created visualization application card.
[0033] It also includes setting deployment time intervals, with a visual timeline configuration component on the system interface, organized by week. The component horizontally displays seven entries from Monday to Sunday. Clicking "Monday" brings up a 24-hour timeline. Users can drag and drop color blocks on the timeline to represent deployment periods. Deployment periods can also be discontinuous, allowing users to draw multiple color blocks. For example, drawing two time periods: 08:00-12:00 and 14:00-18:00, with 12:00-14:00 designated as a lunch break. Users can also make differentiated settings, setting completely different deployment strategies for Saturdays and Sundays, such as disarming all day on weekends. Users can also fine-tune warning thresholds through the interface, with a slider control allowing them to adjust the "sensitivity" or "confidence threshold." For example, the confidence level for open flame detection can be increased from the default 0.6 to 0.8 to reduce false alarms caused by sunlight reflection.
[0034] The system converts these deep configurations of the running strategies into cron expressions or time interval objects, stores them in the database, and allows runtime processes to query them.
[0035] In one embodiment, receiving at least one target algorithm module selected by the user from a list includes: receiving at least two target algorithm modules selected by the user; receiving data flow logic defined by the user for the at least two target algorithm modules; wherein the data flow logic is used to specify the execution order and data interaction method between the target algorithm modules, including at least one of the following: serial cascading logic: using the detection area or feature data output by the preceding target algorithm module as the input constraint of the subsequent target algorithm module, and triggering the subsequent module to perform analysis only when the preset conditions of the preceding module are met; parallel collaborative logic: concurrently inputting the same video stream to multiple target algorithm modules, and performing weighted fusion or logical AND / OR operations on the output results of each module to generate a comprehensive judgment result; conditional branching logic: dynamically routing to different subsequent target algorithm modules for processing according to the recognition category or confidence level of the preceding target algorithm module. Based on the data flow logic defined above, the system constructs a corresponding memory data pipeline or call graph within the generated independent application runtime process to ensure that each algorithm module works efficiently and collaboratively according to the logic specified by the user.
[0036] In this embodiment, when the user selects at least two target algorithm modules (e.g., "Module A: Face Detection" and "Module B: Mask Wearing Detection"), the system not only loads these two modules but also receives user-defined data flow logic to construct complex business scenarios. Cascade Scenario: Configuration: The user sets module A as the preceding node and module B as the succeeding node. Logic: The system first runs module A to scan the entire image. Only when module A detects a face and outputs a bounding box (ROI) with a confidence level higher than a threshold is the image region cropped from that bounding box passed to module B for mask detection. Effect: If module A does not detect a face, module B does not start at all. This "pruning" mechanism significantly reduces computational power consumption and avoids ineffective mask detection algorithms in unoccupied environments.
[0037] Parallel Scenario: Configuration: The user configures module A (smoke detection) and module C (channel congestion detection) to run in parallel. Logic: The same video frame is simultaneously fed into the input ports of both modules. The system sets up a fusion gateway at the output end. A red alert (the highest level) is triggered only when both module A and module C report "channel congestion" (logical AND operation); if only one occurs, a regular yellow alert is triggered. Effect: Through multi-dimensional feature fusion, the false alarm rate caused by a single feature is significantly reduced.
[0038] Conditional Branching Scenario: Configuration: The preceding module is "Vehicle Type Recognition," and the subsequent branches are "Hazardous Goods Vehicle Detection" and "Ordinary Passenger Vehicle Speeding Detection." Logic: The system dynamically determines the data flow based on the output category (Label) of the preceding module. If identified as a "Hazardous Goods Vehicle," the data is routed to the dedicated detection module; if identified as a "Ordinary Passenger Vehicle," the data is routed to the speeding detection module. Effect: This implements an adaptive analysis strategy based on scenario content, improving the system's flexibility and targeting.
[0039] Step S3: Display the camera resource list, receive the user's batch selection operation for multiple target cameras, generate the association mapping table between cameras and target algorithm modules, and instantiate the target algorithm module as an independent application runtime process on the computing node according to the computing resource requirements. Users can select N (N≥1) different cameras at once using a checkbox mechanism. They can also select multiple regions or one or more cameras within a single region based on the camera's location. Upon receiving the user's camera selection, the system automatically generates a mapping table between the selected cameras and the target algorithm module, establishing a "many-to-one" or "one-to-many" analysis relationship.
[0040] Display a list of camera resources and accept batch selection operations from users for multiple target cameras. This includes: providing a list interface with multiple dimensions, such as: channel number, video IP address, video name, and video protocol type; video names include any one or more of the following: server room, A-block elevator, basement parking garage exit, 1st floor lobby reception, freight elevator, takeout locker, exhibition hall entrance, parking garage exit, building entrance, and basement express delivery storage area; and providing a checkbox mechanism that, in response to a user selecting N different cameras at once, maps the video streams of these N cameras to the same independent application runtime process, where N is not less than 1.
[0041] Based on the computational resource requirements defined in the algorithm module, the system dynamically pulls the corresponding container image on a compute node, such as a Kubernetes cluster or Docker host, starts a new container instance, and instantiates the selected target algorithm module as an independent application runtime process. This process has its own independent namespace, file system view, and network stack, and is strictly isolated from other processes.
[0042] When multiple target algorithm modules are selected, in response to the user's combination command, multiple target algorithm modules are selected from the algorithm unit library, and the data flow logic between each target algorithm module is defined, such as serial cascading, parallel collaboration, or conditional branching. The user's configuration of the running parameters of the target algorithm modules is received, including detection thresholds, areas of interest, alarm strategies, etc. Based on the data flow logic and running parameters, the selected multiple target algorithm modules are dynamically assembled and instantiated into an independent custom application process to generate the specific intelligent analysis application required by the user.
[0043] In one embodiment, the system reads all registered cameras from the database and displays them in a list. List fields include: Channel Number (internal channel index of the device); Video IP Address (network address of the camera); Video Name (a user-friendly location description, such as "Server Room," "Staircase A," "Basement 1 Garage Exit," "1st Floor Lobby Reception," "Freight Elevator," "Takeout Locker," "Exhibition Hall Entrance," "Garage Exit," "Building Entrance," "Basement 1 Express Delivery Storage Area," etc.); Video Protocol: RTSP / ONVIF / GB28181, etc.; Status: Online or Offline.
[0044] Each row in the list has a checkbox. Users can select cameras for N areas at once, such as selecting cameras at the three key locations: "Exhibition Hall Entrance," "Building Entrance," and "Basement 1 Package Storage Area." The system backend immediately generates a mapping table between cameras and target algorithm modules. This table records: {App_ID: "App_001", Algo: "Fire_Detect", Cameras: ["Cam_A", "Cam_B", "Cam_C"]}.
[0045] Next, on the compute node, the target algorithm module is instantiated as an independent application runtime process. The following operations are performed on the compute node: Image pulling: Pull the Docker image corresponding to the "open flame detection" module from a pre-built algorithm image repository, such as Harbor. Container startup: Invoke a container engine such as Docker Engine or Kubelet to start a new container instance. Resource limiting: Strictly limit the CPU and memory usage of this container in the startup parameters, for example, `--cpus=2 --memory=4g`, to prevent it from exhausting the host machine's resources. Configuration mounting: Mount the arming time configuration and threshold configuration generated in step S2, as well as the camera mapping table generated in step S3, to this container instance in the form of configuration files or environment variables. Process isolation: At this point, a new, independent application runtime process is created. This new application runtime process has an independent process ID, network namespace, and file system view. Even if this process crashes, it will not affect other running applications such as "face recognition" or "vehicle detection" in the system.
[0046] Step S4: Based on the association mapping table, configure data routing to input the video stream from the selected target camera to an independent application runtime process in real time, and trigger the process to execute video analysis tasks; Based on the generated association mapping table, the system configures an internal data routing gateway. The gateway inputs the video streams from one or more target cameras selected by the user into the same independent application runtime process in real time, using zero-copy or efficient memory sharing technology. This way, the same video stream only needs to be decoded once and can be consumed simultaneously by multiple algorithm instances within the same independent application runtime process, significantly reducing bandwidth and CPU consumption.
[0047] Once the data routing is established, the system automatically triggers a process to execute a video analysis task. Internally, the process determines whether it is currently in the arming period based on the arming time strategy configured in step S2. If it is in the arming period, real-time inference is performed on the video frames; otherwise, inference is paused or the results are discarded to save computing power.
[0048] In one embodiment, the intelligent video gateway within the system reads an associated mapping table. For each camera in the mapping table, such as the "exhibition hall entrance," the gateway establishes an RTSP connection to acquire the video stream. If multiple applications or multiple logical threads within the same application need to analyze the same video stream, the gateway performs decoding only once, and then distributes the decoded YUV / RGB frame data to multiple consumers simultaneously through shared memory or zero-copy technology. In this embodiment, three video streams—the exhibition hall, the building, and the courier station—are input into the same "open flame detection" container process in real time. This mechanism avoids starting a separate decoding process for each camera, greatly saving CPU resources.
[0049] The container process internally loads the deployment time configuration for real-time verification. Before processing each frame, the process first checks if the current system time falls within the configured deployment interval. If it is within the deployment period, such as Monday at 10 AM, the process calls the AI inference engine to perform a full analysis of the video frames to detect the presence of open flames. If detected, an alert event is generated. If it is outside the deployment period, such as Monday at 3 AM, the process automatically skips the inference step and directly enters sleep mode or only maintains a heartbeat connection, without generating any alerts or consuming GPU computing power. This ensures "on-demand allocation" of computing power and avoids unnecessary computation.
[0050] Step S5: In the application management main interface, generate and display a visual application card for each independent application runtime process. Each visual application card serves as an independent interactive entry point for that process and displays the latest detected event images of that process in real time.
[0051] After binding the target camera, clicking the "Next" button will take you to the next visual graphical interface, which is the main interface for application management. The system generates and displays a visual application display unit for each independent application runtime process. In one embodiment, the visual application display unit is a visual application card. Each card is the unique interactive entry point for that process, displaying the latest detected event images in real time. Multiple events are detected, and the latest detected event images are continuously scrolled. Each latest detected event image is a screenshot of a video frame at the moment the warning was triggered, providing a clear view of the situation on site. This display is constantly updated dynamically, with new events automatically displayed and old events automatically removed, maintaining the real-time nature of the preview.
[0052] Each visualization application card can display events in a multi-level interactive manner. The first level of display, namely the quantity badge, is a numerical badge that is displayed in real time in a preset corner of the visualization application card, such as the upper right corner. This number represents the total number of events detected so far, or the number of unprocessed alert events. This allows users to get a quick overview of the application's activity level.
[0053] The second level of display is the event handling preview: When a user hovers over the cover or name area of a visual application card, the system dynamically expands and displays the total number of detected events, the number of events processed, and the number of events not yet processed in an adjacent area of the card, such as the bottom of the interface or in a pop-up overlay in the interface.
[0054] The third level of display involves detailed navigation. When a user clicks on a visual application card, the system immediately redirects to thumbnails of all events corresponding to that event. Clicking on a thumbnail leads to its details page. The details page deeply integrates multi-dimensional information, including the precise time and location of the event, the name and location description of the associated camera, the event type (e.g., "open flame detection"), and complete pre- and post-warning video clips. Users can perform operations such as confirmation, order dispatch, and export on this page.
[0055] In one embodiment, on the application management main interface, the system generates a visual application display unit, i.e., a visual application card, for the previously created "Chemical Plant A Area Open Flame Monitoring" application. In the upper right corner of the card, the system overlays a prominent numerical badge in real time. This number represents the "current cumulative total number of detected events." For example, it displays the red number "10." Users can see from the main interface that 10 suspected open flame events have occurred in that area without clicking any buttons.
[0056] When a user hovers their mouse over a card, a primary interaction is triggered. An interactive box is displayed below the card, including buttons for editing the application, setting up the application, configuring the layout, and deleting the application. Alternatively, a pop-up layer smoothly expands, displaying a statistical chart of detected events, including the total number of detected events, the number processed, and the number unprocessed. When the application is under control, the card cover displays a screenshot of the video frame at the moment the latest detected alert was triggered, clearly showing the location of the fire. For example, when the process detects the 11th event, the system automatically takes a screenshot and displays it on the application card cover.
[0057] When a user clicks the application card cover or the aforementioned buttons for editing, deploying, layouting, or deleting the application, a secondary interaction is triggered. When the user clicks the application card cover, the browser or client immediately redirects to a page displaying images of all events detected by the application. Clicking on one of these event image pages then brings up its details page, displaying the time of the open flame event in a structured manner: 2026-03-20 10:15:32; the location of the open flame event: associated camera name "Exhibition Hall Entrance," and its location on a map. The event type is open flame detection with a confidence level of 0.92. Video clips are displayed, with a built-in player that automatically plays 15 seconds of video recording before and after the alert, which can be paused, dragged, and displayed in full screen. Operation buttons can also be provided, including "Confirm True," "Mark as False Alarm," "Generate Work Order," and "Export Video." The user completes the final confirmation and action on this page.
[0058] Based on the deployment status of the application runtime process, the display mode of the visual application card is dynamically switched, including: when the application runtime process is in an undeployed state, the visual application card displays a preset image and an undeployed indicator; when the application runtime process is in a deployed state, the step of displaying the latest detected event image in real time is executed.
[0059] In one embodiment, a user discovered that the "Open Flame Monitoring in Area A of the Chemical Plant" application occasionally triggered false alarms at night. The user decided to adjust its deployment time, enabling it only during the day and increasing the confidence threshold. The user locates the "Open Flame Monitoring in Area A of the Chemical Plant" visual application card on the main interface. Clicking the "Edit" button below the card enters the configuration page. First, the deployment time is modified, changing the deployment period from Monday to Sunday to 06:00-18:00. Then, the threshold is modified, increasing the confidence level from 0.6 to 0.75. "Save" is clicked. The backend API receives the update request and verifies its validity. Based on the App_ID, the system accurately locates the specific running container instance Process_A. The system generates a new configuration file, config_updated.json. The new configuration is dynamically injected into the runtime environment of Process_A through container runtime interfaces such as docker exec or K8s ConfigMap. Alternatively, a SIGUSR1 signal is sent to notify Process_A to reload the configuration. Upon receiving the new configuration, Process_A immediately updates the deployment time logic and threshold variables in memory. This process does not involve container restarts, the process PID remains unchanged, and the network connection is maintained. Meanwhile, other applications running in the system, such as the "face recognition attendance application" (Process_B) and the "vehicle parking violation application" (Process_C), are completely unaffected. Their video streams continue to be analyzed, and their configurations remain unchanged. The front-end interface receives a successful update receipt, the card status is refreshed, and the new policy takes effect immediately.
[0060] In one embodiment, a user discovered the "Open Flame Monitoring in Area A of the Chemical Plant" application, which was already linked to cameras in five areas. During actual operation, users discovered two problems: firstly, occasional false alarms were caused by light reflection at night; secondly, two new cameras were installed in different areas and needed to be included in the monitoring. The user decided to adjust its deployment schedule, enabling high-precision detection only during the day and reducing sensitivity or disabling at night, and adding the new cameras to the application.
[0061] Users can locate the "Chemical Plant A Zone Open Flame Monitoring" visualization application card on the main interface. Click the "Edit" button below the card to access the configuration page. Modify the deployment and threshold settings. On the weekly timeline component, adjust the deployment period from Monday to Sunday to 06:00-18:00 (daytime, high precision). Add a new rule for the 18:00-06:00 (nighttime) period, adjusting the confidence threshold from 0.6 to 0.85. In the camera list, additionally select all cameras in the newly installed "Chemical Plant A Zone New Unloading Platform" and "Chemical Plant A Zone Backup Storage Tank" areas. Click "Save" to save the modified application.
[0062] During this process, the backend API receives an update request and verifies the online status and configuration validity of the new cameras. Based on the App_ID, the system precisely locates the specific running container instance (Process_A) and generates a new merge configuration file, config_updated_v2.json, containing a list of cameras for the two new areas and time-segmented threshold policies. The new configuration is dynamically injected into the runtime environment of Process_A through container runtime interfaces, such as updating mounted volumes via K8s ConfigMap or writing to a file using docker exec. Simultaneously, a SIGUSR1 signal is sent to Process_A (or via the internal HTTP control port) to notify it to reload the configuration. The streaming media management thread within Process_A receives the signal, parses the new configuration, and identifies all camera IPs in the two newly added areas. Without interrupting the video stream processing threads of the original five areas, it asynchronously initiates RTSP connections and shared memory mappings for all cameras in the two new areas. Once the new streams are ready, the corresponding inference sub-thread is immediately started. The time-segmentation engine within Process_A atomically replaces the deployment schedule and threshold variables in memory. A read-write lock mechanism is used to ensure that ongoing inference frames complete using the old values, and newly arriving frames immediately use the new values. This process does not involve container restarts, the process PID remains unchanged, the network connection of the original 5 area cameras remains unbroken, and video analysis is uninterrupted.
[0063] Meanwhile, the "Face Recognition Attendance Application" (Process_B) and the "Vehicle Parking Violation Application" (Process_C) running in the system remain completely unaffected. Their video streams continue to be analyzed, their configurations remain unchanged, and users are unaware of their actions. Each application is independent and does not affect the others.
[0064] In one embodiment, a "tank area fire early warning" project in a large chemical industrial park faces a severe challenge: the area contains numerous metal pipes and reflective surfaces, which generate strong reflections and sparks when exposed to sunlight or during welding operations. If only a single "open flame detection algorithm" is deployed, the system will misinterpret reflections and welding sparks as fires, resulting in dozens of false alarms daily, causing security personnel to become complacent and fearful. A composite detection application must be built that can simultaneously identify "open flames" and confirm the presence of "smoke." Only when both conditions are met simultaneously (or occur sequentially according to a specific timeframe) will the highest level alarm be triggered, thereby completely eliminating false alarms caused by environmental interference.
[0065] The user locates the running visualization application card for "Open Flame Monitoring in Area A of the Chemical Plant" on the main interface. Clicking the edit button below the card enters the configuration page. The user selects the first target algorithm module from the library: "High-Sensitivity Open Flame Detection Module" and the second target algorithm module: "Dynamic Smoke Texture Analysis Module". On the visualization interface, the user selects "High-Sensitivity Open Flame Detection Module" as the preceding node and "Dynamic Smoke Texture Analysis Module" as the following node. Cascading logic ROI constraints are configured, with the user selecting a rule that the system only extracts the coordinate region (ROI) of the flame frame when the confidence level output by "High-Sensitivity Open Flame Detection Module" is > 0.6. These ROI coordinates are then dynamically passed to the "Dynamic Smoke Texture Analysis Module" and the "Smoke Analysis Module" as spatial constraints. The user clicks "Generate Application" and names it "Dual Confirmation Fire Monitoring in Tank Area". The system starts a unique, independent application runtime process on the edge computing node. The dynamic library files for the two algorithms (open flame and smoke) are loaded into the process's memory. A zero-copy memory channel is established within the process. After the video frame is decoded once, the open flame algorithm processes it and directly outputs the coordinate pointer to the smoke algorithm. Based on the association mapping table, the four high-definition camera video streams from the tank area are concurrently input to this process. Internally, the process uses a multi-threaded mechanism to execute the aforementioned cascaded logic on each of the four streams.
[0066] In a second aspect, this disclosure proposes a configurable video analysis system based on modular algorithm units, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the method as described above.
[0067] Within the scope of knowledge and ability of those skilled in the art, the various embodiments or technical features mentioned herein can be combined with each other as other optional embodiments without conflict. These finite number of optional embodiments, which are not listed one by one and are formed by combining a finite number of technical features, still fall within the scope of the technology disclosed herein and are also derived by those skilled in the art from the accompanying drawings and the foregoing.
[0068] In addition, the descriptions of most embodiments are based on different focuses. For further understanding of the parts not described in detail, reasonable inference can be made by referring to the relevant content of the prior art, other relevant descriptions in this document, or the inventive intent.
[0069] To reiterate, the embodiments listed above are typical and preferred embodiments of this disclosure, and are only used to describe and explain the technical solutions of this disclosure in detail to facilitate the reader's understanding. They are not intended to limit the scope or application of the protection claimed in this disclosure. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure to obtain technical solutions should be covered within the scope of protection claimed in this disclosure.
Claims
1. A configurable video analysis method based on modular algorithm units, characterized in that, include: A pre-built modular algorithm unit library is provided, which stores multiple independent algorithm module definitions. Each definition includes a unique module identifier, input and output data specifications, and computing resource requirements. In response to the user's command to create a new application on the graphical interface, the algorithm selection list is displayed, the user selects at least one target algorithm module from the list, and the running strategy configuration for the target algorithm module is received. The running strategy configuration includes the deployment time interval setting. Display a list of camera resources, receive batch selection operations from users for multiple target cameras, generate an association mapping table between cameras and target algorithm modules, and instantiate the target algorithm modules as independent application runtime processes on computing nodes according to computing resource requirements. Based on the association mapping table, configure data routing to input the video stream from the selected target camera to an independent application runtime process in real time, and trigger the process to execute video analysis tasks; In the main application management interface, a visual application card is generated and displayed for each independent application runtime process. Each visual application card serves as an independent interactive entry point for the application runtime process and displays the latest detected event images of the application runtime process in real time.
2. The method according to claim 1, characterized in that, The runtime policy configuration also includes: Select an application cover image; receive an icon selected by the user from a preset icon library or a custom image uploaded, and set it as the display cover for the visual application card. Set the application name, receive the text string input by the user, and set it as the title of the visual application card; Set early warning thresholds and receive user-defined sensitivity parameters or confidence thresholds to control the early warning triggering conditions of independent application runtime processes.
3. The method according to claim 1, characterized in that, The specific deployment time interval settings include: Provides a timeline configuration component based on weeks, which contains seven time dimension entries from Monday to Sunday; For each day, an editable time period slider or input box is provided to set one or more non-contiguous arming time intervals; The standalone application runtime process is configured to perform analysis and alerts on the input video stream only when the current system time falls within the deployment time interval, and to pause analysis or ignore alert results when the current system time does not fall within the deployment time interval.
4. The method according to claim 1, characterized in that, The algorithm selection list is displayed using a hierarchical classification structure, including: The first level is classified by business scenario, including: personnel management, vehicle management, environmental monitoring, industry customization, and visual large model algorithms; The second level consists of specific algorithm functions, which belong to the corresponding business scenario categories, including: face recognition, uniform recognition, loitering detection, area intrusion detection, crowd gathering detection, fall detection, open flame detection, mobile phone use detection, motor vehicle intrusion detection, fire lane blockage detection, non-motorized vehicle capture, license plate recognition, off-duty detection, grass trampling detection, electric bicycle helmetless detection, elderly scavenger monitoring, and sleeping on duty detection.
5. The method according to claim 1, characterized in that, Displays a list of camera resources and allows users to select multiple target cameras in batches, including: Provides a list interface with multiple dimensions, including: channel number, video IP address, video name and video protocol type; video name includes any one or more of the following: server room, A-block elevator, basement parking garage exit, 1st floor lobby reception, freight elevator, takeaway locker, exhibition hall entrance, parking garage exit, building entrance, basement express delivery storage area; Provide a checkbox mechanism that responds to a user selecting N different cameras at once, mapping the video streams of these N cameras to the same independent application runtime process, where N is not less than 1.
6. The method according to claim 1, characterized in that, The visual application cards execute interactive events, including: In the preset area of the visualization application card, a numerical superscript is overlaid to display the current cumulative total number of detected events or the number of unprocessed events; In response to a user's hovering action over the cover or name area of a visual application card, statistical results of the detected events are displayed in an adjacent area of the card or in a pop-up overlay. In response to a user's click on a visual application card, display a thumbnail of the detected event; In response to a user's click on any thumbnail, the event details page is displayed, showing the time, location, type of event, and a complete warning video clip.
7. The method according to claim 1, characterized in that, Based on the deployment status of the application's runtime processes, the display mode of the visual application card dynamically switches, including: When the application's runtime process is in an uncontrolled state, the visual application card displays a preset image and an uncontrolled indicator; When the application runtime process is in a controlled state, execute the step of displaying the latest detected event images in real time.
8. The method according to claim 1, characterized in that, It also includes setting a global search component in the main application management interface, which includes a text input area and an image upload trigger control; In response to natural language keywords or phrases entered by the user in the text input area, or in response to the target image uploaded by the user via the image upload trigger control, a corresponding search request is generated; Based on the search request, perform a multimodal matching search in the historical event database and display a list of historical detected event images that match the search request, along with their associated attribute information.
9. The method according to claim 1, characterized in that, It also includes, In response to user commands to edit visual application cards, update the runtime strategy configuration of independent application runtime processes or associated cameras.
10. A configurable video analysis system based on modular algorithm units, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor executes the program to implement the method as described in any one of claims 1 to 9.