Multi-channel intelligent mixed stream video method based on AI intelligent card port
By combining AI-powered smart checkpoints and smart NVRs, bandwidth optimization and all-weather adaptability for multi-channel video surveillance are achieved, solving the problems of high bandwidth consumption and false positives and false negatives caused by changes in lighting conditions in traditional methods, and providing an efficient, economical and robust monitoring solution.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING XISHAN EXPERIMENTAL FOREST FARM
- Filing Date
- 2026-02-04
- Publication Date
- 2026-05-05
AI Technical Summary
Existing multi-channel video surveillance methods suffer from problems such as high bandwidth consumption, exacerbated issues when multiple terminals are used concurrently, low resource utilization, difficulty in balancing static scene data retention and timeliness, and false detections and missed detections due to changes in lighting.
A multi-channel intelligent mixed-stream video method based on AI smart checkpoints is adopted. A lightweight AI model is used for dynamic detection to distinguish between valid and invalid targets. Combined with the differentiated data scheduling and light perception of the intelligent NVR, real-time video streams or timestamped screenshots are uploaded to ensure the real-time performance and effectiveness of monitoring.
Significantly reduces bandwidth consumption of multi-channel video surveillance, improves the operating efficiency and environmental adaptability of the monitoring system, ensures the accuracy and reliability of all-weather monitoring, and meets the low-consumption and high-efficiency needs of traffic management, park security, smart cities and other fields.
Smart Images

Figure CN121985147A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a multi-channel intelligent mixed-stream video method based on AI intelligent checkpoints. Background Technology
[0002] Existing multi-channel video surveillance methods are mainly based on traditional full-stream transmission from cameras, storage-and-forwarding from ordinary network video recorders (NVRs), or simple motion detection and transmission technologies. In traditional full-stream transmission schemes, the front-end cameras continuously upload complete video streams to the back-end or NVR regardless of whether there is any valid movement in the monitored scene, such as people, vehicles, or animals. This approach generates significant bandwidth consumption when multiple cameras are running simultaneously and multiple back-end terminals are viewing concurrently—a single 1080P real-time video stream requires 2-4 Mbps of bandwidth, 10 cameras transmitting concurrently require 20-40 Mbps, and 3 back-end terminals viewing simultaneously require a combined bandwidth demand of 60-120 Mbps. This easily leads to network congestion and transmission delays, and the transmission of a large number of meaningless dynamics, such as wind blowing through grass and changes in light and shadow, is a waste of resources, increasing storage and network costs.
[0003] Other methods use fixed-interval screenshots to replace part of the video transmission to reduce bandwidth consumption. However, such solutions lack front-end dynamic recognition linkage, and the fixed screenshot intervals do not embed real-time timestamps. In static scenarios, screenshot intervals that are too short still waste bandwidth, while intervals that are too long can easily miss key information. Furthermore, the back-end cannot intuitively judge the timeliness of the screenshots and needs to manually check the time, which affects monitoring efficiency. In addition, the lack of a time synchronization mechanism between the front-end and the NVR in this solution makes it easy for screenshot timestamps to deviate, further reducing the reliability of monitoring data.
[0004] The limitations of these traditional methods include: high bandwidth consumption for full transmission, exacerbating the problem when multiple terminals are concurrent; incomplete bandwidth saving; lack of flexibility and time correlation in fixed screenshot schemes, resulting in insufficient monitoring effectiveness; low resource utilization due to the lack of differentiated scheduling and single-stream multi-push capability in ordinary NVRs; difficulty in balancing static scene data retention and timeliness, affecting the efficiency of post-event traceability and real-time monitoring; failure to consider the impact of ambient light changes on background modeling, leading to false detections in strong light and low light scenarios, such as light interference being judged as moving targets, and missed detection of real targets due to insufficient light, resulting in poor robustness of all-weather monitoring. Summary of the Invention
[0005] This invention provides a multi-channel intelligent mixed-stream video method based on AI smart checkpoints, which can reduce bandwidth consumption in multi-channel video surveillance scenarios while ensuring the real-time performance and effectiveness of monitoring.
[0006] To achieve the above objectives, this invention provides a multi-channel intelligent mixed-stream video method based on AI intelligent checkpoints, which, crucially, includes the following steps:
[0007] Step 1: Construct a multi-channel intelligent mixed-stream video system based on AI smart checkpoints. This multi-channel intelligent mixed-stream video system is equipped with M AI smart checkpoint cameras. All AI smart checkpoint cameras are connected to the same intelligent NVR module. The intelligent NVR module is also connected to N video viewing terminals.
[0008] Step 2: Each AI smart checkpoint camera collects video streams of the corresponding monitoring area in real time, and then uses the built-in lightweight AI model to perform dynamic detection on the video stream. If the dynamic detection result is that there are no moving objects, it is determined to be a static scene; if the dynamic detection result is that there are moving objects, it is determined to be a dynamic scene.
[0009] Then, moving objects in the dynamic scene are classified as valid targets or invalid targets; when a moving object is a valid target, such as a person, vehicle, or animal, the AI smart checkpoint camera uploads a real-time video stream to the smart NVR module.
[0010] When the moving object is an invalid target, such as trees or grass being blown by the wind, or when there is a static scene, the AI smart checkpoint camera generates a screenshot with a real-time timestamp at a set time interval, such as 30 seconds, and uploads it to the smart NVR module.
[0011] Step 3: The intelligent NVR module classifies and stores the received real-time video streams or screenshots with real-time timestamps;
[0012] Step 4: The video viewing terminal obtains a viewing instruction, generates a viewing request for the target camera based on the viewing instruction, and then sends the viewing request to the smart NVR module;
[0013] Step 5: The intelligent NVR module queries the target camera status according to the viewing request. If the target camera status is a valid dynamic scene, the intelligent NVR module forwards the video stream to the video viewing terminal in real time.
[0014] If the scene is invalid (either dynamic or static), the smart NVR module pushes screenshots with real-time timestamps to the video viewing terminal at set time intervals.
[0015] Step 6: The video viewing terminal receives real-time video streams or screenshots with real-time timestamps pushed from the intelligent NVR module. If it is a video stream, it will be played in real time; if it is a screenshot, it will be displayed in a carousel or other format. At the same time, the real-time video stream can be manually triggered for viewing to meet different monitoring needs.
[0016] Through the above design, this invention leverages the dynamic recognition of front-end intelligent AI checkpoint cameras, the efficient data scheduling of intelligent NVRs, and the on-demand viewing mechanism in the back-end to achieve bandwidth optimization and intelligent transmission in multi-channel video surveillance scenarios. This method achieves a breakthrough in addressing the problem of high bandwidth consumption in traditional multi-channel video surveillance. By accurately distinguishing effective dynamics such as people, animals, and vehicles from ineffective dynamics such as wind blowing through vegetation through front-end AI, and by providing differentiated data pushes from intelligent NVRs—namely, real-time video streams or timestamped screenshots—it achieves end-to-end bandwidth savings from the data acquisition source to the transmission stage, while simultaneously ensuring the effectiveness of monitoring.
[0017] This invention constructs a full-link optimization system encompassing front-end intelligent sensing, mid-end differentiated scheduling, and back-end precise adaptation. It deeply integrates the dynamic recognition capabilities of front-end AI intelligent checkpoints with the differentiated scheduling functions of intelligent NVRs, achieving intelligent linkage from data acquisition to transmission. By accurately distinguishing between valid targets and redundant information, it significantly reduces bandwidth consumption in multi-channel video surveillance scenarios. Introducing a light intensity factor, it uses light sensors to collect ambient light data in real time, building dedicated adaptive background models for different lighting scenarios (strong daylight, dim nightlight, low light at dawn and dusk, etc.), effectively solving the false detection and missed detection problems caused by changes in lighting in traditional solutions. While ensuring the real-time performance and effectiveness of monitoring, it achieves stable adaptation to all-weather scenarios. This solution provides more efficient, economical, and robust technical support for fields requiring 24 / 7 uninterrupted multi-channel video surveillance, such as traffic monitoring, park security, and smart cities, combining technological innovation with practical application.
[0018] Preferably, in step 2, the lightweight AI model uses a fusion algorithm of frame difference method + background modeling to perform dynamic detection and obtain the final motion region mask value.
[0019] When both the frame difference method and the background modeling method determine that the region is a moving region, it is considered a real moving region, and the final moving region mask value is 1; otherwise, it is a static region, and the final moving region mask value is 0. The expression is:
[0020] ;
[0021] in, This is the final motion region mask value; This represents the motion region mask using the frame difference method. This indicates that the frame difference method determines the motion region. Indicates the current frame With background model The difference, The background difference threshold is set to a grayscale value between 0 and 40. The result of the background modeling method is considered to be the motion area.
[0022] The front-end module of this invention is an AI-powered intelligent checkpoint camera, which is the core data acquisition and processing unit for realizing a multi-channel intelligent mixed-stream video method. This module uses a built-in lightweight AI model to dynamically detect and classify video content in the monitored area, thereby determining whether to upload a real-time video stream or a screenshot with a real-time timestamp, thus optimizing bandwidth usage from the data source.
[0023] This lightweight AI model uses a fusion algorithm of frame difference method and background modeling to achieve accurate recognition of moving objects. Its core is to distinguish real moving targets from interference factors such as light and shadow and shaking through mathematical formula calculation and model iteration.
[0024] As a preferred embodiment, the frame difference method is used for preliminary localization of the motion region, and the specific determination process is as follows:
[0025] The lightweight AI model calculates the pixel difference between consecutive video frames using the frame difference method to quickly locate potential motion regions.
[0026] The first step is to calculate the difference between adjacent frames, that is, to calculate the difference between the current frame and the adjacent frame. With the previous frame The difference and the current frame With the next frame The difference The expression is:
[0027] ;
[0028] ;
[0029] in, Representing an image In coordinates The grayscale value at the location, with a grayscale value range of 0-255, and t represents the current time;
[0030] Then, motion regions are filtered out by taking the intersection of two frames to obtain the motion region mask. The formula is:
[0031] ;
[0032] Where T represents the pixel threshold, which can be adjusted from 0 to 50 grayscale values depending on the scene complexity;
[0033] When the difference between two frames exceeds the pixel threshold T, the pixel is determined to be a moving region, and the mask value is 1; otherwise, it is a static region, and the mask value is 0.
[0034] As a preferred option: the AI smart checkpoint camera also has a built-in high-precision light sensor. This light sensor is used to collect ambient light intensity data L of the monitored area in real time while collecting video stream. The light intensity data L is accurately associated with the corresponding video frame through timestamps, and the time synchronization error is ≤10ms.
[0035] The AI-powered smart checkpoint camera categorizes light intensity into four scene types and assigns a unique adaptation strategy to each scene, providing an environmental adaptation foundation for subsequent dynamic detection.
[0036] (1) Strong light scene, i.e., L≥10000lux: Enable strong light suppression mode, adjust camera exposure parameters and anti-halo function, set the frame difference method pixel threshold T range to [35,45], and the background difference threshold The range is [30,35], which improves the accuracy of moving target contour recognition in strong light environments;
[0037] (2) Normal lighting scene, i.e., 1000 lux < L < 10000 lux: adopt the standard adaptation mode, keep the camera's default imaging parameters, and set the frame difference method pixel threshold T range to [25, 35] and the background difference threshold to The range is [20,25], which is suitable for most common monitoring environments;
[0038] (3) Low light scene, i.e., 100 lux ≤ L ≤ 1000 lux: Trigger low light enhancement mode, optimize light sensitivity uniformity and target grayscale difference, set the frame difference method pixel threshold T range to [15, 25], and the background difference threshold The range is [10, 15];
[0039] (4) Low-light scene, i.e., L < 100 lux: Activate low-light noise reduction mode to reduce noise interference in dark scenes, and adjust the camera's automatic exposure (AE) and automatic white balance (AWB) parameters in conjunction. Set the frame difference method pixel threshold T range to [10, 20] and the background difference threshold to [10, 20]. The range is [5,10], ensuring the effectiveness of target detection in low-light environments.
[0040] Through the above-mentioned light acquisition and hierarchical adaptation, the subsequent dynamic detection algorithm can adjust parameters in real time according to the ambient light, reducing the risk of false detection and missed detection caused by changes in light from the source, and at the same time adding light scene attributes to the output data.
[0041] This invention solves the problems of false detection and missed detection caused by changes in lighting, such as strong light during the day and dim light at night, by collecting ambient light intensity in real time and constructing a hierarchical adaptive background model, thus ensuring the effectiveness of all-weather monitoring.
[0042] As a preferred embodiment, the background modeling method is used for adaptive background modeling to eliminate interfering factors, and its background model construction and determination process is as follows:
[0043] (1) Background model initialization: When the system starts, the average gray value of the first 100 frames of motionless images obtained through manual annotation or initial static scene determination is taken as the initial background model. The expression is:
[0044] ;
[0045] Where k represents the k-th frame, k∈[1,100];
[0046] (2) Background model dynamic update: After processing each frame of image, the background model is updated according to the current frame. Compared with the previous background model To mitigate the impact of non-target motion on the background, a weighted averaging strategy is employed to update the background model, as expressed in the following expression:
[0047] ;
[0048] in, To update the weights, Update the background only for static areas;
[0049] (3) Motion region verification: Calculate the current frame With background model The difference The expression is:
[0050] ;
[0051] Then the difference Threshold of difference from background When comparing, If the condition is met, it is determined to be a moving area; otherwise, it is determined to be a static area.
[0052] Through the aforementioned mathematical model and process, the front-end AI camera can accurately distinguish moving objects, providing a reliable basis for subsequent target classification and differentiated uploading.
[0053] Preferably, the intelligent NVR module receives real-time video streams or screenshots with real-time timestamps, and also receives corresponding metadata, including light intensity, camera ID and its status label.
[0054] The viewing request includes the target camera ID, viewing type, and lighting scene. The viewing type includes two types: real-time viewing and historical viewing. The lighting scene includes strong light, normal light, weak light, and dark light scenes.
[0055] The intelligent NVR module is the core scheduling unit connecting the front-end AI intelligent checkpoint and the back-end viewing terminal, undertaking the key functions of "data reception and storage, status management, and on-demand push". This module receives differentiated data transmitted from the front end, namely real-time video streams or timestamped screenshots, and combines this with front-end status tags and back-end requests to achieve dynamic bandwidth allocation and precise data push, solving the pain point of traditional NVRs where "full forwarding leads to bandwidth aggregation".
[0056] Preferably, the intelligent NVR module classifies and stores the data according to the status labels (valid dynamic / static scenes) of the data collected by the AI intelligent checkpoint camera, i.e., the associated illumination level labels (strong light / normal light / weak light / dark light).
[0057] For video stream data with a status label of valid dynamic scene, the intelligent NVR module adopts a hierarchical storage strategy: the data is stored in a solid-state drive (SSD) according to a three-level directory structure of "camera ID / valid dynamic scene / light level" to meet the real-time access requirements; only the most recent 24 hours of data are retained in the SSD, and the excess data is automatically migrated to a mechanical hard drive (HDD) with a RAID5 architecture for archiving storage, and the archive data retention period is configured to be 30 days by default, and historical playback function is supported;
[0058] For screenshots of dynamic or static scenes with invalid status tags, the intelligent NVR module directly stores them to the hard drive according to a three-level directory structure of "camera ID / static / light level". A corresponding index file is automatically generated every hour to record the timestamp and storage path information of the screenshots within that time period. The default retention period for static screenshots is 90 days, which can be adjusted according to business needs to support long-term tracking of static scenes.
[0059] Preferably, after receiving the viewing request, the intelligent NVR module queries the "camera status table" based on the camera ID in the request parameters to obtain the current scene status and the latest data timestamp of the corresponding monitoring area.
[0060] When the query result indicates that the current scene status is valid dynamic: If the viewing request does not specify lighting scene filtering conditions, the intelligent NVR module retrieves real-time video stream data from the corresponding storage path (camera ID / valid dynamic / light level) and synchronously associates lighting scene annotation information such as "strong light scene - threshold T=40". The video forwarding module then pushes the "video stream + lighting scene annotation" to the video viewing terminal in a protocol-based manner, such as the real-time streaming protocol RTSP or the full-duplex communication protocol WebSocket. If the backend request includes lighting scene filtering conditions, such as only viewing "dark light scene" videos, the intelligent NVR module quickly filters the corresponding video stream based on the light level tags in the storage directory, associates the annotations, and pushes it to the video viewing terminal.
[0061] When multiple video viewing terminals simultaneously request to view the same camera, the intelligent NVR module adopts a "single-stream multi-push" multiplexing mechanism, that is, it maintains only one video source input and enables multiple terminals to view concurrently through multi-path forwarding, so as to significantly reduce bandwidth consumption.
[0062] When the query result indicates that the current scene status is static or an invalid dynamic scene: If the viewing request does not specify lighting scene filtering conditions, the intelligent NVR module reads the most recent screenshot file with a real-time timestamp and the associated lighting information from the corresponding storage path (camera ID / static / lighting level), and verifies the difference between the screenshot timestamp and the current system time of the video viewing terminal; if the time difference is less than or equal to the set time interval, the "screenshot + lighting annotation" is determined to be valid data and is directly pushed to the video viewing terminal; if the time difference exceeds the set time interval, the front-end camera is triggered to update the screenshot and real-time lighting data before pushing, to ensure that the static image received by the backend has time validity; if the backend request includes lighting scene filtering conditions, the intelligent NVR module filters the latest valid screenshot under the corresponding lighting level based on the lighting intensity information in the index file, associates the annotation, and pushes it to the video viewing terminal.
[0063] The video viewing terminal can be manually triggered to view real-time video streams and supports filtering historical data by lighting scene, further improving the efficiency of viewing monitoring data and meeting the needs of accurate monitoring under different lighting scenes.
[0064] Preferably, the intelligent NVR module has a built-in real-time synchronization module, which is used to synchronize the video stream or screenshots acquired by the front-end AI intelligent checkpoint camera with the output screen of the back-end video viewing terminal, so as to ensure that the back-end video viewing terminal accurately displays the current status of the monitored area.
[0065] The intelligent NVR outputs images to the backend terminal that match the current scene status of the camera: real-time video streams for valid dynamic scenes and timestamped screenshots for static / invalid dynamic scenes. Furthermore, it obtains the same current standard time as the backend terminal through the NVR's built-in real-time time synchronization module (refreshing once per second), achieving the same "real-time time visualization" effect as the real-time video. This ensures that the backend terminal accurately displays the current status of the monitored area, achieving the monitoring goal of "on-demand retrieval, low power consumption, and high efficiency."
[0066] Preferably, the video viewing terminal obtains the viewing type (real-time / historical), target camera ID (single or multiple selections), lighting scene (strong light / normal light / weak light / dark light) and time parameters (historical mode only) through the interactive interface. The video viewing terminal generates a viewing request based on the obtained viewing type, target camera ID, lighting scene and time parameters, and sends it to the smart NVR module.
[0067] After receiving the data returned by the intelligent NVR module, the video viewing terminal parses the data and maps the data from different cameras to the corresponding display windows, thereby enabling parallel loading of multiple video feeds.
[0068] The video viewing terminal serves as the user interaction terminal of the monitoring system. With "on-demand retrieval and scene adaptation" as its core, it enables flexible viewing of multiple cameras through a progressive process of "request initiation - data reception - screen display - interactive operation," while ensuring the real-time nature of time information and the effectiveness of monitoring.
[0069] The video viewing terminal can intelligently adjust the display method according to the data type and scene status: for dynamic scenes, the system plays continuous images in real time to ensure the continuity and immediacy of monitoring; for static scenes, it automatically displays the latest timestamped screenshots to save bandwidth and keep information up-to-date. Simultaneously, each display window overlays unified time information and camera status indicators, allowing users to quickly determine the type and timeliness of the footage. The terminal also supports multi-window parallel display and single-window zoom switching, allowing users to flexibly view different monitoring screens as needed, achieving optimal information visualization and bandwidth resource allocation.
[0070] The beneficial effects of this invention are as follows: By introducing front-end AI dynamic classification and recognition, lighting environment perception and hierarchical adaptive background modeling, intelligent NVR differentiated scheduling, end-to-end time synchronization, and back-end real-time time and lighting scene linkage display into the entire video processing process, this method significantly improves bandwidth utilization, monitoring data reliability, and environmental adaptability in complex lighting and multi-terminal concurrent scenarios such as traffic checkpoints, park security, and smart cities. This method can effectively reduce the network and storage costs of multi-channel monitoring, avoid the problems of invalid dynamic triggering redundant transmission, multi-terminal concurrent bandwidth superposition, and recognition deviation caused by lighting interference in traditional solutions, and achieve real-time accurate monitoring of effective targets (people, vehicles, animals) and low-bandwidth efficient coverage of static / complex lighting scenarios. By combining source data filtering and lighting adaptation, intermediate intelligent scheduling and data reuse, and back-end on-demand viewing and scenario-based display, the operating efficiency, economy, and environmental robustness of multi-channel monitoring systems are significantly improved, meeting the core needs of traffic management, security inspection, urban operation and maintenance, and other fields for long-term, multi-terminal, all-weather, low-consumption, and efficient monitoring. Attached Figure Description
[0071] Figure 1 This is a flowchart of the method of the present invention;
[0072] Figure 2 This is a flowchart illustrating the workflow of the AI smart checkpoint camera in this embodiment.
[0073] Figure 3 This is a flowchart of the intelligent NVR module in the embodiment. Detailed Implementation
[0074] The present invention will be further described in detail below with reference to the accompanying drawings and specific examples. The following embodiments or drawings are used to illustrate the present invention, but are not intended to limit the scope of the present invention.
[0075] like Figure 1 As shown, a multi-channel intelligent mixing video method based on AI smart checkpoints includes the following steps:
[0076] Step 1: Construct a multi-channel intelligent mixed-stream video system based on AI smart checkpoints. This multi-channel intelligent mixed-stream video system is equipped with M AI smart checkpoint cameras. All AI smart checkpoint cameras are connected to the same intelligent NVR module. The intelligent NVR module is also connected to N video viewing terminals.
[0077] Step 2: Each AI smart checkpoint camera collects video streams of the corresponding monitoring area in real time, and then uses the built-in lightweight AI model to perform dynamic detection on the video stream. If the dynamic detection result is that there are no moving objects, it is determined to be a static scene; if the dynamic detection result is that there are moving objects, it is determined to be a dynamic scene.
[0078] Then, the built-in target detection model is used to classify moving objects in the dynamic scene into valid targets or invalid targets; when the moving object is a valid target such as a person, vehicle or animal, the AI smart checkpoint camera uploads a real-time video stream to the smart NVR module.
[0079] When the moving object is an invalid target, such as trees or grass being blown by the wind, or when there is no moving object in the static scene, the AI smart checkpoint camera generates a screenshot with a real-time timestamp at a set time interval, such as 30 seconds, and uploads it to the smart NVR module.
[0080] Step 3: The intelligent NVR module classifies and stores the received real-time video streams or screenshots with real-time timestamps;
[0081] Step 4: The video viewing terminal obtains a viewing instruction, generates a viewing request for the target camera based on the viewing instruction, and then sends the viewing request to the smart NVR module;
[0082] Step 5: The intelligent NVR module queries the target camera status according to the viewing request. If the target camera status is a valid dynamic scene, the intelligent NVR module forwards the video stream to the video viewing terminal in real time.
[0083] If the scene is invalid (either dynamic or static), the smart NVR module pushes screenshots with real-time timestamps to the video viewing terminal at set time intervals.
[0084] Step 6: The video viewing terminal receives real-time video streams or screenshots with real-time timestamps pushed from the intelligent NVR module. If it is a video stream, it will be played in real time; if it is a screenshot, it will be displayed in a carousel or other format. At the same time, the real-time video stream can be manually triggered for viewing to meet different monitoring needs.
[0085] This invention combines front-end AI dynamic classification and recognition, illumination environment perception and hierarchical adaptive background modeling, intelligent NVR differentiated scheduling and real-time time synchronization technology. It utilizes AI recognition for precise filtering of valid dynamics, illumination factors for adaptation to different scene background models, and NVR intelligent control of data transmission to achieve bandwidth optimization from the source to the transmission stage and improve the accuracy of monitoring in all-weather scenarios. By distinguishing between valid and invalid dynamics at the front end, collecting illumination intensity in real time and matching it with the corresponding background model, pushing video streams or timestamped screenshots to the NVR on demand, and single-stream multi-push in multi-terminal concurrent scenarios, this method effectively solves the problems of bandwidth overload, resource waste, and poor adaptability to illumination scenes in traditional solutions. It also ensures the real-time nature, timeliness, traceability, and recognition accuracy of monitoring data in complex environments. This method provides low-power, efficient, intelligent, and more robust all-weather technical support for multi-channel video surveillance scenarios such as traffic monitoring, park security, and smart cities, meeting the practical needs of large-scale, multi-terminal, and long-term stable monitoring.
[0086] The front-end module of this invention is an AI-powered intelligent checkpoint camera, serving as the core data acquisition, environmental perception, and processing unit for implementing a multi-channel intelligent mixed-stream video method. This module uses a built-in AI model to dynamically detect and classify video content within the monitored area. Simultaneously, it integrates a light sensor to collect real-time ambient light intensity data and perform hierarchical adaptation, matching corresponding background models for different lighting scenarios to avoid misjudgments / missed detections caused by light interference. Finally, combining the dynamic recognition results with light adaptation logic, it determines whether to upload a real-time video stream synchronized with lighting information or a screenshot with real-time timestamps and lighting annotations, thereby optimizing bandwidth and ensuring the effectiveness of 24 / 7 monitoring from the data source. Figure 2 As shown, the specific workflow of the AI smart checkpoint camera is as follows:
[0087] Step A1: Input stage: The input of the front-end AI smart checkpoint camera is the real-time video stream of its monitored area, which is the basic data source for all subsequent processing.
[0088] Step A2: The AI smart checkpoint camera also has a built-in high-precision light sensor. This light sensor is used to collect ambient light intensity data L of the monitored area in real time while collecting video stream. The light intensity data L is accurately associated with the corresponding video frame through timestamps, and the time synchronization error is ≤10ms.
[0089] The AI-powered smart checkpoint camera categorizes light intensity into four scene types and assigns a unique adaptation strategy to each scene, providing an environmental adaptation foundation for subsequent dynamic detection.
[0090] (1) Strong light scene, i.e., L≥10000lux: Enable strong light suppression mode, adjust camera exposure parameters and anti-halo function, set the frame difference method pixel threshold T range to [35,45], and the background difference threshold The range is [30,35], which improves the accuracy of moving target contour recognition in strong light environments;
[0091] (2) Normal lighting scene, i.e., 1000 lux < L < 10000 lux: adopt the standard adaptation mode, keep the camera's default imaging parameters, and set the frame difference method pixel threshold T range to [25, 35] and the background difference threshold to The range is [20,25], which is suitable for most common monitoring environments;
[0092] (3) Low light scene, i.e., 100 lux ≤ L ≤ 1000 lux: Trigger low light enhancement mode, optimize light sensitivity uniformity and target grayscale difference, set the frame difference method pixel threshold T range to [15, 25], and the background difference threshold The range is [10, 15];
[0093] (4) Low-light scene, i.e., L < 100 lux: Activate low-light noise reduction mode to reduce noise interference in dark scenes, and adjust the camera's automatic exposure (AE) and automatic white balance (AWB) parameters in conjunction. Set the frame difference method pixel threshold T range to [10, 20] and the background difference threshold to [10, 20]. The range is [5,10], ensuring the effectiveness of target detection in low-light environments.
[0094] Through the above-mentioned light acquisition and hierarchical adaptation, the subsequent dynamic detection algorithm can adjust parameters in real time according to the ambient light, reducing the risk of false detection and missed detection caused by changes in light from the source, and at the same time adding light scene attributes to the output data.
[0095] Step A3: Dynamic Detection Stage: The camera's built-in AI model first performs dynamic detection on the input real-time video stream to determine whether there are moving objects in the monitored area. This AI model uses a fusion algorithm of "frame difference method + background modeling" to achieve accurate identification of moving objects. The core is to distinguish between real moving targets and interference factors such as light and shadow, and jitter through mathematical formula calculations and model iteration. The specific method is as follows:
[0096] 1. Frame difference method: Initially locates the moving area.
[0097] The lightweight AI model calculates the pixel difference between consecutive video frames using the frame difference method to quickly locate potential motion regions.
[0098] The first step is to calculate the difference between adjacent frames, that is, to calculate the difference between the current frame and the adjacent frame. With the previous frame The difference and the current frame With the next frame The difference The expression is:
[0099] ;
[0100] ;
[0101] in, Representing an image In coordinates The grayscale value at the location, with a grayscale value range of 0-255, and t represents the current time;
[0102] Then, motion regions are filtered out by taking the intersection of two frames to obtain the motion region mask. The formula is:
[0103] ;
[0104] Where T represents the pixel threshold, which can be adjusted from 0 to 50 grayscale values depending on the scene complexity;
[0105] When the difference between two frames exceeds the pixel threshold T, the pixel is determined to be a moving region, and the mask value is 1; otherwise, it is a static region, and the mask value is 0.
[0106] 2. Adaptive background modeling eliminates interfering factors. The background model construction and judgment process is as follows:
[0107] (1) Background model initialization: When the system starts, the average gray value of the first 100 frames of motionless images obtained through manual annotation or initial static scene determination is taken as the initial background model. The expression is:
[0108] ;
[0109] Where k represents the k-th frame, k∈[1,100];
[0110] (2) Background model dynamic update: After processing each frame of image, the background model is updated according to the current frame. Compared with the previous background model To address the differences, a weighted averaging strategy is employed to update the background model, reducing the impact of non-target motion on the background. The expression is as follows:
[0111] ;
[0112] in, To update the weights, Update the background only for static areas;
[0113] (3) Motion region verification: Calculate the current frame With background model The difference The expression is:
[0114] ;
[0115] Then the difference Threshold of difference from background When comparing, If the condition is met, it is determined to be a moving area; otherwise, it is determined to be a static area.
[0116] When both the frame difference method and the background modeling method determine that the region is a moving region, it is considered a real moving region, and the final moving region mask value is 1; otherwise, it is a static region, and the final moving region mask value is 0. The expression is:
[0117] ;
[0118] in, This is the final motion region mask value; This represents the motion region mask using the frame difference method. This indicates that the frame difference method determines the motion region. Indicates the current frame With background model The difference, The background difference threshold is set to a grayscale value between 0 and 40. The result of the background modeling method is considered to be the motion area.
[0119] Step A4: Classification and Processing of Moving Objects Branch:
[0120] No moving object branch: If the dynamic detection result indicates no moving objects, the scene is considered static. In this case, the camera initiates a mechanism to capture screenshots at a fixed frequency, such as every 30 seconds, and embeds a real-time timestamp into the screenshot. This timestamp is synchronized with a standard time server via the NTP network time protocol, with an error of no more than 1 second to ensure time accuracy, and then a screenshot with a real-time timestamp is generated.
[0121] Moving Object Detection Branch: When a moving object is detected, the built-in object detection model, such as MobileNet-YOLOv5s, further classifies the moving object to determine whether it is a valid target, such as a person, animal, or vehicle. If the classification result is a valid target, a real-time video stream is generated so that the backend can monitor key dynamics in real time; if the classification result is an invalid dynamic, such as wind blowing through grass, the process of taking screenshots by frequency and embedding real-time timestamps is also initiated to generate screenshots with real-time timestamps.
[0122] Step A5: Output Stage: Depending on the processing branch, the front-end module ultimately outputs a real-time video stream with illumination information or a screenshot with illumination information and a real-time timestamp, providing data support for subsequent intelligent NVR scheduling and back-end viewing, and realizing bandwidth optimization and time information retention from the data acquisition source.
[0123] The intelligent NVR module is the core scheduling unit connecting the front-end AI intelligent checkpoint and the back-end viewing terminal, undertaking the key functions of "data reception and storage - status management - on-demand push". This module receives differentiated data transmitted from the front end, namely real-time video streams or screenshots with timestamps and associated lighting scene information. Combined with front-end status tags including lighting level tags and back-end requests, it achieves dynamic bandwidth allocation, lighting scene-related scheduling, and precise data push, solving the pain point of traditional NVRs where "full forwarding leads to bandwidth aggregation". Figure 3 As shown, the specific workflow of the intelligent NVR module is as follows:
[0124] Step B1: Input Stage: The input to a smart NVR includes two core types of data:
[0125] Front-end data: Receives data uploaded by the front-end AI smart checkpoint, namely real-time video streams or screenshots with real-time timestamps, and simultaneously receives accompanying metadata such as light intensity, camera ID, and front-end status tags.
[0126] Backend Request: Receives viewing requests from the backend. The request content includes the target camera ID, viewing type, and lighting scene. Viewing type includes real-time mode and historical mode, and lighting scene includes strong light, normal light, weak light, and dark light scenes.
[0127] Step B2: Data Storage and Status Management Stage: The intelligent NVR module classifies and stores the data according to the status tags and associated illumination level tags of the data collected by the AI intelligent checkpoint camera, ensuring accurate binding of illumination information with video streams / screenshots.
[0128] For video stream data with a status label of "valid dynamic scene", the intelligent NVR module adopts a hierarchical storage strategy: the data is stored in a three-level directory structure of "camera ID / valid dynamic scene / light level" on a solid-state drive (SSD) to meet real-time access requirements; only the most recent 24 hours of data are retained in the SSD, and the excess data is automatically migrated to a mechanical hard drive (HDD) with a RAID5 architecture for archiving storage, and the archive data retention period is configured, with a default of 30 days. It supports the historical playback function of filtering by light level, which facilitates the quick location of dynamic video under specific lighting conditions.
[0129] For screenshots of dynamic or static scenes with invalid status tags, the intelligent NVR module directly stores them to the hard disk according to a three-level directory structure of "camera ID / static / light intensity". It automatically generates a corresponding index file every hour to record the timestamp, storage path and light intensity information of the screenshots within that time period. The static screenshot retention period is 90 days by default, which can be adjusted according to business needs. It supports tracing static data by lighting scene to improve the efficiency of historical data query.
[0130] Step B3: Request Processing and Image Push Stage: After receiving the viewing request, the intelligent NVR module queries the "Camera Status and Illumination Level Table" based on the camera ID, viewing type, and lighting scene in the request parameters to obtain the current scene status, latest data timestamp, and real-time illumination level information of the corresponding monitoring area.
[0131] When the query result indicates that the current scene status is valid dynamic: If the viewing request does not specify lighting scene filtering conditions, the intelligent NVR module retrieves real-time video stream data from the corresponding storage path (camera ID / valid dynamic / light level) and synchronously associates lighting scene annotation information such as "strong light scene - threshold T=40". The video forwarding module then pushes the "video stream + lighting scene annotation" to the video viewing terminal in a protocol-based manner, such as the real-time streaming protocol RTSP or the full-duplex communication protocol WebSocket. If the backend request includes lighting scene filtering conditions, such as only viewing "dark light scene" videos, the intelligent NVR module quickly filters the corresponding video stream based on the light level tags in the storage directory, associates the annotations, and pushes it to the video viewing terminal.
[0132] When multiple video viewing terminals simultaneously request to view the same camera, the intelligent NVR module adopts a "single-stream multi-push" multiplexing mechanism, that is, it maintains only one video source input and enables multiple terminals to view concurrently through multi-path forwarding, so as to significantly reduce bandwidth consumption.
[0133] When the query result indicates that the current scene status is static or an invalid dynamic scene: If the viewing request does not specify lighting scene filtering conditions, the intelligent NVR module reads the most recent screenshot file with a real-time timestamp and the associated lighting information from the corresponding storage path (camera ID / static / lighting level), and verifies the difference between the screenshot timestamp and the current system time of the video viewing terminal; if the time difference is less than or equal to the set time interval, the "screenshot + lighting annotation" is determined to be valid data and is directly pushed to the video viewing terminal; if the time difference exceeds the set time interval, the front-end camera is triggered to update the screenshot and real-time lighting data before pushing, to ensure that the static image received by the backend has time validity; if the backend request includes lighting scene filtering conditions, the intelligent NVR module filters the latest valid screenshot under the corresponding lighting level based on the lighting intensity information in the index file, associates the annotation, and pushes it to the video viewing terminal.
[0134] Step 2.4 Output Stage: The intelligent NVR outputs images to the backend terminal that match the current scene status and lighting level of the camera. For valid dynamic scenes, it outputs real-time video streams and lighting scene annotations; for static / invalid dynamic scenes, it outputs timestamped screenshots and lighting scene annotations. The annotations include the lighting scene type, such as strong light scene or low light scene. Furthermore, through the NVR's built-in real-time time synchronization module, it obtains the same current standard time as the backend terminal (refreshed once per second), achieving the same "real-time time visualization" effect as the real-time video. This ensures that the backend terminal accurately displays the current status of the monitored area, ambient lighting information, and time validity, achieving the monitoring goals of on-demand retrieval, lighting correlation, and low power consumption with high efficiency.
[0135] The video viewing terminal serves as the user interaction terminal of the monitoring system. With "on-demand retrieval and scene adaptation" as its core, it enables flexible viewing of multiple cameras through a progressive process of "request initiation - data reception - screen display - interactive operation," while ensuring the real-time nature of time information and the effectiveness of monitoring.
[0136] The video viewing terminal obtains the viewing type (real-time / historical), target camera ID (single or multiple selection), and time parameters (historical mode only) through the interactive interface. The video viewing terminal generates a viewing request based on the obtained viewing type, target camera ID, and time parameters, and sends it to the smart NVR module.
[0137] After receiving the data returned by the intelligent NVR module, the video viewing terminal parses the data and maps the data from different cameras to the corresponding display windows, thereby enabling parallel loading of multiple video feeds.
[0138] The video viewing terminal may be a large monitoring screen, a mobile terminal, a PC client, or any combination of the three.
[0139] The video viewing terminal can intelligently adjust the display method according to the data type and scene status: for dynamic scenes, the system plays continuous images in real time to ensure the continuity and timeliness of monitoring; for static scenes, it automatically displays the latest timestamped screenshots to save bandwidth and keep information up-to-date. Simultaneously, each display window overlays unified time information and camera status indicators, allowing users to quickly determine the type and timeliness of the footage. The terminal also supports multi-window parallel display and single-window zoom switching, allowing users to flexibly view different monitoring screens as needed, achieving optimal information visualization and bandwidth resource allocation.
[0140] This invention proposes a multi-channel intelligent mixed-stream video method based on AI-powered smart checkpoints. This method enables intelligent front-end data filtering, transmission bandwidth optimization, and differentiated terminal display in multi-camera monitoring environments, while simultaneously improving monitoring accuracy in complex lighting conditions. Its beneficial effects include:
[0141] (1) Front-end AI intelligent checkpoint data acquisition and screening: The intelligent checkpoint camera, which embeds a lightweight AI model and a high-precision light sensor, achieves accurate detection of moving targets and dynamic scene recognition through an algorithm that combines frame difference analysis and adaptive background modeling. On the other hand, it collects ambient light intensity data in real time and dynamically adjusts the detection threshold and imaging parameters according to four types of scenes: strong light (L≥10000lux), normal light (1000lux<L<10000lux), weak light (100lux≤L≤1000lux), and dark light (<100lux), thereby reducing the risk of false detection and missed detection from the perspective of environmental adaptation. At the same time, it combines a lightweight target detection model to identify effective objects such as people, vehicles, and animals, and outputs real-time video streams containing light information only for effective dynamic scenes. For static or invalid dynamic scenes, it periodically outputs screenshots with light information and timestamps, thereby reducing invalid transmission from the data source and significantly reducing bandwidth consumption.
[0142] (2) Intelligent NVR Data Reception, Illumination-Related Storage and Scheduling Management: For differentiated data containing illumination information transmitted from the front end, an association storage strategy of "scene type + illumination level" and a hierarchical storage architecture are adopted (dynamic video streams are stored on SSD, static screenshots are stored on HDD), and the index file synchronously records illumination intensity information. The system maintains the front end status in real time through a camera status mapping table containing illumination level fields, supports data filtering by illumination scene, and combines a "single stream multi-push" multiplexing mechanism to realize multi-channel real-time forwarding of dynamic scenes (including illumination annotation) and periodic synchronous updates of static scenes, avoiding the bandwidth superposition caused by the full forwarding of traditional NVRs, while ensuring the end-to-end continuity of illumination information.
[0143] (3) On-demand viewing and adaptive display on the backend terminal: The terminal automatically adjusts the display mode based on the scene type. For dynamic scenes, continuous images are displayed in real time; for static or invalid dynamic scenes, the latest screenshot is displayed and the time information and camera status are displayed synchronously. The system supports multi-window parallel display and single-window zoom switching, and users can flexibly view according to monitoring needs, realizing efficient visualization of information and intelligent scheduling of resources.
[0144] In summary, this invention combines key technologies such as artificial intelligence, video encoding, intelligent data scheduling, and ambient light perception. Compared to traditional full-video transmission methods, it achieves significant bandwidth savings in scenarios involving multi-channel video surveillance and concurrent viewing by multiple terminals. Simultaneously, by using a light sensor to collect ambient light data in real time and construct a hierarchical background model, it solves the problems of false detection and missed detection in complex lighting scenarios such as strong light and low light, further ensuring the real-time performance of effective dynamic target monitoring, the timeliness of static scene information, and adaptability to all-weather scenarios. This method provides intelligent, low-power, efficient, and more robust all-weather technical support for fields such as traffic management, park security, and smart cities that rely on multi-channel video surveillance.
[0145] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A multi-channel intelligent mixed-stream video method based on AI intelligent checkpoint, characterized in that, Includes the following steps: Step 1: Construct a multi-channel intelligent mixed-stream video system based on AI smart checkpoints. This multi-channel intelligent mixed-stream video system is equipped with M AI smart checkpoint cameras. All AI smart checkpoint cameras are connected to the same intelligent NVR module. The intelligent NVR module is also connected to N video viewing terminals. Step 2: Each AI smart checkpoint camera collects video streams of the corresponding monitoring area in real time, and then uses the built-in lightweight AI model to perform dynamic detection on the video stream. If the dynamic detection result is that there are no moving objects, it is determined to be a static scene; if the dynamic detection result is that there are moving objects, it is determined to be a dynamic scene. Then, moving objects in the dynamic scene are classified as valid targets or invalid targets; when a moving object is a valid target, the AI smart checkpoint camera uploads a real-time video stream to the smart NVR module; When the moving object is an invalid target or the scene is static, the AI smart checkpoint camera generates screenshots with real-time timestamps at set time intervals and uploads them to the smart NVR module. Step 3: The intelligent NVR module classifies and stores the received real-time video streams or screenshots with real-time timestamps; Step 4: The video viewing terminal obtains a viewing instruction, generates a viewing request for the target camera based on the viewing instruction, and then sends the viewing request to the smart NVR module; Step 5: The intelligent NVR module queries the target camera status according to the viewing request. If the target camera status is a valid dynamic scene, the intelligent NVR module forwards the video stream to the video viewing terminal in real time. If the scene is invalid (either dynamic or static), the smart NVR module pushes screenshots with real-time timestamps to the video viewing terminal at set time intervals. Step 6: The video viewing terminal receives a real-time video stream or a screenshot with a real-time timestamp pushed from the intelligent NVR module. If it is a video stream, it is played in real time; if it is a screenshot, it is displayed in a carousel format.
2. A multi-channel intelligent mixed-stream video method based on AI intelligent checkpoint as described in claim 1, characterized in that: In step 2, the lightweight AI model uses a fusion algorithm of frame difference method + background modeling to perform dynamic detection and obtain the final motion region mask value. When both the frame difference method and the background modeling method determine that the region is a moving region, it is judged as a real moving region, and the final moving region mask value is 1. Otherwise, it is a static region, and the final motion region mask value is 0, expressed as: ; in, This is the final motion region mask value; This represents the motion region mask using the frame difference method. This indicates that the frame difference method determines the motion region. Indicates the current frame With background model The difference, Use the background difference threshold; The result of the background modeling method is considered to be the motion area.
3. A multi-channel intelligent mixed-stream video method based on AI intelligent checkpoint as described in claim 2, characterized in that: The specific determination process of the frame difference method is as follows: The lightweight AI model calculates the pixel difference between consecutive video frames using the frame difference method to locate existing motion regions. The first step is to calculate the difference between adjacent frames, that is, to calculate the difference between the current frame and the adjacent frame. With the previous frame The difference and the current frame With the next frame The difference The expression is: ; ; in, Representing an image In coordinates The grayscale value at that location, where t represents the current time; Then, motion regions are filtered out by taking the intersection of two frames to obtain the motion region mask. The formula is: ; Where T represents the pixel threshold; When the difference between two frames exceeds the pixel threshold T, the pixel is determined to be a moving region, and the mask value is 1; otherwise, it is a static region, and the mask value is 0.
4. A multi-channel intelligent mixed-stream video method based on AI intelligent checkpoint as described in claim 3, characterized in that: The AI smart checkpoint camera also has a built-in light sensor, which is used to collect ambient light intensity data L of the monitored area in real time while collecting video streams. The AI-powered smart checkpoint camera categorizes light intensity into four scene types and assigns a unique adaptation strategy to each scene, providing an environmental adaptation foundation for subsequent dynamic detection. (1) Strong light scene, i.e., L≥10000lux: Enable strong light suppression mode, set the frame difference method pixel threshold T range to [35,45], and the background difference threshold The range is [30, 35]; (2) Normal lighting scene, i.e., 1000 lux < L < 10000 lux: adopt the standard adaptation mode, set the frame difference method pixel threshold T range to [25, 35], and the background difference threshold The range is [20, 25]; (3) Low light scene, i.e., 100 lux ≤ L ≤ 1000 lux: Trigger low light enhancement mode, set the frame difference method pixel threshold T range to [15, 25], and the background difference threshold The range is [10, 15]; (4) Low-light scene, i.e., L < 100 lux: Start low-light noise reduction mode, set the frame difference method pixel threshold T range to [10, 20], and the background difference threshold The range is [5, 10].
5. A multi-channel intelligent mixed-stream video method based on AI intelligent checkpoint as described in claim 2, characterized in that: The background model construction and determination process of the background modeling method is as follows: (1) Background model initialization: When the system starts, the average gray value of the first 100 frames of images without motion is taken as the initial background model. The expression is: ; Where k represents the k-th frame, k∈[1,100]; (2) Background model dynamic update: After processing each frame of image, the background model is updated according to the current frame. Compared with the previous background model The difference is used to update the background model using a weighted average strategy, expressed as: ; in, To update the weights, Update the background only for static areas; (3) Motion region verification: Calculate the current frame With background model The difference The expression is: ; Then the difference Threshold of difference from background When comparing, If the condition is met, it is determined to be a moving area; otherwise, it is determined to be a static area.
6. A multi-channel intelligent mixed-stream video method based on AI intelligent checkpoint as described in claim 1, characterized in that: While receiving real-time video streams or screenshots with real-time timestamps, the intelligent NVR module also receives corresponding metadata, including light intensity, camera ID, and its status label. The viewing request includes the target camera ID, viewing type, and lighting scene. The viewing type includes two types: real-time viewing and historical viewing. The lighting scene includes strong light, normal light, weak light, and dark light scenes.
7. A multi-channel intelligent mixed-stream video method based on AI intelligent checkpoint as described in claim 6, characterized in that: The intelligent NVR module categorizes and stores the data based on the status tags, i.e., the associated light level tags, of the data collected by the AI intelligent checkpoint camera: For video stream data with a status label of valid dynamic scene, the intelligent NVR module adopts a hierarchical storage strategy: the data is stored in a solid-state drive (SSD) according to a three-level directory structure of "camera ID / valid dynamic / light level"; only the data of the most recent 24 hours is retained in the SSD, and the excess part is automatically migrated to a mechanical hard drive (HDD) with a RAID5 architecture for archiving storage. For screenshots of dynamic or static scenes with invalid status tags, the intelligent NVR module directly stores them to the hard disk according to a three-level directory structure of "camera ID / static / light level". It automatically generates a corresponding index file every hour to record the timestamp and storage path information of the screenshots within that time period.
8. A multi-channel intelligent mixed-stream video method based on AI intelligent checkpoint as described in claim 6, characterized in that: After receiving the viewing request, the intelligent NVR module queries the "camera status table" based on the camera ID in the request parameters to obtain the current scene status and the latest data timestamp of the corresponding monitoring area. When the query result indicates that the current scene status is valid and dynamic: if the viewing request does not specify lighting scene filtering conditions, the intelligent NVR module retrieves real-time video stream data from the corresponding storage path and synchronously associates it with lighting scene annotation information. The "video stream + lighting scene annotation" is then pushed to the video viewing terminal via the video forwarding module in a protocol-based manner. If the backend request contains video with lighting scene filtering conditions, the intelligent NVR module quickly filters the corresponding video stream based on the lighting level tags in the storage directory, associates the annotations, and pushes it to the video viewing terminal. When multiple video viewing terminals simultaneously request to view the same camera, the intelligent NVR module adopts a "single-stream multi-push" multiplexing mechanism, that is, it maintains only one video source input and achieves concurrent viewing by multiple terminals through multi-stream forwarding; When the query result indicates that the current scene status is static or an invalid dynamic scene: if the viewing request does not specify lighting scene filtering conditions, the intelligent NVR module reads the most recent screenshot file with a real-time timestamp and the associated lighting information from the corresponding storage path, and verifies the difference between the screenshot timestamp and the current system time of the video viewing terminal. If the time difference is less than or equal to the set time interval, the "screenshot + lighting annotation" is determined to be valid data and is directly pushed to the video viewing terminal; If the time difference exceeds the set time interval, the front-end camera will be triggered to update the screenshot and real-time lighting data before being pushed, so as to ensure that the static image received by the back-end has time validity. If the backend request includes lighting scene filtering conditions, the intelligent NVR module uses the lighting intensity information in the index file to filter the latest valid screenshots under the corresponding lighting level, associates and annotates them, and then pushes them to the video viewing terminal.
9. A multi-channel intelligent mixed-stream video method based on AI intelligent checkpoint as described in claim 1, characterized in that: The intelligent NVR module has a built-in real-time time synchronization module, which is used to synchronize the video stream or screenshots acquired by the front-end AI intelligent checkpoint camera with the output screen of the back-end video viewing terminal, so as to ensure that the back-end video viewing terminal accurately displays the current status of the monitored area.
10. A multi-channel intelligent mixed-stream video method based on AI intelligent checkpoint as described in claim 1, characterized in that: The video viewing terminal obtains the viewing type, target camera ID, lighting scene, and time parameters through the interactive interface. The video viewing terminal generates a viewing request based on the obtained viewing type, target camera ID, lighting scene, and time parameters, and sends it to the smart NVR module. After receiving the data returned by the intelligent NVR module, the video viewing terminal parses the data and maps the data from different cameras to the corresponding display windows, thereby enabling parallel loading of multiple video feeds.