A brain-inspired bio-inspired asynchronous event-driven sparse computing method for massive video swarm intelligence
By employing a brain-inspired, bio-inspired, asynchronous event-driven, massive video swarm intelligence sparse computing method, an edge-cloud collaborative architecture was constructed, solving the resource overload problem of city-level visual swarm intelligence perception systems and achieving efficient, low-power video processing and improved perception accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HANGZHOU DIANZI UNIV
- Filing Date
- 2026-04-16
- Publication Date
- 2026-06-02
AI Technical Summary
Existing city-level visual crowd sensing systems suffer from resource overload due to their fixed-frequency synchronous processing architecture, making them unable to effectively handle high-value events. This results in wasted computing, storage, and transmission resources, and a lack of global adaptive feedback and multi-level collaboration, leading to system crashes.
We employ a brain-inspired, bio-inspired, asynchronous event-driven massive video swarm intelligence sparse computing method. By generating sparse data on the edge side, extracting features through dual-stream collaboration on the edge side, spatiotemporal fusion in the middle layer, and global situational awareness aggregation in the cloud, we construct an edge-cloud hierarchical collaborative architecture. This allows us to dynamically adjust event triggering and pulse emission thresholds to achieve adaptive control of resources across the entire link.
It reduces system power consumption and resource overhead, improves sensing accuracy and efficiency, enables low-latency processing of high-value events, and supports large-scale deployment of massive video devices at the city level.
Smart Images

Figure CN122135181A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of neuromorphic computing technology, and in particular to a neuromorphic bionic asynchronous event-driven method for massive video swarm intelligence sparse computing. Background Technology
[0002] In recent years, city-level video big data applications such as intelligent transportation, security monitoring, and city brain have developed rapidly. The deep integration of artificial intelligence algorithms with edge-cloud distributed computing architecture has driven the large-scale deployment of intelligent visual perception technology. Video, as a core big data form in the city's perception system, has seen widespread deployment of collaborative sensing based on video processing. Typical scenarios include cross-camera image search in traffic monitoring, traffic situation awareness, and public safety target tracking and semantic retrieval. City-level applications require parallel analysis of data from thousands to tens of thousands of cameras to complete target detection, tracking, and feature extraction. Edge-cloud collaboration supports large-scale retrieval of terabytes of data and hundreds of millions of samples, forming the core perception foundation of the city brain.
[0003] Current mainstream city-level visual crowd sensing systems adopt a fixed-frequency synchronous processing architecture: edge cameras acquire and fully encode video at a fixed frame rate, while edge and cloud systems perform full inference at fixed intervals. Regardless of whether there are valid events or high-value targets in the footage, computing, storage, and transmission resources are continuously consumed. When tens of thousands of cameras are deployed on a large scale, the system's end-to-end resource overhead is extremely high, making it impossible to guarantee low-latency processing of high-value events.
[0004] Spiking neural networks (SNNs) in the field of neuromorphic computing mimic the spiking mechanism of biological neurons, possessing event-driven and sparse computing characteristics, and providing new solutions to industry pain points. The hierarchical asynchronous perception mechanism of the human visual system—retina-lateral geniculate nucleus-visual cortex—generates neural impulses only in response to scene changes, achieving efficient bandwidth compression and low-power perception.
[0005] Existing edge-cloud event-driven technologies and SNN event-aware methods have the following drawbacks: 1. Traditional fixed-frequency synchronous architecture leads to overload of full-link resources: indiscriminate transmission of static background redundant video streams results in wasted uplink bandwidth, and the cloud faces a computing power bottleneck of "unable to transmit, unable to store, and unable to compute".
[0006] 2. Strong coupling between computation and communication and performance loss of hierarchical semantic awareness: Existing solutions are mostly based on simple threshold triggering, which means full computation and full transmission. Traditional convolutional networks have difficulty decoupling temporal dynamics and spatial semantics, and are prone to misjudging compressed artifacts as valid features, resulting in decreased accuracy and wasted computing power.
[0007] 3. Lack of global adaptive feedback and multi-level collaboration: SNN is limited to single-device edge perception and does not realize the asynchronous mapping of biological vision pathways in the edge-cloud architecture; there is no closed-loop control of real-time load and task priority, and the fixed threshold can easily crowd out the computing power of high-value tasks during data surges, causing system crashes.
[0008] In summary, existing technologies cannot overcome the resource bottleneck across the entire edge-cloud chain while ensuring perception accuracy. There is an urgent need for a brain-like asynchronous event-driven sparse computing method to improve the energy efficiency and performance of city-level swarm intelligence sensing systems. Summary of the Invention
[0009] The purpose of this invention is to provide a brain-inspired bionic asynchronous event-driven massive video swarm intelligence sparse computing method, which can reduce end-to-end resource overhead, improve edge-side energy efficiency while taking into account perception accuracy, realize adaptive regulation of system resources, and form an efficient end-edge-cloud collaborative swarm intelligence perception and processing paradigm.
[0010] To achieve the above objectives, this invention provides a brain-inspired bio-inspired asynchronous event-driven sparse computing method for massive video swarm intelligence, implemented based on an edge-cloud hierarchical collaborative architecture consisting of the edge, middleware, and cloud. The steps are as follows: S1. Edge-side event triggering and sparse data generation: Based on a brain-inspired bionic mechanism, effective motion events are detected, sparse video data blocks are generated and uploaded to edge nodes, and data transmission in areas without event background is suppressed. S2. Edge-side dual-stream collaborative feature extraction: A dual-stream architecture that integrates spiking neural network (SNN) and artificial neural network (ANN) is adopted to asynchronously process sparse video data blocks, generate and upload semantic metadata; S3, Intermediate Layer Spatiotemporal Fusion: Only when a target cross-domain movement event is detected, cross-edge node trajectory stitching and regional situation fusion are performed to generate a regional situation summary; S4, Cloud-based Global Situation Aggregation: Receives semantic metadata and regional situation summaries from each edge node, completes cross-node target association, trajectory stitching and macro-situation analysis, and outputs collective intelligence perception results; During execution, the cloud monitors the resource load across the entire chain in real time and dynamically adjusts the event trigger threshold of S1 and the pulse release threshold of S2 to achieve adaptive allocation of global resources.
[0011] Preferably, step S1 specifically includes: The cumulative motion saliency is calculated at the pixel level or region level based on the video frame sequence, and the cumulative motion saliency is updated using a leakage integral dynamics model. The cumulative amount of motion significance is compared with a dynamic threshold. If it exceeds the dynamic threshold, it is determined to be a valid motion event and a trigger pulse is generated. The system extracts the region of interest in motion in response to the trigger pulse, and generates a sparse video data block after slicing and compression encoding; the sparse video data block is uploaded only when the uplink bandwidth utilization rate on the end side is lower than a preset threshold.
[0012] Preferably, step S2 specifically includes: SNN stream processing of sparse video data blocks provides temporal information and outputs a sequence of impulse events. Using pulse event sequences as gate signals, the ANN stream is activated only during valid events to extract spatial semantic features; The extracted features are encoded into semantic metadata containing target category, location, and feature vectors, and then transmitted upstream instead of the original video stream.
[0013] Preferably, in step S2, the SNN stream uses a convolutional spiking neural network and an IE-LIF neuron model to extract temporal features and determine valid events by accumulating pulse energy; the ANN stream remains physically dormant when there is no pulse gating signal.
[0014] Preferably, the triggering condition for step S3 is: the number of consecutive frames lost from the target trajectory reaches a threshold determined adaptively based on the target's moving speed and the sensor's frame rate, and the target disappears at a location within the sensor's field of view boundary area. At this time, the cross-sensor semantic target association module is activated to complete trajectory stitching and situation summary generation.
[0015] Preferably, step S4 specifically includes: completing cross-node target association and trajectory splicing based on semantic metadata, constructing a spatiotemporal knowledge graph, completing macro-situation analysis through graph neural networks, and outputting the results of collective intelligence perception.
[0016] Preferably, real-time monitoring of the entire link resource load is implemented, specifically: real-time monitoring of processor utilization, memory usage, and network bandwidth usage at each level; and dynamic adjustment of the dynamic threshold of S1 and the SNN pulse firing threshold of S2 through an exponential adaptive function based on the difference between the resource load and the security load threshold, combined with the task priority weight.
[0017] Preferably, when the system resource load exceeds the safe load threshold, the event trigger threshold and pulse issuance threshold are increased to filter low-value events and prioritize the computing power and bandwidth resources of high-priority tasks.
[0018] Preferably, steps S1-S4 all adopt a standardized event-driven process of input buffering-calculation activation gating-state update-transmission activation gating; when no valid event is triggered, the processing unit corresponding to each step remains in a dormant state.
[0019] A computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the above-described method.
[0020] According to specific embodiments provided by the present invention, the present invention discloses the following technical effects: (1) A brain-like bionic leakage integral dynamics detection and dynamic threshold triggering mechanism is adopted on the end side. Only the areas with effective motion in the picture are extracted, encoded and transmitted. The sparse communication mode of no transmission without events is realized from the data acquisition source. The massive static background redundant data brought by the traditional fixed frequency full transmission is completely eliminated, and the uplink communication bandwidth occupation on the end side is greatly reduced. The systemic resource overload bottleneck of "cannot be transmitted, cannot be stored and cannot be calculated" in the scenario of tens of thousands of concurrent video in the city is fundamentally alleviated.
[0021] (2) SNN is used on the edge side The ANN dual-stream decoupled gating architecture uses a low-power SNN stream to process temporal dynamic information and a high-computing-power ANN stream to wake up only when triggered by a pulse to perform spatial feature extraction. This enables ANN units to physically sleep when there are no valid events, significantly reducing the average computational power consumption of edge nodes. At the same time, this architecture can effectively decouple video temporal changes from spatial semantic features, resist compression artifact interference, and ensure the accuracy and stability of target detection and feature extraction while greatly improving computational energy efficiency.
[0022] (3) The middle layer and the cloud adopt an event-triggered spatiotemporal fusion and global situation aggregation mechanism. Cross-node trajectory splicing and regional situation fusion are only initiated when the target moves across domains, avoiding invalid data interaction and computing power waste. The cloud constructs a spatiotemporal knowledge graph based on semantic metadata to complete macro-situation analysis, realizes accurate association and trajectory continuation of cross-camera targets across the entire domain, and greatly improves the global perception and situation judgment capabilities of urban intelligent perception.
[0023] (4) Construct a brain-like bionic asynchronous event-driven architecture for the entire edge-cloud chain. The entire process adopts a sparse computing mode of "default sleep and event wake-up" to form a complete low-power, high-efficiency and high-robust crowd intelligence perception processing paradigm. It can support the large-scale deployment of massive video perception devices at the city level and comprehensively improve the system operation efficiency and practical value of scenarios such as smart transportation, security monitoring, and city brain.
[0024] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0025] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0026] Figure 1This is a schematic diagram of the brain-inspired bionic perception architecture of edge-cloud collaboration according to an embodiment of the present invention; Figure 2 This is a detailed flowchart illustrating the terminal event triggering and sparse data generation process according to an embodiment of the present invention. Figure 3 This is a schematic diagram of the edge-side SNN-ANN dual-stream decoupled feature extraction network according to an embodiment of the present invention; Figure 4 This is a schematic diagram illustrating the control principle of cloud-based global adaptive feedback gating in an embodiment of the present invention. Detailed Implementation
[0027] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0028] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0029] Example To further overcome the bottlenecks in computing power, bandwidth, and energy consumption caused by massive concurrent sensing data at the city level, this invention proposes a brain-inspired bionic asynchronous event-driven massive video swarm intelligence sparse computing method. It constructs a four-level hierarchical architecture: edge-side, middle layer, and cloud layer. This architecture maps one-to-one with the retina, optic nerve, primary visual cortex, and higher visual cortex of the human visual system. The edge-side simulates the early photosensitivity and redundancy filtering functions of the retina, responsible for extracting pixel-level motion saliency and generating sparse visual signals. The edge-side simulates the high-frequency signal transmission and preliminary semantic encoding functions of the optic nerve, responsible for converting sparse visual signals into compact semantic metadata and transmitting it upwards. The middle layer simulates the cross-regional spatiotemporal fusion function of the primary visual cortex, responsible for completing cross-node trajectory stitching and generating regional situation summaries. The cloud layer simulates the global cognition and decision-making functions of the higher visual cortex, responsible for constructing a spatiotemporal knowledge graph and outputting a macroscopic swarm intelligence situation. Meanwhile, through a cloud-based global adaptive gating mechanism, the activation thresholds at each level are dynamically adjusted based on the real-time load and task priority of the entire link, thereby achieving global overall optimization of resources across the entire link. This significantly reduces system power consumption and overhead while ensuring the detection accuracy and response priority of high-value semantic events.
[0030] like Figure 1 As shown, the overall optimization scheme design includes simulating the early redundancy filtering process of the retina, simulating the small-scale high-frequency semantic compensation process of the optic nerve, and simulating the macroscopic coordination process of the brain. The specific steps are as follows: S1, End-side pixel-level sparse sensing and region of interest (RoI) extraction ( ) S1.1 Determine the input structure: The end-side intelligent camera is in... The raw high frame rate multi-source signal data stream (such as video stream) acquired in real time is denoted as... At the same time, a local state buffer is maintained in the node's memory. It is used to store the background Gaussian mixture model (GMM) parameters or reference frame from the previous moment, in order to prepare data for subsequent event detection.
[0031] S1.2, Perform activation and saliency evaluation: extract the current frame. With local state buffer The pixel-level L1 norm difference is used as the semantic saliency evaluation value for the task. Compare this evaluation value with the currently set baseline calculation threshold. In comparison, if If the frame is active, the lightweight motion detection operator on the end side is activated; otherwise, it is determined to be a static background or a minor disturbance, and the frame is kept in sleep mode and discarded directly.
[0032] S1.3, Complete State Update and Biomimetic Leakage Integration: After the motion detection operator is activated, first extract the bounding rectangle containing the moving target in the image, i.e., the region of interest (RoI); then simulate the membrane potential decay characteristics of biological neurons, utilizing the leakage time constant. Update the background model, and the higher the layer... The larger the value, the larger it is. To accommodate longer memory decay periods and the need for long-term storage in higher-level semantics, the update expression is:
[0033] in For the state update function, modeling the GMM background on the endpoint, Online updates are performed only on the GMM parameters within the extracted RoI region, while the model parameters in the static background region remain unchanged, reducing unnecessary computational overhead.
[0034] S1.4, Conduct transmission activation judgment and sparse data output: First, evaluate the information value of the extracted RoIs. (e.g., target pixel area ratio) Then perform the double check if and only if Furthermore, when the uplink bandwidth utilization rate on the device side is lower than a preset threshold (typically 60%), the adaptive frame truncation and encoding operators are driven to compress only these RoI slices, generating sparse video data blocks after redundancy removal. Data is sent to the edge; otherwise, no data is reported, physically blocking the uplink transmission of redundant background data. The specific process is as follows: Figure 2 As shown.
[0035] S2. Asynchronous extraction and local tracking of target-level semantic features at the edge ( ) This stage is completed at the edge nodes, and its core is as follows: Figure 3 The SNN-ANN dual-stream decoupled architecture shown can theoretically achieve zero power consumption of the ANN stream when there are no events by using the SNN stream pulse-gated ANN stream sleep-wake mechanism. Compared with the traditional continuously running dual-stream architecture, the average computing power consumption of edge nodes can be reduced by orders of magnitude. At the same time, high-fidelity feature extraction is only performed when an event is triggered, ensuring the detection accuracy of semantic features.
[0036] S2.1 Determine the input structure and maintain the trajectory pool: Receive sparse video data blocks from one or more end-side nodes in parallel. Simultaneously, local state buffering. The system maintains a local semantic target trajectory pool, which records the target features and location information that have recently appeared at the intersection or within the area in real time, providing a basis for subsequent feature matching.
[0037] S2.2 Execution of computational activation and dual-stream decoupling verification: The SNN stream is responsible for extracting the temporal variation features of the input sparse video blocks. This SNN stream is implemented using a convolutional spiking neural network, and the neurons adopt the IE-LIF model. Its membrane potential update rule is as follows:
[0038] in, express All presynaptic neurons receive input pulses at any given time. After synaptic weight Weighted total input stimulus; This is an information enhancement term, calculated based on the local spatiotemporal gradient of the current input sparse block. It is used to enhance the response to subtle target motion changes and improve the SNN's sensitivity to low-salience events. The spiking threshold for SNN neurons; This indicates that the neuron at the previous time ( The emitted pulse: If Pulse was emitted at all times Otherwise, it is 0; this item enables the membrane potential to be automatically reset to the resting level after the neuron fires a pulse, which is in line with the firing characteristics of biological neurons.
[0039] S2.3, Feature Determination and ANN Branch Wake-up Control: After the SNN stream outputs the pulse accumulated energy at each spatial location, the statistical characteristics of the pulse accumulated energy are used as the global temporal features of the input data block. Then, the cosine distance / prediction error is calculated between the global temporal features output by the SNN stream and the historical features of the corresponding spatial locations in the trajectory pool, and finally quantified into the edge-side activation semantic saliency evaluation value. The activation threshold is calculated on the edge side using... The principle is to perform generalized calculations without fixed parameters; the formula is as follows:
[0040] in, This represents the historical average prediction error. This represents the standard deviation of historical forecast errors. If the abrupt change in time series prediction error satisfies If the event is identified as a valid semantic event, a pulse event sequence is generated and used as a gating signal to wake up the dormant artificial neural network (ANN) branch. After the ANN branch is woken up, it performs spatial feature extraction, including target detection, high-dimensional ReID desensitization identity feature extraction, target attribute recognition, local trajectory fitting and other tasks. Finally, it outputs target-level structured semantic features to provide a data foundation for subsequent trajectory pool updates and uplink transmission. If no valid semantic event is detected, the ANN branch remains physically dormant and does not perform any computational operations.
[0041] Note: Here The activation threshold is calculated at the macroscopic level of the edge nodes, and the threshold is calculated at the neuron level within the SNN stream. It is a two-level progressive filtering relationship: the former controls the wake-up of the entire ANN branch and is the core decision switch at the node level; the latter only controls the pulse firing of a single neuron inside the SNN and is the filtering rule for low-level feature extraction. The two are completely independent in function, level and target, forming a collaborative mechanism that first filters low-level noise and then makes decisions on high-power computations, thereby minimizing the overhead of ineffective computing power.
[0042] S2.4. Perform trajectory pool state update: Utilize the high-dimensional features extracted from the ANN branch (such as the ReID de-identification feature vector) to buffer the local state. The trajectory pool is updated using a larger leakage constant. This enables smooth fusion and long-term memory of cross-frame features, enhancing resistance to video compression artifacts.
[0043] S2.5, Perform transport activation judgment and semantic metadata output: First, perform feature fusion processing on the extracted information and convert it into a unified, highly compressed semantic metadata format. This metadata contains core information such as target category, bounding box, and identity identifier; when the target exhibits a complete cross-domain movement intention, i.e. At that time, the upward transmission of cross-modal fusion features is triggered. ;in This represents the percentage of frames in which the target trajectory continuously exists within the current node's edge field of view. If the percentage is greater than the preset threshold, it means that the target has successfully crossed the domain and entered the current region.
[0044] S3, Intermediate layer spatiotemporal feature fusion and trajectory stitching ( ) S3.1 Determine the input structure and maintain the state database: Receive semantic metadata streams from multiple lower-layer edge nodes in parallel. Meanwhile, local state buffering Maintaining a weighted directed graph of the multi-sensor network topology covering a larger physical area, as well as a spatiotemporal state database of regional targets, provides spatial and data support for cross-node trajectory stitching.
[0045] S3.2 Execution Activation and Cross-Node Trajectory Stitching: First, detect the loss of the target trajectory. When it is detected that the target's trajectory has been lost for ≥ N consecutive frames (N is a positive integer) at the current node, further judgment is made based on the target's disappearance position. The adaptive rule is as follows:
[0046] in As the baseline target's moving speed, For sensor frame rate, Estimate the target's movement speed. This indicates rounding up; the higher the frame rate and the slower the target moves, the larger the value of N becomes. If the location where the target disappears is within the sensor's field of view boundary, it is determined to be a triggered event. The cross-sensor semantic target association algorithm module is immediately activated. After activation, the module calculates the matching probability based on temporal continuity, spatial rationality, and feature correlation. Based on the matching results, it links the cross-camera trajectory sequence and synchronously updates the global state library. .
[0047] S3.3 Conduct transmission activation judgment and regional situation summary output: Gradually merge the discrete cross-sensor target individual trajectories that have been linked into a whole group activity pattern; then generate a regional situation summary in the form of a high-dimensional feature vector. When the number of active targets or the magnitude of situation change in the region exceeds the preset transmission threshold, the situation summary is aggregated and sent to the cloud.
[0048] S4, Global Situation Aggregation of Cloud Brain ( ) Knowledge graph construction and global situation aggregation are performed: After receiving the regional situation summary collected by each local intermediate layer in the cloud, the multi-source situation data is synthesized into a city-level semantic spatiotemporal knowledge graph through a preset graph construction function. This graph contains three core elements: entity nodes, relation edges, and dynamic attributes. Based on this knowledge graph, macro-scenario-based situation analysis is then completed, and the results of crowd intelligence perception, such as crowd gathering and traffic operation, are finally output.
[0049] During the execution of S1-S4, global adaptive feedback control is carried out synchronously, the resource load of the entire link is monitored in real time, and the event trigger threshold of S1 and the pulse issuance threshold of S2 are dynamically adjusted to achieve global resource adaptive allocation. Specifically, this includes: Global Resource-Value Adaptive Gating: Real-time monitoring of processor utilization, memory usage, computing costs, and network bandwidth load across the entire value chain, from the edge to the middleware layer and the cloud. ,like Figure 4 As shown. Using the difference between the preset safe load level and the actual load, an adjustment command is issued in real time through an exponential adaptive function. The command formula is:
[0050]
[0051] in, and The basic activation threshold of the m-th layer. The system has a preset safe load level (typically 70%), and k is an adjustment coefficient. Task priority weights: High priority tasks (Used to lower the threshold and improve sensitivity), normal tasks Low priority tasks (Used to raise the threshold and filter out low-value disturbances).
[0052] Load handling strategy: When faced with massive concurrent data surges (such as abnormal load increases), the cloud immediately instructs lower-layer edge and endpoint nodes to exponentially increase activation thresholds, proactively filtering low-value disturbance events and reducing invalid data and computation; simultaneously, in conjunction with the issued... Global task priority prioritizes the allocation of computing power and network access for high-value core semantic events (such as tracking blacklisted targets), achieving global multi-objective coordinated optimization of computing, storage, and transmission costs.
[0053] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A brain-inspired, bio-inspired, asynchronous event-driven sparse computing method for massive video swarm intelligence, characterized in that, The implementation is based on a hierarchical collaborative architecture consisting of the endpoint, edge, middleware, and cloud, and the steps are as follows: S1. Edge-side event triggering and sparse data generation: Based on a brain-inspired bionic mechanism, effective motion events are detected, sparse video data blocks are generated and transmitted to the edge side, and data transmission in areas without event background is suppressed. S2. Edge-side dual-stream collaborative feature extraction: The edge side receives sparse video data blocks from S1 and uses a dual-stream architecture that integrates a spiking neural network (SNN) and an artificial neural network (ANN) to asynchronously process the sparse video data blocks, generate semantic metadata, and upload them to the middle layer and the cloud respectively. S3, Intermediate Layer Spatiotemporal Fusion: Only when a target cross-domain movement event is detected, cross-edge node trajectory stitching and regional situation fusion are performed to generate a regional situation summary; S4, Cloud-based Global Situation Aggregation: Receives semantic metadata and regional situation summaries from each edge node, completes cross-node target association, trajectory stitching and macro-situation analysis, and outputs collective intelligence perception results; During execution, the cloud monitors the resource load across the entire chain in real time and dynamically adjusts the event trigger threshold of S1 and the pulse release threshold of S2 to achieve adaptive allocation of global resources.
2. The brain-inspired bionic asynchronous event-driven massive video swarm intelligence sparse computing method according to claim 1, characterized in that: Step S1 specifically includes: The cumulative motion saliency is calculated at the pixel level or region level based on the video frame sequence, and the cumulative motion saliency is updated using a leakage integral dynamics model. The cumulative amount of motion significance is compared with a dynamic threshold. If it exceeds the dynamic threshold, it is determined to be a valid motion event and a trigger pulse is generated. The system extracts the region of interest in motion in response to the trigger pulse, and generates a sparse video data block after slicing and compression encoding; the sparse video data block is uploaded only when the uplink bandwidth utilization rate on the end side is lower than a preset threshold.
3. The brain-inspired bionic asynchronous event-driven massive video swarm intelligence sparse computing method according to claim 1, characterized in that: Step S2 specifically includes: SNN stream processing of sparse video data blocks provides temporal information and outputs a sequence of impulse events. Using pulse event sequences as gate signals, the ANN stream is activated only during valid events to extract spatial semantic features; The extracted features are encoded into semantic metadata containing target category, location, and feature vectors, and then transmitted upstream instead of the original video stream.
4. The brain-inspired bionic asynchronous event-driven massive video swarm intelligence sparse computing method according to claim 3, characterized in that: In step S2, the SNN stream uses a convolutional spiking neural network and an IE-LIF neuron model to extract temporal features and determine valid events by accumulating pulse energy; the ANN stream remains physically dormant when there is no pulse gating signal.
5. The brain-inspired bionic asynchronous event-driven massive video swarm intelligence sparse computing method according to claim 1, characterized in that: The triggering condition for step S3 is: the number of consecutive frames lost from the target trajectory reaches a threshold determined adaptively based on the target's moving speed and the sensor's frame rate, and the target disappears at a location within the sensor's field of view boundary area. At this time, the cross-sensor semantic target association module is activated to complete trajectory stitching and situation summary generation.
6. The brain-inspired bionic asynchronous event-driven massive video swarm intelligence sparse computing method according to claim 1, characterized in that: Step S4 specifically includes: completing cross-node target association and trajectory splicing based on semantic metadata, constructing a spatiotemporal knowledge graph, completing macro-situation analysis through graph neural networks, and outputting the results of collective intelligence perception.
7. The brain-inspired bionic asynchronous event-driven massive video swarm intelligence sparse computing method according to claim 1, characterized in that: Real-time monitoring of end-to-end resource load includes: real-time monitoring of processor utilization, memory usage, and network bandwidth usage at each level; and dynamic adjustment of the dynamic threshold of S1 and the SNN pulse firing threshold of S2 using an exponential adaptive function based on the difference between resource load and security load threshold, combined with task priority weights.
8. The brain-inspired bionic asynchronous event-driven massive video swarm intelligence sparse computing method according to claim 7, characterized in that: When the system resource load exceeds the safe load threshold, the event trigger threshold and pulse issuance threshold are increased to filter low-value events and prioritize the computing power and bandwidth resources of high-priority tasks.
9. The brain-inspired bionic asynchronous event-driven massive video swarm intelligence sparse computing method according to claim 1, characterized in that: Steps S1-S4 all adopt a standardized event-driven process of input buffering, calculation activation gating, state update, and transmission activation gating; when no valid event is triggered, the processing unit corresponding to each step remains in a dormant state.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1 to 9.