Intelligent security terminal cooperative monitoring system with edge computing capability
By employing multimodal sensing, edge computing, and distributed collaborative communication technologies, the problems of collaborative processing and communication stability of multimodal security data have been solved, enabling efficient and accurate anomaly detection and resource optimization in multi-terminal collaborative monitoring systems.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-25
- Publication Date
- 2026-04-07
AI Technical Summary
Existing technologies are insufficient for the collaborative processing of multimodal security data, cannot adapt to the complex collaborative processing needs of multiple terminals and multiple types of data, and do not guarantee the communication stability between multiple security terminals or the accuracy of multimodal data fusion decision-making.
Multimodal sensing units are used to collect visual, acoustic, and environmental physical signals. CPU and FPGA computing tasks are allocated through dynamic feature distillation and hardware load feedback technology of edge computing processing units. Combined with signal attenuation gradient prediction and frequency band dynamic switching technology of distributed collaborative communication units, multimodal data collaborative processing and stable transmission are realized. The spatiotemporal alignment and dynamic threshold algorithm of intelligent decision-making units are used to optimize decision accuracy.
It enables effective collaborative processing of multimodal security data, improves the processing efficiency of edge-side sensing data, ensures the stability of communication links and the accuracy of abnormal event judgment when multiple security terminals collaborate, and optimizes the allocation efficiency of edge-side computing power.
Smart Images

Figure CN120881654B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of terminal monitoring, in particular to a smart security terminal cooperative monitoring system with edge computing capability. BACKGROUND
[0002] With the development of Internet of Things and artificial intelligence technology, the real-time and reliability of multi-terminal cooperative monitoring in the field of smart security are put forward higher requirements. Edge computing gradually becomes the core technology direction of smart security terminal system with the advantages of local data processing and low delay response. How to make multi-modal sensing security terminals efficiently cooperate and accurately make decisions on the edge side has become a problem to be solved.
[0003] Chinese patent CN202510593421.8 discloses an intelligent terminal monitoring system and method based on edge computing. For traffic monitoring scene, through calculating the important coefficient of camera and the adaptive coefficient of edge computing center, the data migration and load distribution of fault edge computing center are realized, which guarantees the stability and accuracy of traffic data processing. However, this scheme focuses on edge computing load scheduling in specific traffic scenarios and does not involve multi-modal security data cooperative sensing and dynamic decision-making, which is difficult to be directly applied to more complex multi-terminal smart security scenarios. Chinese patent CN202411961913.X discloses a security monitoring method based on a hybrid of large models and neural network algorithms. Through the real-time detection of light network and the fine analysis of cloud deep model, combined with sample pool differential update strategy to optimize model performance, the balance between edge resource constraint and detection accuracy is achieved. However, this method mainly focuses on edge-cloud cooperation and model optimization at the algorithm level, and does not design the communication cooperation mechanism between multiple terminals, the spatio-temporal alignment of multi-modal data on the edge side, and the dynamic threshold decision of smart security terminal cooperation.
[0004] Despite the design advantages of the aforementioned technical solutions, they also suffer from the following technical shortcomings: First, they lack the ability to collaboratively process multimodal security data. Chinese patent CN202510593421.8 focuses on load scheduling of edge computing centers in traffic monitoring scenarios, addressing only data from a single traffic camera and failing to address the collaborative perception and dynamic decision-making of multimodal data such as visual, acoustic, and environmental data required for smart security. This makes it unsuitable for the complex collaborative processing needs of multiple terminals and data types. Second, they fail to guarantee communication stability and the accuracy of multimodal data fusion decisions during multi-terminal collaboration. While Chinese patent CN202411961913.X achieves edge-cloud collaboration and model optimization at the algorithm level, it does not design a communication collaboration mechanism between multiple security terminals, nor does it address issues such as spatiotemporal alignment and dynamic threshold decision-making for multimodal data at the edge, making it difficult to guarantee communication stability and the accuracy of fusion decisions during multi-terminal collaboration. Therefore, we propose a smart security terminal collaborative monitoring system with edge computing capabilities. Summary of the Invention
[0005] The purpose of this invention is to provide a smart security terminal collaborative monitoring system with edge computing capabilities to solve the problems mentioned in the background art.
[0006] To address the aforementioned technical problems, the present invention aims to provide a smart security terminal collaborative monitoring system with edge computing capabilities, comprising:
[0007] A multimodal sensing unit is used to collect visual, acoustic, and environmental physical signals of the security area to obtain multi-dimensional basic sensing data.
[0008] The edge computing processing unit, based on multimodal sensing data, uses dynamic feature distillation technology and hardware load feedback technology to allocate CPU and FPGA computing tasks, uses incremental convolution kernel pruning technology to optimize the model computing process, judges the computing power requirements and model redundancy of the sensing data processing task, and outputs simplified feature data to improve the efficiency of sensing data processing.
[0009] The distributed collaborative communication unit analyzes wireless transmission quality based on edge terminal interaction data, uses signal attenuation gradient prediction technology and frequency band dynamic switching technology, utilizes fragmented verification relay technology to ensure abnormal event data transmission, judges communication link stability, triggers optimal transmission strategy, and realizes cross-terminal data collaboration.
[0010] The intelligent decision-making unit, based on the simplified feature data output by the edge computing processing unit, spatiotemporally aligns visual dynamic features, voiceprint features and environmental physical features, adjusts the judgment criteria of each feature through a dynamic threshold algorithm based on the real-time environmental baseline, optimizes the decision weight of each modality feature, and achieves accurate judgment of abnormal events.
[0011] A multi-level response execution unit, which triggers physical response actions such as audible and visual alarms, access control, video recording, and lighting linkage based on the anomaly determination result.
[0012] As a further improvement to this technical solution, the multimodal perception unit includes a visual acquisition module, a voiceprint acquisition module, an environmental sensing module, and a data synchronization module, wherein:
[0013] The visual acquisition module includes a CMOS image sensor and an automatic exposure control circuit, used to acquire visual signals of the security area, adapt to a wide range of lighting conditions by adjusting the shutter time, and output standard format video frames carrying optical center coordinates.
[0014] The voiceprint acquisition module includes a MEMS microphone array and a bandpass filter, used to acquire voiceprint signals, enhance the target sound source through specific frequency band filtering and beamforming algorithms, and output an audio stream with a sampling rate adapted to voiceprint feature extraction.
[0015] The environmental sensing module includes a temperature and humidity sensing component and a smoke sensing component, used to collect environmental physical signals;
[0016] The data synchronization module is used to add a unified time stamp to visual, voiceprint and environmental signals, and establish spatial coordinate association based on the installation positions of the visual acquisition module, voiceprint acquisition module, environmental sensing module and data synchronization module, so that the multi-dimensional basic perception data has spatiotemporal matching.
[0017] As a further explanation of this embodiment, the implementation of the unified time stamp of the data synchronization module relies on the real-time clock (RTC) module built into the multimodal sensing unit. The RTC module is synchronized with the global clock of the system to ensure that the time base of all sub-modules is consistent. When the data synchronization module receives video frames from the visual acquisition module, audio streams from the voiceprint acquisition module, and environmental data from the environmental sensing module, it reads the current time of the RTC module in real time and adds a timestamp to each data segment. Specifically, video frames are timestamped frame by frame, audio streams are timestamped according to preset time segments, and environmental data is timestamped according to the acquisition time, so as to avoid time stamp deviation caused by data transmission delay.
[0018] As a further improvement to this technical solution, the edge computing processing unit includes a computing power scheduling module, a feature distillation module, and a model optimization module, wherein:
[0019] The computing power scheduling module collects the operating status parameters of the CPU and FPGA in real time, combines the computational complexity analysis of multimodal data, and completes the targeted allocation of computing tasks between the CPU and FPGA based on hardware load feedback technology.
[0020] The feature distillation module is used to map the original features of visual, voiceprint, and environmental multimodal features to a unified feature space, and extract high-value features and remove redundant feature dimensions through dynamic feature distillation technology.
[0021] The model optimization module is used to analyze the feature contribution of convolution kernels in the computational model to identify redundant convolution kernels, remove redundant parts through incremental convolution kernel pruning technology, and optimize the computing power consumption of the model operation.
[0022] As a further improvement to this technical solution, the task allocation of the computing power scheduling module includes the following steps:
[0023] S210.1, Real-time CPU Utilization Monitoring FPGA resource utilization and memory usage This forms a hardware load status dataset;
[0024] S210.2, Synchronously acquire visual frame size parameters Voiceprint sampling scale Number of environmental data points and the dimension of features to be extracted Determine the baseline for computational complexity of multimodal data;
[0025] S210.3. Based on hardware load feedback technology, the computational complexity benchmark is matched and analyzed with the hardware load status dataset. When the matching result exceeds the parallel processing threshold... When the feature extraction task is assigned to the FPGA, and the matching result does not exceed the parallel processing threshold, the task is performed accordingly. At that time, logical judgment tasks are assigned to the CPU.
[0026] As a further improvement to this technical solution, the collaborative processing of the feature distillation module and the model optimization module includes the following steps:
[0027] S220.1 The feature distillation module receives multimodal basic perception data and converts visual features... Voiceprint characteristics Environmental characteristics Mapping to a unified feature space to generate a fused feature matrix ;
[0028] Calculation using dynamic characteristic distillation technology The correlation between various characteristics and security anomalies Retain features with a correlation degree no lower than the feature importance threshold. Feature subset ;
[0029] The model optimization module processes... The computational model is used to evaluate convolution kernels and calculate the output characteristic fluctuation value of each convolution kernel. and gradient contribution The marker simultaneously satisfies Fluctuation threshold and Contribution threshold Redundant convolution kernels ;
[0030] Remove Post-validation model output error ,when Error allowable threshold At the same time, maintain the simplified model structure and output based on Simplified feature data.
[0031] As a further improvement to this technical solution, the distributed cooperative communication unit includes a transmission quality analysis module, a data transmission guarantee module, and a cooperative strategy triggering module, which collects the signal reception power between terminals based on edge terminal interaction data. Data packet loss rate and transmission delay By combining signal attenuation gradient prediction technology with the signal reception power between terminals in continuous sampling periods, Calculate the signal attenuation gradient To predict signal change trends, and simultaneously switch between preset communication frequency bands using dynamic frequency band switching technology, and test the signal-to-noise ratio of each frequency band. The final output is divided into "Excellent" ,middle ,Difference Level 3 transmission quality ;
[0032] The data transmission protection module, for abnormal event data, breaks down the data into pre-defined chunks and generates a unique checksum for each chunk. Mark the fragment number Simultaneously, data can be relayed through nearby terminals, and the relay node verifies the checksum after receiving the fragments. When the verification fails, the corresponding sequence number is triggered. Fragmented retransmission requests;
[0033] The collaborative strategy triggering module receives the transmission quality level. ,pass With stability threshold The comparison is used to judge the stability of the communication link. (correspond When the direct transmission strategy is triggered, (correspond , When selecting a neighboring terminal whose transmission quality meets the preset requirements as a relay node, the relay fragmentation strategy is triggered. At the same time, the fragmentation reception status of each terminal is synchronized based on the cross-terminal data interaction protocol. After the target terminal receives all fragments and passes the verification, the data reassembly instruction is triggered to realize cross-terminal data collaboration.
[0034] As a further improvement to this technical solution, the transmission quality analysis process of the transmission quality analysis module includes the following steps:
[0035] S310.1, Signal reception power between acquisition terminals Data packet loss rate and transmission delay Determine the method used to calculate the signal attenuation gradient. The number of consecutive sampling periods;
[0036] S310.2, Inter-terminal signal reception power within a preset number of continuous sampling periods Calculate the statistical value of the power difference between adjacent cycles, and use the statistical value of the power difference between adjacent cycles as the signal attenuation gradient. and based on The positive and negative values and absolute value range can be used to predict the trend of subsequent signal changes;
[0037] S310.3. Switch to each preset communication frequency band using dynamic frequency band switching technology. The test duration for each frequency band is controlled according to preset standards. During the test, interference signals from non-target frequency bands are shielded, and the signal-to-noise ratio of each frequency band is recorded. The statistical values, and the signal-to-noise ratio of each frequency band. The statistical value is used as the effective signal-to-noise ratio of the corresponding frequency band;
[0038] S310.4, combined with signal attenuation gradient Range, effective signal-to-noise ratio of each frequency band Range and transmission delay Scope of transmission quality level :when Within the preset stable range, effective signal-to-noise ratio Within the preset high priority range and transmission delay When within the preset low latency range, it is judged as excellent. ;when Within the preset basic stable range, effective signal-to-noise ratio Within the preset medium range and transmission delay When the delay is within the preset acceptable range, it is judged as medium. All other cases are judged as poor. .
[0039] As a further improvement to this technical solution, the collaborative working process of the data transmission guarantee module and the collaborative strategy triggering module includes the following steps:
[0040] S320.1 The data transmission guarantee module obtains the total amount of abnormal event data, matches and uses a preset segment of the corresponding size according to the preset range in which the total amount of data is located, breaks down the abnormal event data according to the segment size, and generates a unique check code for each segment. And mark the fragment number. ;
[0041] S320.2 The collaborative strategy triggering module takes the current terminal as the center and first filters the transmission quality level within the preset short-range communication range. The nearest terminal; if no suitable terminal is found within this range, the range is expanded to a preset medium-range communication range, and the transmission quality level is selected. The nearest terminals; for the selected relay candidate terminals within the medium-range communication range, combined with The transmission quality level and communication distance are weighted and evaluated, and the relay nodes are assigned priorities (transmission quality level) based on the evaluation results. The weighting of communication distance can be adjusted according to the needs of the scenario.
[0042] S320.3, the collaborative strategy triggering module is based on a cross-terminal data interaction protocol (including terminal identity identifier and fragment sequence number). Verification code (and receiving status field), synchronize the receiving and verification results of the fragments of each terminal, and each terminal provides real-time feedback information through the cross-terminal data interaction protocol;
[0043] S320.4 When the target terminal returns a certain fragment sequence number When the verification fails and the number of failures reaches a preset threshold, the collaborative strategy triggering module switches to the next highest priority relay node, controlling the data transmission guarantee module to re-initiate the segment sequence number. The transmission request continues until the target terminal passes the verification.
[0044] As a further improvement to this technical solution, the intelligent decision-making unit includes a spatiotemporal alignment module and a decision-making module, wherein:
[0045] The spatiotemporal alignment module is used to receive the simplified feature data output by the edge computing processing unit and call the unified time stamp added by the data synchronization module. Spatial coordinates Visual dynamic features Voiceprint characteristics and environmental physical characteristics Mapping to the same spatiotemporal coordinate system generates a multimodal feature set with spatiotemporal consistency. ;
[0046] The decision-making module is based on a real-time environmental baseline. (Generated from historical environmental physical characteristic data), using a dynamic threshold algorithm for... Each modality feature matching has its own specific judgment threshold: targeting visual dynamic features Set abnormal motion threshold Targeting voiceprint features Set the abnormal voiceprint threshold Targeting environmental physical characteristics Set environmental anomaly thresholds Simultaneously, initial decision weights are assigned to each modality feature. The system optimizes the weight allocation based on the correlation between each modal feature and the abnormal event, and finally outputs the abnormal event judgment result by combining the comparison results of each modal feature and the corresponding threshold and the optimized weight.
[0047] As a further improvement to this technical solution, the multi-level response execution unit includes a response triggering module and an action execution module, wherein:
[0048] The response triggering module is used to receive the abnormal event judgment result output by the intelligent decision-making unit, and generate corresponding response instructions according to the severity of the abnormal event. When the abnormal event is judged to be a general abnormality, the basic response combination instruction is triggered; when the abnormal event is judged to be a serious abnormality, the full response combination instruction is triggered.
[0049] The action execution module receives the instruction from the response trigger module and executes the corresponding physical response; when executing lighting linkage, it combines the real-time illumination data of the environmental sensing module in the multimodal perception unit to adjust the brightness of the lighting equipment to a preset range suitable for abnormal scene monitoring.
[0050] The physical response executed by the action execution module also includes:
[0051] When an audible and visual alarm is triggered, the preset audible and visual alarm devices and status indicator lights in the security area are activated, and audible and visual signals of a preset frequency are output.
[0052] When performing access control, a door lock / unlock signal is sent to the access control controller, and the access control operation log is recorded simultaneously;
[0053] When recording video, the high-definition recording mode of the visual acquisition module in the multimodal perception unit is triggered to store video data from abnormal periods.
[0054] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0055] 1. This invention collects visual, acoustic, and environmental physical signals of the security area through a multimodal sensing unit, and adds a unified time stamp and spatial coordinate association to the multi-dimensional signals by relying on the data synchronization module. Then, combined with the collaboration of the feature distillation module and the model optimization module in the edge computing processing unit, the multimodal features are mapped to a unified feature space and high-value features are selected and redundant convolution kernels are identified and removed. This enables effective collaborative processing of multimodal security data and improves the processing efficiency of edge-side sensing data.
[0056] 2. This invention collects parameters such as signal receiving power and packet loss rate through the transmission quality analysis module of the distributed collaborative communication unit and predicts the signal change trend. Combined with the fragmentation verification mechanism of the data transmission guarantee module and the relay node hierarchical screening and failure retransmission logic of the collaborative strategy trigger module, it can ensure the stability of the communication link when multiple security terminals are working together and reduce the risk of data loss during abnormal event data transmission.
[0057] 3. This invention maps the multimodal simplified features output by edge computing to the same spatiotemporal coordinate system through the spatiotemporal alignment module of the intelligent decision unit. Then, with the help of the decision judgment module, the judgment threshold of each modality feature is dynamically adjusted based on the real-time environmental baseline. Combined with the self-learning mechanism to correct the decision weight, the multimodal features can be more in line with the actual security scenario to participate in the anomaly judgment, thereby improving the accuracy of anomaly event judgment.
[0058] 4. This invention collects CPU and FPGA operating status parameters in real time through the computing power scheduling module of the edge computing processing unit, allocates computing tasks based on the complexity of multimodal data operations, and optimizes the computing power consumption of the computing model with the model optimization module. This can avoid the ineffective occupation of computing power resources of edge computing devices and optimize the allocation efficiency of computing power on the edge side. Attached Figure Description
[0059] Figure 1 This is a schematic diagram of the system framework of the present invention;
[0060] The meanings of the labels in the diagram are as follows:
[0061] 100. Multimodal sensing unit; 110. Visual acquisition module; 120. Voiceprint acquisition module; 130. Environmental sensing module; 140. Data synchronization module;
[0062] 200. Edge computing processing unit; 210. Computing power scheduling module; 220. Feature distillation module; 230. Model optimization module;
[0063] 300. Distributed cooperative communication unit; 310. Transmission quality analysis module; 320. Data transmission guarantee module; 330. Cooperative strategy triggering module;
[0064] 400. Intelligent Decision-Making Unit; 410. Spatiotemporal Alignment Module; 420. Decision Judgment Module;
[0065] 500. Multi-level response execution unit; 510. Response triggering module; 520. Action execution module. Detailed Implementation
[0066] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0067] like Figure 1 As shown, this embodiment provides a smart security terminal collaborative monitoring system with edge computing capabilities, including:
[0068] The multimodal sensing unit 100 is used to collect visual, acoustic, and environmental physical signals in the security area to obtain multi-dimensional basic sensing data.
[0069] Understandably, the multimodal sensing unit 100 is suitable for indoor and outdoor security areas such as park entrances and exits, building corridors, and warehouses. During installation, the layout can be adjusted according to the spatial structure of the area: in open areas, the visual acquisition module 110, the voiceprint acquisition module 120, and the environmental sensing module 130 can be integrated into the same security terminal and suspended high; in narrow areas, multiple sets of sub-terminals are arranged at intervals along the path, and the data synchronization module 140 is integrated into the main control board of each sub-terminal to avoid sensing blind spots.
[0070] In this embodiment, the multimodal sensing unit 100 includes a visual acquisition module 110, a voiceprint acquisition module 120, an environmental sensing module 130, and a data synchronization module 140, wherein:
[0071] The visual acquisition module 110 includes a CMOS image sensor and an automatic exposure control circuit, which is used to acquire visual signals of the security area, adapt to a wide range of lighting conditions by adjusting the shutter time, and output standard format video frames carrying optical center coordinates.
[0072] As a further explanation of this embodiment, the selection of CMOS image sensors needs to match the monitoring distance: medium resolution sensors are selected for close-range monitoring (such as inside an elevator), and wide dynamic range sensors are selected for long-range monitoring (such as the perimeter of a park).
[0073] Furthermore, the automatic exposure control circuit is linked with the CMOS sensor to acquire light intensity signals in real time: when the light intensity is higher than the preset upper limit, the shutter speed is shortened to prevent overexposure; when the light intensity is lower than the preset lower limit, the shutter speed is extended and frame accumulation is used to increase brightness. The output video frame format (such as YUV420) is adapted to the feature extraction interface of the edge computing processing unit 200, and the optical center coordinates are determined according to the sensor installation angle and lens parameters (focal length, field of view).
[0074] The voiceprint acquisition module 120 includes a MEMS microphone array and a bandpass filter for acquiring voiceprint signals, enhancing the target sound source through specific frequency band filtering and beamforming algorithms, and outputting an audio stream with a sampling rate adapted to voiceprint feature extraction.
[0075] As a further explanation of this embodiment, the MEMS microphone array is arranged with 4 elements (suitable for quiet corridors) or 8 elements (suitable for noisy entrances and exits) depending on the scenario; the bandpass filter frequency band is set to the main distribution frequency band of abnormal sound patterns (such as glass breakage, metal impact, etc.), filtering low-frequency noise below 200Hz and high-frequency interference above 8000Hz. The beamforming algorithm determines the sound source location by calculating the time difference of the sound pattern signals of each element, adjusts the phase and gain of the element signals to form a directional beam to suppress interference, and outputs an audio stream sampling rate (16kHz or 32kHz) adapted to the subsequent sound pattern feature extraction, without the need for secondary processing.
[0076] The environmental sensing module 130 includes a temperature and humidity sensing component and a smoke sensing component, used to collect environmental physical signals;
[0077] As a further explanation of this embodiment, the temperature and humidity sensing component can be installed at the ventilation point on the side of the terminal, avoiding interference sources such as air conditioning vents and water pipes. Under normal conditions, it collects data once every 30 seconds, and when it exceeds the normal range, it increases to once every 5 seconds. The smoke sensing component adopts the photoelectric principle and is installed below the ceiling. The dust cover is adapted to the terminal structure. When the concentration reaches the warning threshold, it sends a warning signal to the unit main controller.
[0078] The data synchronization module 140 is used to add a unified time stamp to visual, voiceprint and environmental signals, and establish spatial coordinate association based on the installation positions of the visual acquisition module 110, voiceprint acquisition module 120, environmental sensing module 130 and data synchronization module 140, so that the multi-dimensional basic perception data has spatiotemporal matching.
[0079] As a further explanation of this embodiment, the implementation of the unified time stamp of the data synchronization module 140 relies on the real-time clock (RTC) module built into the multimodal sensing unit 100. The RTC module is synchronized with the global clock of the system to ensure that the time base of all sub-modules is consistent. When the data synchronization module 140 receives video frames from the visual acquisition module 110, audio streams from the voiceprint acquisition module 120, and environmental data from the environmental sensing module 130, it reads the current time of the RTC module in real time and adds a timestamp to each data segment. The video frames are timestamped frame by frame, the audio streams are timestamped according to preset time segments (such as every 100ms), and the environmental data is timestamped according to the acquisition time, so as to avoid time stamp deviation caused by data transmission delay.
[0080] Furthermore, the data synchronization module 140 establishes spatial coordinate associations through the following steps:
[0081] During the system installation phase, the three-dimensional coordinates of the visual acquisition module 110, the voiceprint acquisition module 120, and the environmental sensing module 130 in the security area's preset coordinate system (e.g., with the lower left corner of the area as the origin, the horizontal direction as the X-axis, the vertical direction as the Y-axis, and the height direction as the Z-axis) are measured using a laser rangefinder or coordinate calibration tool, and the coordinate data is pre-stored in the storage unit of the data synchronization module 140.
[0082] During data synchronization, the data synchronization module 140 retrieves the three-dimensional coordinates of the corresponding visual acquisition module 110, voiceprint acquisition module 120, and environmental sensing module 130 from the storage unit according to the data source (i.e., which sub-module it comes from), and binds the coordinates to the data itself to form a related data structure of "data content-time stamp-spatial coordinates".
[0083] To ensure spatiotemporal matching, a synchronization verification is performed during the system initialization phase: a test signal is triggered at a known coordinate point in the security area (such as a pre-marked test point) (e.g., playing a standard test tone or placing a test object). After receiving the test data collected by each module, the data synchronization module 140 checks whether the timestamp deviation of the visual acquisition module 110, the voiceprint acquisition module 120, and the environmental sensing module 130 for the same test event is within the preset allowable range (e.g., ≤10ms), and whether the spatial coordinates match the coordinates of the test point. If the deviation exceeds the range, the RTC module is automatically calibrated or the coordinate parameters of the visual acquisition module 110, the voiceprint acquisition module 120, and the environmental sensing module 130 are corrected until the spatiotemporal matching requirements are met.
[0084] The edge computing processing unit 200 is based on multimodal sensing data. It uses dynamic feature distillation technology and hardware load feedback technology to allocate CPU and FPGA computing tasks, uses incremental convolution kernel pruning technology to optimize the model computing process, judges the computing power requirements and model redundancy of the sensing data processing task, and outputs simplified feature data to improve the efficiency of sensing data processing.
[0085] In this embodiment, the edge computing processing unit 200 includes a computing power scheduling module 210, a feature distillation module 220, and a model optimization module 230, wherein:
[0086] The computing power scheduling module 210 collects the operating status parameters of the CPU and FPGA in real time, combines the computational complexity analysis of multimodal data, and completes the targeted allocation of computing tasks between the CPU and FPGA based on hardware load feedback technology.
[0087] The feature distillation module 220 is used to map the original features of visual, voiceprint and environmental multimodal features to a unified feature space, and extract high-value features and remove redundant feature dimensions through dynamic feature distillation technology;
[0088] The model optimization module 230 is used to analyze the feature contribution of convolution kernels in the computation model to identify redundant convolution kernels, remove redundant parts through incremental convolution kernel pruning technology, and optimize the computing power consumption of the model operation.
[0089] In this embodiment, the targeted allocation of computing tasks by the computing power scheduling module 210 includes the following steps:
[0090] S210.1, Real-time CPU Utilization Monitoring FPGA resource utilization and memory usage This forms a hardware load status dataset;
[0091] S210.2, Synchronously acquire visual frame size parameters Voiceprint sampling scale Number of environmental data points and the dimension of features to be extracted Determine the baseline for computational complexity of multimodal data;
[0092] S210.3. Based on hardware load feedback technology, the computational complexity benchmark is matched and analyzed with the hardware load status dataset. When the matching result exceeds the parallel processing threshold... When the feature extraction task is assigned to the FPGA, and the matching result does not exceed the parallel processing threshold, the task is performed accordingly. At that time, logical judgment tasks are assigned to the CPU.
[0093] As a further explanation of this embodiment, the CPU utilization rate in this embodiment... Read via the operating system kernel interface, with a value range of 0-100%; FPGA resource utilization. The FPGA configuration interface reads the resource usage ratio of DSP, BRAM, etc., with a value range of 0-100%.
[0094] As a further explanation of this embodiment, the visual frame size parameters are acquired synchronously. (Determined directly by the video frame resolution output by the multimodal sensing unit 100), speaker sampling scale (Calculated from the sampling rate and sampling duration of the voiceprint acquisition module 120), number of environmental data points (Calculated from the acquisition frequency and acquisition duration of the environmental sensing module 130) and the dimension of the features to be extracted. (Determined based on the multimodal feature dimension requirements commonly used in security scenarios, such as 256 dimensions for visual features and 40 dimensions for voiceprint features), and the computational complexity of each modality's data is quantified based on hardware computing power:
[0095] Visual data computational complexity ,in: The video frame rate (determined by the output frame rate of the visual acquisition module 110). Calculated based on the CPU's ability to process visual pixels per cycle (refer to the "Single-cycle pixel processing capacity" parameter in the CPU chip manual, for example, the CPU's single-cycle processing capacity). pixels, then );
[0096] Voiceprint data computation complexity ,in Calculated based on the CPU's ability to process voiceprint sampling points per cycle;
[0097] Environmental data computational complexity ,in The environmental data acquisition frequency (determined by the acquisition frequency of the environmental sensing module 130). Calculated based on the CPU's ability to process environmental data points per single cycle;
[0098] Total computational complexity , which serves as a benchmark for computational complexity.
[0099] Furthermore, the parallel processing threshold in this embodiment The value is set to 60% of the theoretical hardware computing power (the theoretical hardware computing power refers to the peak computing power indicated in the CPU and FPGA chip manuals); and the computational complexity benchmark is matched and analyzed with the hardware load state dataset.
[0100] like Feature extraction tasks (such as visual frame feature and voiceprint feature extraction) are assigned to FPGAs based on the "parallel computing efficiency advantage" specified in the FPGA chip manual (FPGAs are more efficient than CPUs in processing matrix operation tasks).
[0101] like : Logical judgment tasks (such as preliminary judgment of environmental data anomalies) are assigned to the CPU, based on the fact that the CPU's processing latency for serial logic operations is lower than that of the FPGA;
[0102] After task allocation, the computing power scheduling module 210 continuously monitors the hardware load. This triggers the "task queue buffer + dynamic downsampling" mechanism (the video frame rate is reduced to the lowest effective frame rate supported by the visual acquisition module 110, and the audioprint sampling rate is reduced to the lowest effective sampling rate supported by the audioprint acquisition module 120) to avoid hardware overload (the threshold setting refers to the "safe operating load limit" marked in the CPU and FPGA chip manual).
[0103] In this embodiment, the collaborative processing of the feature distillation module 220 and the model optimization module 230 includes the following steps:
[0104] S220.1, Feature distillation module 220 receives multimodal basic perception data and converts visual features... (Extracted using publicly available lightweight CNN models such as MobileNet) Voiceprint features (Extracted using the publicly available MFCC algorithm, with algorithm parameters referencing industry standards for speech signal processing), Environmental characteristics (Feature vectors composed directly of data such as temperature, humidity, and smoke concentration) are mapped to a unified feature space to generate a fused feature matrix. ;
[0105] In this step, , : ; Original video features Pre-trained transformation matrix The unified dimensional feature representation obtained after linear transformation; transformation matrix Based on publicly available smart security anomaly datasets (such as the UCF-Crime dataset, launched in 2018 by a research team at the University of Central Florida (UCF), which focuses on anomaly detection, this dataset provides synchronized video frames, ambient audio streams, and anomaly event labels, and supports...) and The correlation learning is obtained through cross-modal contrastive learning training, and the training process follows the publicly available contrastive learning algorithm flow;
[0106] , : The training dataset and training process are the same. ; Audio features After transformation matrix The feature representation obtained after linear transformation;
[0107] , : The training dataset and training process are the same. ; Audio features After transformation matrix The feature representation obtained after linear transformation;
[0108] Finally, the fused feature matrix is obtained. .
[0109] Calculation using dynamic characteristic distillation technology The correlation between various characteristics and security anomalies Retain features with a correlation degree no lower than the feature importance threshold. Feature subset ;
[0110] Model optimization module 230 processes The computational model is used to evaluate convolution kernels and calculate the output characteristic fluctuation value of each convolution kernel. and gradient contribution The marker simultaneously satisfies Fluctuation threshold and Contribution threshold Redundant convolution kernels ;
[0111] Remove Post-validation model output error ,when Error allowable threshold At the same time, maintain the simplified model structure and output based on Simplified feature data.
[0112] As a further explanation of this embodiment, in this step, the correlation between each feature in F and the security anomaly event is calculated using dynamic feature distillation technology. ( (For feature indexes), the cosine similarity formula is used to calculate the relevance:
[0113] ;
[0114] in, , The height and width of the output feature map of the convolutional layer are determined by the structure of the computational model; refer to the convolutional layer parameters of the open-source MobileNet model. for The feature map after convolution operation is in The pixel value of the location, The mean of the feature map; set the fluctuation threshold. This threshold is determined based on the minimum effective fluctuation range of the convolutional layer output features (when...). (At times, the feature fluctuations are too small to distinguish between normal and abnormal scenarios).
[0115] right The convolution kernel is then used to further calculate its gradient contribution. The formula is: ;in, The number of parameters for the convolution kernel (determined by the kernel size, such as the number of parameters for a 3×3 convolution kernel). ), The cross-entropy loss function is used for anomaly detection (following the publicly available definition of loss functions for classification tasks). For the convolution kernel One parameter; setting the contribution threshold. This threshold is determined based on industry-standard model pruning criteria (when...). At that time, the contribution of the convolution kernel to the model output is lower than the industry minimum effective contribution), and the label simultaneously satisfies Redundant convolution kernels ;
[0116] In removing redundant convolution kernels Then, verify the model output error. (Calculated by comparing the anomaly detection results output by the model with the labeled results of the public dataset), and an error tolerance threshold is set. This threshold is determined based on the anomaly detection accuracy requirements for smart security scenarios (the industry standard requirement is anomaly detection accuracy ≥ 95%, corresponding to an error of...). );like Maintain the simplified model structure and output based on Simplified feature data; if The gradient contribution of the backtracking retained portion is close to The convolution kernel is re-validated until the error requirement is met.
[0117] It should be added that, in this embodiment, the computing power scheduling module 210 completes hardware load monitoring and computing task allocation based on the parameters in the CPU and FPGA chip manuals, and the feature distillation module 220, based on the allocated computing power resources, uses a transformation matrix trained on a public dataset to perform feature mapping and filtering on multimodal data to generate a high-value feature subset. Subsequently, the model optimization module 230, based on industry pruning standards and anomaly detection accuracy requirements, optimized the processing... The computational model undergoes redundant convolution kernel pruning and optimization to reduce computational power consumption. Finally, the optimized model outputs simplified feature data, which is transmitted to the subsequent distributed collaborative communication unit 300 through an internal interface that conforms to the chip interconnect standard, providing lightweight and accurate feature support for anomaly decision-making in the smart security system.
[0118] The distributed collaborative communication unit 300 analyzes wireless transmission quality based on edge terminal interaction data, uses signal attenuation gradient prediction technology and frequency band dynamic switching technology, uses fragmented verification relay technology to ensure abnormal event data transmission, judges the stability of communication links, triggers the optimal transmission strategy, and realizes cross-terminal data collaboration.
[0119] In this embodiment, the distributed cooperative communication unit 300 includes a transmission quality analysis module 310, a data transmission guarantee module 320, and a cooperative strategy triggering module 330, wherein:
[0120] The transmission quality analysis module 310 collects the signal reception power between terminals based on the interaction data of the edge terminals. Data packet loss rate and transmission delay By combining signal attenuation gradient prediction technology with the signal reception power between terminals in continuous sampling periods, Calculate the signal attenuation gradient To predict signal change trends, and simultaneously switch between preset communication frequency bands using dynamic frequency band switching technology, and test the signal-to-noise ratio of each frequency band. The final output is divided into "Excellent" ,middle ,Difference Level 3 transmission quality ;
[0121] In this embodiment, the transmission quality analysis process of the transmission quality analysis module 310 includes the following steps:
[0122] S310.1, Signal reception power between acquisition terminals Data packet loss rate and transmission delay Determine the method used to calculate the signal attenuation gradient. The number of consecutive sampling periods;
[0123] As a further explanation of this embodiment, the core parameter acquisition in the transmission quality analysis module 310 of this embodiment specifically includes: acquiring the signal reception power between terminals. (Acquired via the RF chip's RSSI interface, with the acquisition frequency synchronized with the chip's signal detection period), data packet loss rate. (Calculate the ratio of "number of unresponded packets / total number of packets sent" by periodically sending 10 ICMP echo request packets every 5 seconds), transmission delay (After synchronizing the terminal time based on the NTP protocol, through) calculate, Send timestamps for data frames. (received timestamp); and simultaneously determine the data used to calculate the signal attenuation gradient. The number of continuous sampling periods is 3 (each period is 1 second to avoid random errors in single-period sampling).
[0124] S310.2, Inter-terminal signal reception power within a preset number of continuous sampling periods Calculate the statistical value of the power difference between adjacent cycles, and use the statistical value of the power difference between adjacent cycles as the signal attenuation gradient. and based on The positive and negative values and absolute value range can be used to predict the trend of subsequent signal changes;
[0125] Furthermore, the signal attenuation gradient calculation and trend prediction of the transmission quality analysis module 310 in this embodiment includes the following steps:
[0126] First, obtain the signal received power for three consecutive sampling periods. ( This is the power for the first cycle. This is the power for the second cycle. (Power of the 3rd cycle);
[0127] Subsequently, the power difference between adjacent cycles was calculated. ;
[0128] Then, through the formula Calculate the signal attenuation gradient ;
[0129] Next, based on Predicting trends based on numerical range: The time signal trend is stable. The trend of the time signal is strengthening. The signal trend decays over time;
[0130] Finally, the trend prediction results will be used as a reference for classifying transmission quality levels.
[0131] S310.3. Switch to each preset communication frequency band using dynamic frequency band switching technology. The test duration for each frequency band is controlled according to preset standards. During the test, interference signals from non-target frequency bands are shielded, and the signal-to-noise ratio of each frequency band is recorded. The statistical values, and the signal-to-noise ratio of each frequency band. The statistical value is used as the effective signal-to-noise ratio of the corresponding frequency band;
[0132] As a further explanation of this embodiment, the frequency band dynamic switching and signal-to-noise ratio test in the transmission quality analysis module 310 of this embodiment specifically include: preset communication frequency bands based on the selection of legal frequency bands supported by hardware; LoRa RF chip presets 433MHz and 868MHz frequency bands; Wi-Fi RF chip presets 2.4GHz and 5GHz frequency bands; the test duration for each frequency band is set to 5 seconds; before the test, non-target frequency band interference is shielded by microcontroller instructions; and the signal-to-noise ratio of each frequency band is collected by the built-in module of the RF chip. The average value of data collected 10 times for each frequency band is used as the effective signal-to-noise ratio for that frequency band.
[0133] S310.4, combined with signal attenuation gradient Range, effective signal-to-noise ratio of each frequency band Range and transmission delay Scope of transmission quality level :when Within the preset stable range, effective signal-to-noise ratio Within the preset high priority range and transmission delay When within the preset low latency range, it is judged as excellent. ;when Within the preset basic stable range, effective signal-to-noise ratio Within the preset medium range and transmission delay When the delay is within the preset acceptable range, it is judged as medium. All other cases are judged as poor. .
[0134] As a further explanation of this embodiment, the transmission quality level classification of the transmission quality analysis module 310 in this embodiment includes the following steps:
[0135] First, determine the intervals for each parameter: signal attenuation gradient. The stable interval is The basic stable range is Effective signal-to-noise ratio The high priority interval is The middle range is Transmission delay The low latency range is Acceptable delay range is ;
[0136] Then, determine whether the parameter combination meets the optimal requirements. : In the stable range, In the high-optimal range and If the condition is met in the low-latency range, it is considered as... ;
[0137] Then, determine whether it meets the requirements. : Within the basically stable range, In the middle range and If the acceptable delay range is met, it is considered as... ;
[0138] Next, all other parameter combinations were determined to be differences. ;
[0139] Finally, output the determined transmission quality level. .
[0140] For abnormal event data, the data transmission protection module 320 breaks down the data into pre-defined chunk sizes and generates a unique checksum for each chunk. Mark the fragment number Simultaneously, data can be relayed through nearby terminals, and the relay node verifies the checksum after receiving the fragments. When the verification fails, the corresponding sequence number is triggered. Fragmented retransmission requests;
[0141] In this embodiment, the collaborative working process between the data transmission guarantee module 320 and the collaborative strategy triggering module 330 includes the following steps:
[0142] S320.1 The data transmission guarantee module 320 obtains the total amount of abnormal event data, matches and uses a preset segment of the corresponding size according to the preset range in which the total amount of data is located, breaks down the abnormal event data according to the segment size, and generates a unique check code for each segment. And mark the fragment number. ;
[0143] As a further explanation of this embodiment, the data fragmentation and identifier generation in the data transmission guarantee module 320 of this embodiment specifically includes: setting the preset fragment size based on the communication protocol MTU value; for LoRa communication, "no fragmentation for total data ≤ 200 bytes, fragment size of 200 bytes for 200 bytes < total data ≤ 1000 bytes, and fragment size of 500 bytes for total data > 1000 bytes"; and generating a unique checksum for each fragment. (Using the CRC32 algorithm, calculated by the microcontroller's built-in CRC module, outputting a 4-byte checksum) and fragment sequence number. (A 16-bit unsigned integer, incrementing from 1).
[0144] S320.2, the collaborative strategy triggering module 330, taking the current terminal as the center, first filters the transmission quality level within the preset short-range communication range. The nearest terminal; if no suitable terminal is found within this range, the range is expanded to a preset medium-range communication range, and the transmission quality level is selected. The nearest terminals; for the selected relay candidate terminals within the medium-range communication range, combined with The transmission quality level and communication distance are weighted and evaluated, and the relay nodes are assigned priorities (transmission quality level) based on the evaluation results. The weighting of communication distance can be adjusted according to the needs of the scenario.
[0145] S320.3, the collaborative strategy triggering module 330 is based on a cross-terminal data interaction protocol (including terminal identity identifier and fragment sequence number). Verification code (and receiving status field), synchronize the receiving and verification results of the fragments of each terminal, and each terminal provides real-time feedback information through the cross-terminal data interaction protocol;
[0146] S320.4 When the target terminal returns a certain fragment sequence number When the verification fails and the number of failures reaches a preset threshold, the collaborative strategy triggering module 330 switches to the next highest priority relay node and controls the data transmission guarantee module 320 to re-initiate the segment sequence number. The transmission request continues until the target terminal passes the verification.
[0147] As a further explanation of this embodiment, the relay node data verification of the data transmission guarantee module 320 in this embodiment includes the following steps:
[0148] First, after receiving the fragmented data, the relay node extracts the checksum from the fragment header. With fragment number ;
[0149] Subsequently, the CRC32 value was recalculated for the fragmented data;
[0150] Then, compare the recalculated CRC32 value with the extracted one. If they match, the verification is considered successful, and the record is made. The corresponding "verification passed" status is then forwarded to the fragment;
[0151] Next, if there is a discrepancy, the verification is deemed to have failed, and a "fragmentation" message is immediately sent back to the sender. The "verification failed" request (including terminal ID and fragment) );
[0152] Finally, wait for the sender to issue a retransmission command or for subsequent fragments.
[0153] Cooperative strategy triggering module 330 receives transmission quality level ,pass With stability threshold The comparison is used to judge the stability of the communication link. (correspond When the direct transmission strategy is triggered, (correspond , When selecting a neighboring terminal whose transmission quality meets the preset requirements as a relay node, the relay fragmentation strategy is triggered. At the same time, the fragmentation reception status of each terminal is synchronized based on the cross-terminal data interaction protocol. After the target terminal receives all fragments and passes the verification, the data reassembly instruction is triggered to realize cross-terminal data collaboration.
[0154] As a further explanation of this embodiment, the transmission strategy determination in the cooperative strategy triggering module 330 of this embodiment specifically includes: setting a stability threshold. (i.e., "excellent" level), when the transmission quality level (Right now When the communication link is stable, the direct transmission strategy is triggered (the current terminal directly sends data to the target terminal); when (Right now or When the link is deemed unstable, a relay fragmentation strategy is triggered (data is forwarded through a nearby terminal).
[0155] Furthermore, the relay node screening and priority evaluation of the coordination strategy triggering module 330 in this embodiment includes the following steps:
[0156] First, set the preset communication range: short-range communication range ≤ 50 meters, medium-range communication range 50~200 meters;
[0157] Subsequently, centered on the current terminal, a "Terminal ID + Transmission Quality Level Request" signal is broadcast, and Q-level feedback from nearby terminals is received. The system then filters data within the immediate vicinity. The terminal;
[0158] Then, if no suitable terminal is found at close range, the search is expanded to a mid-range range. Terminals (excluding) (terminal).
[0159] Next, a weighted evaluation is performed on the candidate mid-range relay terminals. The weighted scoring formula is as follows: ;in, for Rank weight, For communication distance weights, It can be adjusted according to the scenario; for The grade quantification value is set to 0.8; This is a quantified value for communication distance; the closer the distance, the larger the value, which is converted from RSSI value.
[0160] Finally, according to the weighted score Sort the relay nodes from high to low priority and assign them priority.
[0161] As a further explanation of this embodiment, the cross-terminal state synchronization and failure retransmission in the collaborative strategy triggering module 330 of this embodiment specifically includes: a cross-terminal data interaction protocol containing a terminal identity identifier (8 bytes) and a fragment sequence number. (2 bytes), checksum (4 bytes), receive status (1 byte, 0 = not received, 1 = received, 2 = verification failed) field, the protocol frame is transmitted through the UART interface and follows the "send-ACK confirmation" mechanism; and the threshold for the number of fragment verification failures is set to 3. When the target terminal reports a certain fragment When the "verification failed" status (receive status = 2) and the number of failures reaches the threshold, the system switches to the second-highest priority relay node, triggering the data transmission guarantee module 320 to retransmit the fragment. .
[0162] Furthermore, the collaborative working process (data reassembly stage) between the collaborative strategy triggering module 330 and the data transmission guarantee module 320 in this embodiment includes the following steps:
[0163] First, the target terminal counts the sequence numbers of the received fragments in real time. And compare "received fragments" The maximum value and the total number of fragments (the total number of fragments is provided in advance by the data transmission guarantee module 320);
[0164] Subsequently, when the target terminal confirms that it has received all fragments and that all fragments have passed verification (reception status = 1), it sends a "all fragments are ready" signal to the coordination strategy triggering module 330.
[0165] Then, after receiving the signal, the collaborative strategy triggering module 330 sends a data reassembly instruction to the target terminal;
[0166] Next, the target terminal will be sorted by fragment number. Sort by size from smallest to largest and merge all partitions;
[0167] Finally, the target terminal recalculates the CRC32 value of the merged overall data, compares it with the check value of the original data, and completes cross-terminal data collaboration after confirming the integrity.
[0168] The intelligent decision-making unit 400, based on the simplified feature data output by the edge computing processing unit 200, spatiotemporally aligns visual dynamic features, voiceprint features and environmental physical features, and adjusts the judgment criteria of each feature through a dynamic threshold algorithm based on the real-time environmental baseline, optimizes the decision weight of each modality feature, and achieves accurate judgment of abnormal events.
[0169] In this embodiment, the intelligent decision-making unit 400 includes a spatiotemporal alignment module 410 and a decision determination module 420, wherein:
[0170] The spatiotemporal alignment module 410 is used to receive the simplified feature data output by the edge computing processing unit 200 and call the unified time stamp added by the data synchronization module 140. Spatial coordinates Visual dynamic features Voiceprint characteristics and environmental physical characteristics Mapping to the same spatiotemporal coordinate system generates a multimodal feature set with spatiotemporal consistency. ;
[0171] As a further explanation of this embodiment, the "calling the time stamp and spatial coordinates of the data synchronization module 140" in the spatiotemporal alignment module 410 of this embodiment specifically includes:
[0172] Using a pre-defined communication protocol between units (following the interaction protocol between the multimodal sensing unit 100 and the edge computing processing unit 200, including data identifiers, time fields, and coordinate fields), the unified time stamp pre-added by the data synchronization module 140 is extracted from the simplified feature data forwarded by the edge computing processing unit 200. Spatial coordinates (Three-dimensional coordinates (X, Y, Z), using the preset coordinate system of the multimodal sensing unit 100: with the lower left corner of the security area as the origin, the X-axis is horizontal, the Y-axis is vertical, and the Z-axis is vertical); among which, visual dynamic features Related The installation coordinates of the visual acquisition module 110, and the voiceprint features. Related The installation coordinates and environmental physical characteristics of the voiceprint acquisition module 120. Related The coordinates are for the installation of the environmental sensing module 130.
[0173] Furthermore, the spatiotemporal mapping of multimodal features in the spatiotemporal alignment module 410 of this embodiment includes the following steps:
[0174] First, the time window precision for spatiotemporal alignment is set to 100ms (consistent with the time stamp precision of the data synchronization module 140 to avoid time deviation), and all simplified feature data are time-stamped. Categorization, within the same 100ms time window , , Divide into a set of features to be aligned;
[0175] Subsequently, a "feature-spatial coordinate" mapping relationship is established: visual dynamic features The optical center coordinates included (from the vision acquisition module 110) are related to the corresponding Perform coordinate calibration (using a preset coordinate system transformation formula to convert the optical center pixel coordinates into three-dimensional spatial coordinates to ensure consistency with the target coordinate system). (The coordinate system is consistent) voiceprint features Environmental physical characteristics Directly bind to their respective associations ;
[0176] Then, for multimodal features within the same time window, a unified time label (taking the start timestamp of the window) and spatial association identifier (recording the P and relative distance corresponding to each feature, such as...) are applied. and (X-axis distance difference);
[0177] Next, remove data sets that are missing any modal feature within the time window (to avoid judgment bias caused by single modal features).
[0178] Finally, the feature combinations that complete the temporal classification and spatial binding are encapsulated into a multimodal feature set with spatiotemporal consistency. , The structure includes "time tag + spatial association identifier + The data is transmitted to the decision-making module 420 via the SPI interface.
[0179] Decision-making module 420 is based on real-time environmental baseline (Generated from historical environmental physical characteristic data), using a dynamic threshold algorithm for... Each modality feature matching has its own specific judgment threshold: targeting visual dynamic features Set abnormal motion threshold Targeting voiceprint features Set the abnormal voiceprint threshold Targeting environmental physical characteristics Set environmental anomaly thresholds Simultaneously, initial decision weights are assigned to each modality feature. The system optimizes the weight allocation based on the correlation between each modal feature and the abnormal event, and finally outputs the abnormal event judgment result by combining the comparison results of each modal feature and the corresponding threshold and the optimized weight.
[0180] As a further explanation of this embodiment, the real-time environmental baseline in the decision-making module 420 of this embodiment... The generation specifically includes:
[0181] Extract normal environmental physical characteristic data (i.e., periods without triggered abnormal alarms) from the historical data stored in the edge computing processing unit 200 within the last 7 days (the industry-standard environmental baseline statistical period to avoid the impact of short-term data fluctuations). (Including parameters such as temperature, humidity, and smoke concentration).
[0182] Calculate statistical values for each type of environmental parameter: calculate the mean value for temperature and humidity parameters. with standard deviation Calculation of the average value of smoke concentration parameters with standard deviation ;
[0183] The mean ± 1 standard deviation of each parameter is taken as the normal fluctuation range of that parameter. The normal fluctuation ranges of all parameters are combined to form the real-time environmental baseline. ,Right now Baseline It is automatically updated every day at midnight (recalculated based on the latest data from the previous 7 days) to ensure that it matches the real-time environment.
[0184] As a further explanation of this embodiment, the dynamic threshold algorithm in the decision-making module 420 of this embodiment specifically includes: for different modal characteristics, using the real-time environmental baseline... Alternatively, based on the normal feature range, a specific judgment threshold can be set in combination with the feature deviation requirements of abnormal scenarios:
[0185] Visual dynamic features abnormal motion threshold Extracting from historical normal scenes Calculate the maximum value of the motion characteristic values (such as target speed, percentage of motion area). ,set up (Allowing reasonable fluctuations in normal movement to avoid misjudgment), when The motion feature values in the middle exceed At that time, visual abnormalities were determined;
[0186] Voiceprint characteristics abnormal voiceprint threshold Extracting from historical normal scenes The sound pressure level characteristics, calculate the mean with standard deviation ,set up (The maximum fluctuation of normal ambient sound) When The sound pressure level exceeds If the voiceprint frequency falls within a preset abnormal frequency band (such as 1000-2000Hz for the sound of breaking glass), the voiceprint characteristics are determined to be abnormal.
[0187] Environmental physical characteristics Environmental anomaly threshold Based on real-time environmental baseline Setting temperature and humidity parameters Smoke concentration parameters (The smoke concentration should normally be close to 0; the upper limit is relaxed here to avoid misjudgment.) When Any parameter exceeding the corresponding At that time, the environmental characteristics were determined to be abnormal.
[0188] Furthermore, the decision weight setting and optimization in the decision determination module 420 of this embodiment specifically includes:
[0189] Initial decision weights Based on the contribution of each modality to historical anomaly events, and referencing common anomaly types in smart security scenarios (e.g., visual input is highest in intrusion events, and environmental input is highest in fire events), the initial weights are set to... (The weights sum to 1, which conforms to the probability allocation logic);
[0190] Weight optimization is based on calculating the correlation between each modal feature and the abnormal event. correlation Using cosine similarity algorithm: ,in , Typical characteristics of this modality in historical anomalous events are extracted from the anomalous data sample library of the edge computing processing unit 200. The value ranges from 0 to 1, with larger values indicating higher correlation.
[0191] Optimized weight calculation: The weights are adjusted using a normalization formula, resulting in optimized weights. ,make sure .
[0192] The "abnormal event determination" of the decision-making module 420 in this embodiment includes the following steps:
[0193] First, receive the output from the spatiotemporal alignment module 410. Extract them separately , , ;
[0194] Subsequently, the modal features are compared with the corresponding dynamic thresholds to generate single-modal judgment results. ( Indicates an anomaly. (Indicates normal) ,otherwise ; Exceeding but ,otherwise ; Exceeding but ,otherwise ;
[0195] Then, the comprehensive judgment score is calculated by combining the optimized weights. The formula is ;
[0196] Next, set the anomaly detection threshold. (Based on historical abnormal data testing, to ensure normal scenarios) Abnormal scenarios );
[0197] Finally, comparison and :like Output the "abnormal" judgment result, along with a single-modal abnormality label (e.g., "visual + environmental abnormality"); if The system outputs a "normal" result, which is then transmitted to the subsequent multi-level response execution unit 500 via the UART interface.
[0198] The multi-level response execution unit 500 triggers physical response actions such as audible and visual alarms, access control, video recording, and lighting linkage based on the anomaly determination result.
[0199] In this embodiment, the multi-level response execution unit 500 includes a response triggering module 510 and an action execution module 520, wherein:
[0200] The response triggering module 510 is used to receive the abnormal event judgment result output by the intelligent decision unit 400, and generate corresponding response instructions according to the severity of the abnormal event. When the abnormal event is judged to be a general abnormality, the basic response combination instruction is triggered; when the abnormal event is judged to be a serious abnormality, the full response combination instruction is triggered.
[0201] As a further explanation of this embodiment, the response triggering module 510 of this embodiment receives the anomaly judgment result (including "normal / abnormal" label and comprehensive judgment score) output by the intelligent decision-making unit 400. (Single-modal anomaly marker), based on a comprehensive judgment score. The severity is classified based on the following criteria:
[0202] General abnormalities: (Corresponding to single-modal anomalies or mild multimodal anomalies, such as only the ambient temperature and humidity exceeding the threshold, or short-term low-amplitude voiceprint anomalies).
[0203] Serious abnormality: (This corresponds to multimodal anomalies or high-risk single-modal anomalies, such as visual detection of intrusion + voiceprint detection of breaking sounds, or severely excessive smoke concentration).
[0204] Meanwhile, the severity classification thresholds are set based on risk level standards for smart security scenarios, avoiding over-response or under-response.
[0205] As a further explanation of this embodiment, the "response instruction generation" in the response triggering module 510 of this embodiment specifically includes:
[0206] The response command adopts a binary frame structure, containing "abnormal severity code (1 byte, 01=general abnormality, 10=severe abnormality), action control bits (4 bytes, each byte corresponds to 1 action: byte 1=audio-visual alarm, byte 2=access control, byte 3=video recording, byte 4=lighting linkage, 1=trigger, 0=no trigger), check bit (1 byte, CRC8 check)";
[0207] Basic response combination command (general anomaly): Set the action control bit to "1,0,1,0", which triggers the sound and light alarm and video recording, but does not trigger access control and lighting linkage (to reduce interference with normal scenes).
[0208] Full response combined command (serious anomaly): Set the action control bit to "1,1,1,1", which triggers audible and visual alarms, access control, video recording, and lighting linkage (to comprehensively address high-risk scenarios).
[0209] After the instruction is generated, it is transmitted to the action execution module 520 through the internal SPI interface, and the instruction generation time is recorded at the same time (the format is consistent with the time stamp of the data synchronization module 140).
[0210] The "command verification and retransmission" of the response triggering module 510 in this embodiment includes the following steps:
[0211] First, after receiving the response instruction, the action execution module 520 will send a "instruction reception confirmation" signal (including instruction verification result) back to the response trigger module 510.
[0212] Subsequently, the response trigger module 510 judges the verification result in the confirmation signal: if the verification passes (CRC8 value matches), it marks "instruction issued";
[0213] Then, if the verification fails or no acknowledgment signal is received (the timeout is set to 500ms), the response command is regenerated and resent, with the number of resentments not exceeding 3 (to avoid infinite retransmission occupying bus resources).
[0214] Next, if the resending still fails after 3 attempts, a "response command issuance failed" signal will be sent to the intelligent decision unit 400 to indicate that the abnormal processing was interrupted.
[0215] Finally, if the instruction is successfully issued, the system continues to receive feedback on the action execution status from the action execution module 520 until all actions are completed.
[0216] After receiving the instruction from the response trigger module 510, the action execution module 520 executes the corresponding physical response; among them, when executing the lighting linkage, it combines the real-time illumination data of the environment sensing module 130 in the multimodal sensing unit 100 to adjust the brightness of the lighting equipment to a preset range suitable for abnormal scene monitoring.
[0217] As a further explanation of this embodiment, the action execution module 520 also includes executing the corresponding physical response:
[0218] When an audible and visual alarm is triggered, the preset audible and visual alarm devices and status indicator lights in the security area are activated, and audible and visual signals of a preset frequency are output.
[0219] When performing access control, a door lock / unlock signal is sent to the access control controller, and the access control operation log is recorded simultaneously;
[0220] When video recording is performed, the high-definition recording mode of the visual acquisition module 110 in the multimodal perception unit 100 is triggered to store video data during abnormal periods.
[0221] Furthermore, the audible and visual alarm can be an industrial-grade LTE-1101J integrated audible and visual alarm (suitable for security scenarios, supporting 220V AC power supply). The action execution module 520 controls the power supply of the alarm via a relay. When the audible and visual alarm is triggered, the relay engages to start the alarm, outputting an audible and visual signal at a preset frequency (e.g., sound frequency 2Hz, light flashing red at a frequency of 2Hz, meeting the visibility requirements of security alarms). Simultaneously, it links the status indicator lights of the security area (e.g., LED indicators in corridors and entrances), which switch to solid yellow (distinguishing it from a red alarm, indicating "abnormal processing"). The audible and visual alarm continues until the abnormality is resolved (the intelligent decision unit 400 outputs a "normal" judgment result), or can be manually stopped via the security control console.
[0222] The specific implementation of access control includes: the access controller is an industrial-grade controller supporting RS485 communication (such as the K200 access control motherboard), and the action execution module 520 communicates with the access controller via RS485 bus, following the MODBUS protocol commonly used in the access control industry; when access control is triggered, corresponding signals are sent according to the abnormal scenario requirements: if the abnormality is "area intrusion", a "door lock" signal is sent (controlling the electromagnetic lock to de-energize and lock, preventing personnel from entering or exiting); if the abnormality is "personnel trapped", a "door lock unlock" signal is sent (controlling the electromagnetic lock to energize and unlock, facilitating rescue); at the same time, the access control operation log is automatically recorded, and the log content includes "operation time (consistent with the time stamp of the data synchronization module 140), operation type (lock / unlock), trigger reason (abnormality identifier, such as "visual intrusion abnormality"), access control number", the log is stored in the local Flash of the access controller, and is also synchronously uploaded to the edge computing processing unit 200 for backup;
[0223] The video recording process specifically includes: sending a "high-definition recording mode" command to the visual acquisition module 110 of the multimodal perception unit 100 via the UART interface. The command includes "recording start time, estimated recording duration (default 5 minutes, which can be dynamically adjusted based on abnormal duration), resolution parameters (1080P), and frame rate parameters (30fps)". Upon receiving the command, the visual acquisition module 110 immediately switches to high-definition mode and records video data during abnormal periods. The video data storage path is divided into two levels: the first level is stored on the local SD card of the visual acquisition module 110, and the second level is synchronously backed up to the storage module of the edge computing processing unit 200 via the distributed collaborative communication unit 300 (to avoid data loss due to damage to the local SD card). The recording ends when either the estimated recording duration is reached or the intelligent decision unit 400 outputs a "normal" judgment result, with the first condition being met taking precedence.
[0224] Furthermore, as a further explanation of this embodiment, the execution of lighting linkage in the action execution module 520 of this embodiment specifically includes: selecting LED lights that support PWM dimming (adapting to edge-side voltage, such as 12V DC power supply) as lighting equipment; the action execution module 520 controlling the brightness of the lights through the PWM dimming module; before executing the linkage, real-time illumination data is first obtained from the environmental sensing module 130 of the multimodal sensing unit 100 through the I2C interface, and then the brightness is adjusted to a preset range adapted to abnormal scene monitoring based on the illumination data;
[0225] If the real-time illumination data is <50 lux (low-light environment, such as at night or in a basement): adjust the illumination brightness to 80% (preset range 70%-90%, to ensure that the visual acquisition module 110 can clearly capture details).
[0226] If 50 lux ≤ real-time lighting data ≤ 200 lux (medium lighting environment, such as cloudy day or evening): adjust the lighting brightness to 50% (preset range 40%-60%, to avoid strong light reflection affecting the image);
[0227] If the real-time lighting data is >200 lux (high light environment, such as noon on a sunny day): adjust the lighting brightness to 30% (preset range 20%-40%, only supplementing local brightness to avoid overexposure of the image);
[0228] After brightness adjustment, the illumination data of the environmental sensor module 130 is continuously monitored (monitoring cycle 1 second). If the illumination change exceeds 20%, the brightness is readjusted to ensure that it always meets the visual monitoring requirements.
[0229] Those skilled in the art will understand that the process of implementing all or part of the steps of the above embodiments can be carried out by hardware or by a program instructing the relevant hardware.
[0230] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely preferred examples and are not intended to limit the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.
Claims
1. A smart security terminal collaborative monitoring system with edge computing capabilities, characterized in that, include: A multimodal sensing unit (100) is used to collect visual, acoustic and environmental physical signals of the security area to obtain multi-dimensional basic sensing data; The edge computing processing unit (200) is based on multi-dimensional basic perception data. It uses dynamic feature distillation technology and hardware load feedback technology to allocate CPU and FPGA computing tasks, uses incremental convolution kernel pruning technology to optimize the model computing process, judges the computing power requirements and model redundancy of the perception data processing task, and outputs simplified feature data to improve the efficiency of perception data processing. The edge computing processing unit (200) includes a computing power scheduling module (210). The computing power scheduling module (210) collects the operating status parameters of the CPU and FPGA in real time, combines the computational complexity analysis of multi-dimensional basic perception data, and completes the targeted allocation of computing tasks between the CPU and FPGA based on hardware load feedback technology. The targeted allocation of computing tasks by the computing power scheduling module (210) includes the following steps: S210.1, Real-time CPU Utilization Monitoring FPGA resource utilization and memory usage This forms a hardware load status dataset; S210.2, Synchronously acquire visual frame size parameters Voiceprint sampling scale Number of environmental data points and the dimension of features to be extracted Determine the computational complexity benchmark for multi-dimensional basic perception data; S210.
3. Based on hardware load feedback technology, the computational complexity benchmark is matched and analyzed with the hardware load status dataset. When the matching result exceeds the parallel processing threshold... At that time, feature extraction tasks will be specifically assigned to the FPGA; When the matching result does not exceed the parallel processing threshold At that time, logical judgment tasks will be assigned to the CPU; A distributed collaborative communication unit (300) analyzes wireless transmission quality based on edge terminal interaction data, using signal attenuation gradient prediction technology and frequency band dynamic switching technology, utilizes fragmented verification relay technology to ensure data transmission of abnormal events, judges communication link stability, triggers optimal transmission strategies, and realizes cross-terminal data collaboration. The distributed collaborative communication unit (300) includes a transmission quality analysis module (310), a data transmission guarantee module (320), and a collaborative strategy triggering module (330), wherein: The data transmission protection module (320) for abnormal event data breaks down the data into pre-defined fragment sizes and generates a unique checksum for each fragment. Mark the fragment number Simultaneously, data is relayed through nearby terminals, and the relay node verifies the checksum after receiving the fragment. When the verification fails, the corresponding sequence number is triggered. Fragmented retransmission requests; The intelligent decision-making unit (400) aligns visual dynamic features, voiceprint features and environmental physical features in time and space based on the simplified feature data output by the edge computing processing unit (200). It adjusts the judgment criteria of each feature through a dynamic threshold algorithm based on the real-time environmental baseline, optimizes the decision weight of each modality feature, and achieves accurate judgment of abnormal events. The intelligent decision-making unit (400) includes a spatiotemporal alignment module (410) and a decision-making module (420), wherein: The spatiotemporal alignment module (410) is used to receive the simplified feature data output by the edge computing processing unit (200) and call the unified time stamp added by the data synchronization module (140). Spatial coordinates Visual dynamic features Voiceprint characteristics and environmental physical characteristics Mapping to the same spatiotemporal coordinate system generates a multimodal feature set with spatiotemporal consistency. ; The decision-making module (420) is based on the real-time environmental baseline. Using a dynamic threshold algorithm for Each modality feature matching has its own specific judgment threshold: targeting visual dynamic features Set abnormal motion threshold Targeting voiceprint features Set the abnormal voiceprint threshold Targeting environmental physical characteristics Set environmental anomaly thresholds Simultaneously, initial decision weights are assigned to each modality feature. The system optimizes the weight allocation based on the correlation between each modal feature and the abnormal event, and finally outputs the abnormal event judgment result by combining the comparison results of each modal feature and the corresponding threshold and the optimized weight. A multi-level response execution unit (500) triggers physical response actions such as audible and visual alarms, access control, video recording, and lighting linkage based on the anomaly determination result.
2. The intelligent security terminal collaborative monitoring system with edge computing capabilities according to claim 1, characterized in that, The multimodal sensing unit (100) includes a visual acquisition module (110), a voiceprint acquisition module (120), an environmental sensing module (130), and a data synchronization module (140), wherein: The visual acquisition module (110) includes a CMOS image sensor and an automatic exposure control circuit, which is used to acquire visual signals of the security area, adapt to wide range of lighting conditions by adjusting the shutter time, and output standard format video frames carrying optical center coordinates. The voiceprint acquisition module (120) includes a MEMS microphone array and a bandpass filter, used to acquire voiceprint signals, enhance the target sound source through specific frequency band filtering and beamforming algorithms, and output a sampling rate audio stream adapted to voiceprint feature extraction. The environmental sensing module (130) includes a temperature and humidity sensing component and a smoke sensing component, used to collect environmental physical signals; The data synchronization module (140) is used to add a unified time stamp to visual, voiceprint and environmental signals, and establish spatial coordinate association based on the installation positions of the visual acquisition module (110), voiceprint acquisition module (120), environmental sensing module (130) and data synchronization module (140), so that the multi-dimensional basic perception data has spatiotemporal matching.
3. The intelligent security terminal collaborative monitoring system with edge computing capabilities according to claim 2, characterized in that, The edge computing processing unit (200) further includes a feature distillation module (220) and a model optimization module (230), wherein: The feature distillation module (220) is used to map the original features of visual, voiceprint and environmental multimodal features to a unified feature space, and extract high-value features and remove redundant feature dimensions through dynamic feature distillation technology; The model optimization module (230) is used to analyze the feature contribution of convolution kernels in the computation model to identify redundant convolution kernels, remove redundant parts through incremental convolution kernel pruning technology, and optimize the computing power consumption of the model operation.
4. The intelligent security terminal collaborative monitoring system with edge computing capabilities according to claim 3, characterized in that, The collaborative processing of the feature distillation module (220) and the model optimization module (230) includes the following steps: S220.1, Feature distillation module (220) receives multi-dimensional basic perception data and distills visual features. Voiceprint characteristics Environmental characteristics Mapping to a unified feature space to generate a fused feature matrix ; Calculation using dynamic characteristic distillation technology The correlation between various characteristics and security anomalies Retain features with a correlation degree no lower than the feature importance threshold. Feature subset ; The model optimization module (230) processes... The computational model is used to evaluate convolution kernels and calculate the output characteristic fluctuation value of each convolution kernel. and gradient contribution The marker simultaneously satisfies Fluctuation threshold and Contribution threshold Redundant convolution kernels ; Remove Post-validation model output error ,when Error allowable threshold At the same time, maintain the simplified model structure and output based on Simplified feature data.
5. The intelligent security terminal collaborative monitoring system with edge computing capabilities according to claim 4, characterized in that, The transmission quality analysis module (310) collects the signal reception power between terminals based on edge terminal interaction data. Data packet loss rate and transmission delay By combining signal attenuation gradient prediction technology with the signal reception power between terminals in continuous sampling periods, Calculate the signal attenuation gradient To predict signal change trends, and simultaneously switch between preset communication frequency bands using dynamic frequency band switching technology, and test the signal-to-noise ratio of each frequency band. The final output is divided into "Excellent" ,middle ,Difference Level 3 transmission quality The transmission quality analysis process of the transmission quality analysis module (310) includes the following steps: S310.1, Signal reception power between acquisition terminals Data packet loss rate and transmission delay Determine the method used to calculate the signal attenuation gradient. The number of consecutive sampling periods; S310.2, Inter-terminal signal reception power within a preset number of continuous sampling periods Calculate the statistical value of the power difference between adjacent cycles, and use the statistical value of the power difference between adjacent cycles as the signal attenuation gradient. and based on The positive and negative values and absolute value range can be used to predict the trend of subsequent signal changes; S310.
3. Switch to each preset communication frequency band using dynamic frequency band switching technology. The test duration for each frequency band is controlled according to preset standards. During the test, interference signals from non-target frequency bands are shielded, and the signal-to-noise ratio of each frequency band is recorded. The statistical values, and the signal-to-noise ratio of each frequency band. The statistical value is used as the effective signal-to-noise ratio of the corresponding frequency band; S310.4, combined with signal attenuation gradient Range, effective signal-to-noise ratio of each frequency band Range and transmission delay Scope of transmission quality level :when Within the preset stable range, effective signal-to-noise ratio Within the preset high priority range and transmission delay When within the preset low latency range, it is judged as excellent. ;when Within the preset basic stable range, effective signal-to-noise ratio Within the preset medium range and transmission delay When the delay is within the preset acceptable range, it is judged as medium. All other cases are judged as poor. .
6. The intelligent security terminal collaborative monitoring system with edge computing capabilities according to claim 5, characterized in that, The cooperative strategy triggering module (330) receives the transmission quality level. ,pass With stability threshold The comparison is used to judge the stability of the communication link. When the direct transmission strategy is triggered, The system selects neighboring terminals whose transmission quality meets preset requirements as relay nodes and triggers a relay fragmentation strategy. Simultaneously, it synchronizes the fragmentation reception status of each terminal based on a cross-terminal data interaction protocol. Once the target terminal receives all fragments and passes verification, it triggers a data reassembly command, thus achieving cross-terminal data collaboration. The collaborative working process of the data transmission guarantee module (320) and the collaboration strategy triggering module (330) includes the following steps: S320.1, The data transmission guarantee module (320) obtains the total amount of abnormal event data, matches and uses a preset segment of the corresponding size according to the preset interval where the total amount of data is located, breaks down the abnormal event data according to the segment size, and generates a unique check code for each segment. And mark the fragment number. ; S320.2, The collaborative strategy triggering module (330) takes the current terminal as the center and first filters the transmission quality level within the preset short-range communication range. The nearest terminal; if no suitable terminal is found within this range, the range is expanded to a preset medium-range communication range, and the transmission quality level is selected. The nearest terminals; for the selected relay candidate terminals within the medium-range communication range, combined with The transmission quality level and communication distance are weighted and evaluated, and priorities are assigned to relay nodes based on the evaluation results; the candidate relay terminals are determined as relay nodes and priorities are assigned based on the evaluation results. S320.3, the collaborative strategy triggering module (330) is based on the cross-terminal data interaction protocol to synchronize the reception and verification results of each terminal on the fragment, and each terminal provides real-time feedback information through the cross-terminal data interaction protocol; S320.4 When the target terminal returns a certain fragment sequence number When the verification fails and the number of failures reaches a preset threshold, the collaborative strategy triggering module (330) switches to the next highest priority relay node and controls the data transmission guarantee module (320) to re-initiate the segment sequence number. The transmission request continues until the target terminal passes the verification.
7. The intelligent security terminal collaborative monitoring system with edge computing capabilities according to claim 6, characterized in that, The multi-level response execution unit (500) includes a response triggering module (510) and an action execution module (520), wherein: The response triggering module (510) is used to receive the abnormal event judgment result output by the intelligent decision unit (400), generate corresponding response instructions according to the severity of the abnormal event, and trigger the basic response combination instruction when the abnormal event is judged to be a general abnormality; and trigger the full response combination instruction when the abnormal event is judged to be a serious abnormality. The action execution module (520) executes the corresponding physical response after receiving the instruction from the response trigger module (510); wherein, when executing the lighting linkage, the brightness of the lighting equipment is adjusted to a preset range suitable for abnormal scene monitoring by combining the real-time illumination data of the environment sensing module (130) in the multimodal sensing unit (100).
Citation Information
Patent Citations
Safety monitoring method based on mixing of large model and neural network algorithm
CN119380166A
An intelligent terminal monitoring system and method based on edge computing
CN120128697B
Accelerator based on multi-mirror FPGA, accelerator implementation method, terminal equipment and computer readable storage medium
CN116541073A
Intelligent blueberry disease detection method and system based on multi-mode unsupervised learning
CN120451968A