Multi-modal edge computing law enforcement recorder and data processing method thereof

By using a multimodal edge computing law enforcement recorder, the problems of strong network dependence, insufficient multimodal fusion capability, and insufficient data security in existing technologies have been solved. It enables real-time intelligent analysis and secure data storage in network interruption scenarios, thereby improving the intelligent decision-making capability and evidence integrity of law enforcement equipment.

CN122496602APending Publication Date: 2026-07-31GUANGDONG POLICE COLLEGE (GUANGDONG PROVINCIAL PUBLIC SECURITY JUDICIAL MANAGEMENT CADRE COLLEGE)
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGDONG POLICE COLLEGE (GUANGDONG PROVINCIAL PUBLIC SECURITY JUDICIAL MANAGEMENT CADRE COLLEGE)
Filing Date
2026-04-04
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing smart law enforcement recorders rely on cloud computing, and their intelligent analysis functions are paralyzed when network signals are limited. They also lack multimodal data fusion capabilities, edge computing capabilities, and data security, making it difficult to meet the requirements for real-time decision support and the integrity of the evidence chain.

Method used

The multimodal edge computing law enforcement recorder integrates a heterogeneous computing engine, a deep multimodal fusion mechanism, and national cryptographic algorithms plus blockchain anchoring technology to achieve real-time intelligent analysis and secure data storage on the device side. It forms a closed-loop collaborative architecture through multimodal data acquisition, edge computing processing, multimodal fusion analysis, knowledge base judgment, and secure storage.

Benefits of technology

It provides complete intelligent analysis capabilities in network outage or classified scenarios, enhances the comprehensive understanding of complex law enforcement scenarios, ensures the immutability and traceability of law enforcement evidence, shortens decision response time, and optimizes model capabilities through federated learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure REF-OBJ-1775297285762-000001
    Figure REF-OBJ-1775297285762-000001
  • Figure REF-OBJ-1775297285762-000002
    Figure REF-OBJ-1775297285762-000002
  • Figure REF-OBJ-1775297285762-000003
    Figure REF-OBJ-1775297285762-000003
Patent Text Reader

Abstract

This invention discloses a multimodal edge computing law enforcement recorder and its data processing method. The recorder includes a multimodal data acquisition module, an edge computing processing module, a multimodal fusion analysis module, a local knowledge base module, a secure storage module, and an intelligent interaction module. The edge computing processing module adopts a heterogeneous computing architecture of CPU, NPU, and FPGA to achieve real-time inference of deep learning models on the edge. The multimodal fusion analysis module performs deep semantic fusion of video and audio features based on a cross-modal attention mechanism. The secure storage module uses national cryptographic encryption, digital signatures, and blockchain hash anchoring mechanisms to ensure the integrity of the evidence chain. This invention enables the law enforcement recorder to perform independent intelligent analysis in a network-free environment, improving law enforcement efficiency and data security.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent law enforcement equipment technology, and in particular to a multimodal edge computing law enforcement recorder and its data processing method. Background Technology

[0002] With the continuous advancement of public security informatization, body cameras have become indispensable standardized equipment for frontline police officers. From the initial analog cameras to digital high-definition recorders, the core function of this type of equipment has long remained at the level of passive audio and video recording, with very limited ability to proactively analyze complex situations at law enforcement scenes, making it difficult to provide real-time auxiliary decision support for frontline law enforcement personnel. In recent years, artificial intelligence technology has made breakthrough progress in computer vision, natural language processing, and multimodal learning, providing a new technological path for the intelligent upgrading of body cameras.

[0003] In existing technologies, Chinese invention patent application publication number CN119625961A discloses a brain-like edge computing-driven multimodal intelligent security method and system. This solution collects multimodal data such as video, audio, and sensor data, preprocesses and extracts dynamic behavioral features using a brain-like edge computing gateway, and generates a formal warning after secondary confirmation through a lightweight visual large model. However, this solution is mainly aimed at fixed-installation security monitoring scenarios. The size, power consumption, and computing power deployment of its edge computing gateway are not suitable for the harsh working environment of wearable law enforcement equipment for individual soldiers, and it does not address the issues of knowledge base support and evidence chain integrity assurance specific to the law enforcement field. In addition, Chinese invention patent application publication number CN108805783A discloses a law enforcement record management system based on big data co-construction and sharing. This solution focuses on solving the information sharing problem of cross-level law enforcement record data, but its data processing relies entirely on cloud servers. It cannot provide real-time intelligent analysis functions in scenarios with limited network signals or confidentiality, and it lacks multimodal data fusion capabilities at the device end. Further research revealed that most publicly available patents related to smart wearable law enforcement devices limit the role of the device to the data acquisition front end, performing only lightweight processing such as video compression encoding and basic motion detection. Deeper intelligent analysis tasks still need to be uploaded to the backend server for execution. This "front-end acquisition + backend analysis" architecture faces significant latency bottlenecks and availability risks in actual deployment.

[0004] A comprehensive analysis of existing technologies reveals the following key technical bottlenecks in current intelligent law enforcement recording devices: First, they rely excessively on cloud computing. When law enforcement officers are in signal blind spots, underground spaces, or areas with restricted communication, the device's intelligent analysis functions are completely paralyzed, failing to meet real-time decision-making support needs. In reality, network unavailability is a significant issue in law enforcement scenarios, particularly in densely populated urban villages, underground parking lots, and mountainous forest areas—typical high-incidence areas for law enforcement. Second, multimodal data fusion mechanisms are lacking. Existing solutions typically only process a single video modality or simply splice video and audio, failing to establish effective cross-modal spatiotemporal correlation and semantic fusion mechanisms. This results in insufficient understanding of complex law enforcement scenarios, such as when a suspect makes threatening statements. When a statement is accompanied by an aggressive action, a single-modal system cannot correlate the voice content with the visual behavior to make an accurate risk assessment. Third, the heterogeneous computing architecture design at the edge is immature; traditional general-purpose processors struggle to support the parallel inference needs of multiple deep learning models within a limited power budget, preventing the device from simultaneously performing facial recognition and other intelligent functions such as behavioral analysis. Fourth, the secure storage of law enforcement data and the mechanisms for ensuring the integrity of the evidence chain are still inadequate. Existing encryption schemes cannot simultaneously meet the multiple requirements of data tamper-proofing, verifiable traceability, and efficient access. Some schemes, while employing encrypted storage, lack an immutable timestamp proof mechanism, failing to meet the stringent requirements of judicial evidence collection for the integrity of the evidence chain. Therefore, there is an urgent need for a new type of intelligent law enforcement recorder that integrates powerful edge computing capabilities, deep multimodal fusion capabilities, and high-level security capabilities. Summary of the Invention

[0005] To address the technical problems of existing law enforcement recorders, such as strong cloud dependence, insufficient multimodal fusion capabilities, limited edge computing capabilities, and insufficient data security, this invention provides a multimodal edge computing law enforcement recorder and its data processing method. By constructing a heterogeneous computing engine and a deep multimodal fusion mechanism on the device side, it can realize real-time autonomous assessment of the law enforcement scene situation. At the same time, it uses national cryptographic algorithms and blockchain anchoring technology to ensure the integrity and credibility of law enforcement evidence.

[0006] The first aspect of this invention provides a multimodal edge computing law enforcement recorder, comprising: a multimodal data acquisition module for simultaneously acquiring video streams, audio streams, location information, and environmental sensor data from the law enforcement scene, and attaching a unified timestamp to each acquired data stream to generate a multimodal raw data stream; an edge computing processing module, employing a heterogeneous computing architecture, integrating a high-performance processor, a neural network acceleration chip, and a reconfigurable logic unit, for performing real-time feature extraction, target detection, behavior recognition, and speech understanding on the multimodal raw data stream at the device end, and outputting structured analysis results for each modality; and a multimodal fusion analysis module, based on cross-modal injection... The system comprises several modules: a **intention mechanism** for spatiotemporal alignment and semantic fusion of structured analysis results from different modalities to generate a comprehensive event description vector; a **local knowledge base module** for storing law enforcement knowledge graphs and historical case libraries, used to retrieve similar cases and associate legal clauses based on the comprehensive event description vector, generating law enforcement assessment suggestions; a **secure storage module** employing national cryptographic algorithms and blockchain hash anchoring mechanisms for encrypted writing, digital signatures, and integrity verification of law enforcement data; and an **intelligent interaction module** for providing law enforcement assessment suggestions to law enforcement personnel in voice or visual form, and for data interaction with the command center through a secure channel.

[0007] The six modules described above form a deeply coupled closed-loop collaborative architecture: the multimodal data acquisition module provides the edge computing processing module with a raw data stream based on a unified time benchmark; the structured analysis results from the edge computing processing module drive the multimodal fusion analysis module to perform cross-modal semantic fusion; the event description vectors generated by the fusion support the intelligent judgment of the local knowledge base module; and the judgment results are fed back to law enforcement personnel through the intelligent interaction module, triggering the encrypted evidence storage process of the secure storage module. Furthermore, the judgment output of the local knowledge base module can inversely adjust the detection priority and model selection strategy of the edge computing processing module, forming a closed loop of feedforward analysis and feedback optimization, making the overall system performance significantly better than the linear superposition of the independent operation of each module.

[0008] The second aspect of the present invention provides a data processing method for a multimodal edge computing law enforcement recorder, applied to the aforementioned law enforcement recorder. The method includes the following steps: Step S1, synchronously acquiring video streams, audio streams, location information, and environmental sensor data from the law enforcement scene through a multimodal data acquisition module, and generating a multimodal raw data stream by attaching a unified timestamp; Step S2, preprocessing the multimodal raw data, including extracting keyframes from the video, windowing the audio frames, and standardizing the data; Step S3, inputting the preprocessed data into an edge computing processing module, and performing real-time intelligent analysis, including target detection, face recognition, license plate recognition, speech recognition, and behavior analysis, through a pre-deployed deep learning model; Step S4, performing cross-modal spatiotemporal correlation and semantic fusion on the structured analysis results of each modality in Step S3 through a multimodal fusion analysis module, generating a comprehensive event description vector; Step S5, intelligently judging the comprehensive event description vector based on a local knowledge base, and generating law enforcement suggestions and handling plans; Step S6, feeding back the analysis results and law enforcement suggestions in real time through an intelligent interaction module, and storing the law enforcement data after encryption and signing by a secure storage module.

[0009] The beneficial effects of this invention are as follows: It eliminates the reliance on cloud computing power through a heterogeneous edge computing architecture, providing complete intelligent analysis functions even in network outages or confidential scenarios; the multimodal deep fusion technology based on cross-modal attention mechanisms enhances the comprehensive understanding of complex law enforcement scenarios; the multi-layered data protection mechanism combining national cryptographic algorithms with blockchain hash anchoring ensures the immutability and traceability of law enforcement evidence; edge-side intelligent analysis supported by a local knowledge base shortens the response time from situational awareness to decision-making; and the federated learning mechanism achieves continuous collaborative evolution of model capabilities while protecting data privacy. Attached Figure Description

[0010] Figure 1 This is a schematic diagram of the system architecture of the multimodal edge computing law enforcement recorder of the present invention.

[0011] Figure 2 This is a schematic diagram of the heterogeneous computing architecture of the edge computing processing module of the present invention.

[0012] Figure 3 This is a flowchart of the multimodal fusion analysis process of the present invention.

[0013] Figure 4 This is a flowchart of the data processing method of the present invention.

[0014] Figure 5 This is a schematic diagram of the data protection mechanism of the secure storage module of the present invention.

[0015] Figure 6 This is a schematic diagram of the federated learning mechanism of the present invention. Detailed Implementation

[0016] The specific embodiments of the present invention will be further described below with reference to the accompanying drawings. It should be noted that these descriptions are for the purpose of aiding understanding the present invention, but do not constitute a limitation thereof. Furthermore, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0017] like Figure 1 As shown, this invention provides a multimodal edge computing law enforcement recorder. The recorder adopts a portable wearable design, meeting the technical standards for single-officer law enforcement audio-visual recorders set by the Ministry of Public Security. Its casing meets IP68 protection requirements, and the overall dimensions are controlled within approximately 90mm x 62mm x 30mm, with a weight not exceeding 280g. It is suitable for wearing on the chest or shoulder of law enforcement officers without interfering with normal law enforcement operations. The recorder's system architecture includes a multimodal data acquisition module 101, an edge computing processing module 102, a multimodal fusion analysis module 103, a local knowledge base module 104, a secure storage module 105, and an intelligent interaction module 106. These modules are interconnected via a high-speed internal bus, forming a complete closed-loop processing chain from data acquisition, edge intelligent analysis, cross-modal fusion judgment to secure encrypted evidence storage.

[0018] The multimodal data acquisition module 101 undertakes the task of synchronously acquiring multi-source perception data at the law enforcement scene. In one embodiment of the present invention, this module integrates the following sensor components: a high-definition camera array including a front-facing main camera and a wide-angle auxiliary camera for video input, wherein the front-facing main camera uses a 1 / 2.5-inch CMOS image sensor, supports 4K resolution video acquisition, and is equipped with an f / 1.8 large aperture lens and an infrared fill light, which can still acquire clear images in low-light environments (as low as 0.1 lux); the wide-angle auxiliary camera uses a 120-degree field-of-view lens to capture panoramic information of the environment around law enforcement personnel. The directional microphone array consists of four high-sensitivity MEMS microphones, which achieve directional sound pickup through a digital beamforming algorithm, and can improve the signal-to-noise ratio of the main pickup direction by at least 15dB in noisy environments, while supporting 360-degree environmental sound field monitoring. Preferably, the sampling rate of the microphone array is set to 16kHz and the quantization accuracy is 16bit, which can meet the accuracy requirements of speech recognition and voiceprint analysis.

[0019] The positioning unit employs a dual-mode satellite positioning scheme of BeiDou and GPS, providing sub-meter positioning accuracy in open environments. It also incorporates an inertial navigation unit (six-axis accelerometer and gyroscope) for indoor positioning enhancement. Through an inertial navigation recursive algorithm, it can still provide continuous position trajectory estimation within buildings where satellite signals are obstructed, with a position drift rate controlled to no more than 2 meters per minute. The environmental sensor group includes digital temperature and humidity sensors (accuracy of ±0.3 degrees Celsius for temperature and ±2%RH for humidity) and electrochemical gas detection sensors (supporting CO detection range of 0 to 500 ppm and combustible gas detection range of 0 to 100% LEL), used to sense the environmental safety status at law enforcement sites. Preferably, the module also integrates a heart rate and blood oxygen sensor, which monitors the physiological state of law enforcement officers in real time through photoplethysmography (PPG) technology. When an abnormally high heart rate (exceeding 140 bpm for more than 30 seconds) or a low blood oxygen saturation (below 90% for more than 15 seconds) is detected, a physiological abnormality warning signal is automatically triggered, prompting law enforcement officers to pay attention to their own health status or reminding the command center to pay attention to the safety of the law enforcement officers.

[0020] A key technical feature of this invention is that the multimodal data acquisition module 101 applies a unified timestamp to the data from each sensor. Specifically, this module incorporates a high-precision clock chip (with an accuracy better than 1ms). All data frames acquired by the sensors generate timestamps with millisecond-level precision based on this clock source, thus establishing a foundation for cross-modal time alignment at the source of data generation. The video stream is continuously acquired at a frame rate of 30fps, with each frame carrying a corresponding timestamp; the audio stream is acquired through continuous sampling, with frames marked according to a frame length of 25ms and a frame shift of 10ms; location information and environmental data are updated at frequencies of one second and every 10 seconds, respectively. Through this unified timestamp mechanism, subsequent multimodal fusion analysis can accurately establish the temporal correspondence between video frames, audio frames, and location information.

[0021] like Figure 2As shown, the edge computing processing module 102 is the core engine for realizing intelligent analysis on the device side in this invention. This module adopts a heterogeneous computing architecture and is composed of three types of computing units working together. The high-performance CPU core 201 uses the ARM Cortex-A78 architecture, with a main frequency of up to 2.8GHz and a 4-core design. It is mainly responsible for system-level task scheduling, general data preprocessing, and non-intensive computing inference tasks. The neural network acceleration chip (NPU) 202 uses a dedicated AI inference accelerator with a peak computing power of up to 10 TOPS (INT8 precision). It can efficiently execute core deep learning operations such as convolution operations, matrix multiplication, and activation function calculations, and is the main computing engine for video analysis and multimodal fusion. The reconfigurable logic unit (FPGA) 203 provides hardware-level algorithm customization acceleration capabilities, and performs dedicated pipeline acceleration for specific signal processing algorithms (such as beamforming and FFT transformation) and data encryption algorithms (such as SM4 encryption), with latency as low as microseconds.

[0022] Preferably, the three types of computing units interact via a unified shared memory bus. The CPU is responsible for dynamically allocating computing resources based on the current task load: when the video analysis task has the highest priority, the NPU performs video inference at full capacity, while the CPU and FPGA share the audio processing and encryption tasks; when a sudden high-priority recognition request occurs (such as comparison of key personnel), the CPU can temporarily interrupt low-priority tasks and concentrate the NPU's computing power on the face recognition model. This dynamic resource scheduling strategy enables the entire edge computing module to achieve near real-time parallel inference of multiple AI models within a total power consumption budget constraint of no more than 5W.

[0023] In one embodiment of the present invention, the edge computing processing module 102 pre-deploys a group of deep learning models optimized by model compression to achieve efficient inference under limited computing power. The target detection model adopts the YOLOv8n architecture, with an input image size of 640x640 pixels and approximately 3.2M model parameters. The single-frame inference time on the NPU is approximately 15ms, enabling real-time detection and tracking of targets such as people, vehicles, and objects in the video stream at a rate greater than 30fps. The face recognition model uses a lightweight ArcFace network that has undergone knowledge distillation compression, with a feature vector dimension of 512. When performing 1:N face comparison, the time for retrieval and matching with a locally stored 100,000-level face feature database is controlled within 50ms, with a recognition accuracy exceeding 99.5%. The license plate recognition model adopts an improved CRNN architecture, supporting end-to-end recognition of various types of license plates, including blue plates, yellow plates, and new energy vehicle license plates, with a recognition accuracy exceeding 98%. The speech recognition model adopts a lightweight Whisper architecture with knowledge distillation, compressing the number of model parameters to 1 / 4 of the original model. It supports continuous speech recognition in both Chinese and English, and the word error rate is controlled within 5%.

[0024] Preferably, the above model compression process adopts a three-stage strategy: The first stage is knowledge distillation, which uses a large teacher model (such as YOLOv8l, Whisper Large) to guide the training of a compact student model, so that the student model can maintain recognition accuracy close to that of the teacher model under the condition of a significant reduction in the number of parameters; The second stage is weight quantization, which quantizes the model weights from FP32 floating-point precision to INT8 integer precision, increasing the computational throughput by about 3 to 4 times, while controlling the accuracy loss introduced by quantization to within 0.5% through quantization-aware training (QAT); The third stage is structured pruning, which removes the convolutional kernel channels with the lowest contribution based on the L1 norm importance evaluation of the weights of each layer, reducing the model's FLOPs by 30% to 50%, and further shortening the inference latency.

[0025] In terms of behavior analysis, the edge computing processing module 102 also integrates a human behavior recognition model based on a spatiotemporal graph convolutional network (ST-GCN). This model takes the temporal sequence of human skeletal key points as input and captures the spatial topological relationships and temporal movement patterns between various joints of the human body through graph convolution operations. It can recognize various common behavior types in law enforcement scenarios, including running, fighting, falling, and surrendering. In one embodiment of the invention, the input to the behavior recognition model is a sequence of coordinates of 17 human key points across 16 consecutive frames. The graph convolution layer has 9 layers, each containing 64 feature channels. The inference latency of the entire model is approximately 20ms, which meets the requirements for real-time behavior warning. Preferably, the training dataset of the behavior recognition model contains over 50,000 labeled human action sequence samples from law enforcement scenarios, covering 12 behavior categories such as normal walking, running, squatting, lying down, pushing, fighting, and brandishing weapons. The model achieves a classification accuracy of 93.7% on the test set. When the model detects high-risk behavior (fighting, brandishing weapons), the system generates a high-priority behavior warning within 200ms and immediately notifies law enforcement personnel through the intelligent interaction module 106.

[0026] Furthermore, the edge computing processing module 102 employs a priority queue-based task scheduling strategy to coordinate the parallel execution of the aforementioned multiple AI models. In one embodiment of the present invention, the default priorities of each analysis subtask, from high to low, are as follows: abnormal behavior detection (highest priority, due to personal safety concerns), face recognition comparison (high priority, due to key personnel identification), target detection and tracking (medium priority, providing basic scene awareness), speech recognition (medium priority), license plate recognition (low priority), and environmental awareness (low priority). When the NPU's computational load reaches more than 85% of its peak computing power, the task scheduler automatically pauses the inference requests of low-priority tasks, concentrating computing power on high-priority tasks to ensure that safety-critical analysis tasks always receive timely responses. Once the high-priority tasks are completed, the paused low-priority tasks automatically resume execution.

[0027] like Figure 3 As shown, the multimodal fusion analysis module 103 is a key component for achieving cross-modal deep understanding in this invention. Unlike simple decision-level fusion (which directly merges the independent analysis results of each modality), this invention employs a multimodal encoder based on the Transformer architecture to achieve deep semantic fusion of the two main modalities, video and audio, at the feature level. Preferably, this fusion process sequentially executes the following steps: First, the visual semantic feature generation step S4a encodes the features of the video stream. In one embodiment of the present invention, a SlowFast dual-path network is used to extract the spatiotemporal features of the video. The Slow path captures spatial semantic information at a low frame rate (4 frames per second), while the Fast path captures motion temporal information at a high frame rate (32 frames per second). The features of the two paths are fused through lateral connections, ultimately generating a visual semantic feature vector sequence with a dimension of 2048, where each vector corresponds to a time segment (approximately 0.5 seconds).

[0028] Meanwhile, the speech semantic feature generation step S4b encodes the features of the audio stream. In one embodiment of the present invention, a wav2vec 2.0 pre-trained model is used to extract the acoustic and semantic features of the audio. This model first encodes the original audio waveform into a frame-level feature sequence through multiple one-dimensional convolutions, and then captures contextual semantic information through a 12-layer Transformer encoder, finally outputting a speech semantic feature vector sequence with a dimension of 768, with each vector corresponding to an audio frame of approximately 20ms.

[0029] The spatiotemporal alignment step S4c, based on the unified timestamps added during the data acquisition phase, precisely aligns the visual and audio feature sequences along the time axis. Since the two modalities have different temporal resolutions (0.5s for video and 20ms for audio), this step uses a sliding aggregation operation within a time window to perform average pooling on the audio features within the window corresponding to each video time segment, ensuring that the feature sequences of both modalities achieve a consistent length in the temporal dimension.

[0030] The cross-modal attention calculation step S4d is the core of the fusion process. In one embodiment of this invention, a cross-attention mechanism is employed to achieve mutual enhancement between features from two modalities: visual features are used as the query vector, and speech features are used as the key and value vectors. The attention weight of visual features on each frame of speech features is calculated using scaled dot product attention, and then the speech features are weighted and summed according to this weight to obtain the speech-enhanced visual features; conversely, speech features are used as the query vector to focus on visual features to obtain the visual-enhanced speech features. Through this bidirectional cross-attention mechanism, features from each modality incorporate complementary information from the other modality.

[0031] The feature fusion step S4e concatenates the visual and speech features enhanced by cross-attention, and performs dimensionality compression and nonlinear transformation through a fully connected layer to generate a 512-dimensional multimodal fusion feature vector. Finally, step S4f performs global average pooling and normalization on the fusion feature sequence to generate a fixed-dimensional comprehensive event description vector, which will serve as the unified input representation for subsequent knowledge base retrieval and intelligent judgment.

[0032] Preferably, the multimodal fusion analysis module 103 is also equipped with a dynamic weight adjustment mechanism. This mechanism evaluates the signal quality of each modality in real time: for video modalities, a quality score is calculated based on indicators such as image sharpness (Laplacian variance), illumination uniformity, and target occlusion rate; for audio modalities, a quality score is calculated based on indicators such as signal-to-noise ratio, voice activity detection (VAD) results, and reverberation time. When the quality score of a certain modality is lower than a preset threshold (e.g., 0.3, out of 1.0), the weight of that modality in the fusion process will be automatically reduced, thereby avoiding the negative impact of low-quality data on the fusion result. In one embodiment of the present invention, when the environmental noise at the law enforcement scene is extremely high (signal-to-noise ratio below 5dB) causing the voice modality quality score to be only 0.15, the system automatically reduces the fusion weight of the voice modality from the default 0.4 to 0.1, while increasing the weight of the visual modality from 0.6 to 0.9, ensuring that accurate situational assessment can still be performed based on visual information even under unreliable audio conditions.

[0033] The local knowledge base module 104 provides domain knowledge support for intelligent analysis at the edge. In one embodiment of the present invention, the storage capacity of this module is 64GB, and it contains the following knowledge resources: a law enforcement domain knowledge graph containing more than 5,000 entity nodes (covering types such as personnel, organizations, items, locations, and events) and more than 20,000 relationship edges (covering relationship types such as affiliation, possession, occurrence, and involvement). This knowledge graph is stored in a compressed triple format and supports fast retrieval based on graph embedding vectors; a legal and regulatory database containing more than 3,000 commonly used legal and regulatory provisions, and establishing a scenario-legal provision mapping index according to common public security law enforcement scenarios; and a historical case database storing more than 100,000 structured typical law enforcement case records. Each case includes fields such as event type tags, handling process summary, and handling result evaluation, and pre-calculates the semantic embedding vector of the case, supporting fast case retrieval based on vector similarity.

[0034] During the intelligent analysis process, the local knowledge base module 104 receives the comprehensive event description vector output by the multimodal fusion analysis module 103 and performs the following operations in sequence: First, it calculates the cosine similarity between the event description vector and the case embedding vector in the historical case library, and retrieves the top 5 similar cases as a reference; then, it queries the applicable legal provisions in the legal provisions library based on the event type classification results (such as public security disputes, traffic violations, suspected crimes, etc.) and generates a provision summary and application suggestions; finally, it comprehensively assesses the on-site risk level (calculated based on factors such as the number of participants, whether dangerous items are carried, and the degree of emotional agitation) and generates graded disposal suggestions (different suggested disposal strategies correspond to low risk / medium risk / high risk).

[0035] like Figure 5 As shown, the secure storage module 105 employs a multi-layered data protection mechanism to ensure the integrity, authenticity, and immutability of law enforcement data as legal evidence. The data writing process includes the following five sequential processing steps: Step 1 (No. 501): The original law enforcement data to be stored is encrypted using the SM4 national standard symmetric encryption algorithm. The SM4 algorithm uses a 128-bit key length and CBC encryption mode. The initialization vector (IV) is generated by a secure random number generator, and a different IV value is used for each write operation. Step 2 (No. 502): A SHA-256 hash digest is calculated on the encrypted data to generate a 256-bit data fingerprint. This fingerprint uniquely identifies the data content; any minor modification to the data will result in a drastic change in the hash value. Step 3 (No. 503): The SHA-256 hash value is signed using ECDSA elliptic curve digital signature using the device private key stored in the security chip (SE). The signing process is completed inside the security chip, and the private key never leaves the chip, thus ensuring the signature's unforgeability. Step 4 (No. 504): The signed encrypted data is fragmented and stored using the RS(6,4) erasure coding algorithm. Each data block is divided into 4 data fragments and 2 check fragments. Even if any two data fragments are lost, the system can still recover the complete data, significantly improving storage reliability. Step 5 (No. 505): Periodically (once per hour by default), the data hash digest and signature information are synchronized to the public security consortium chain node through a secure channel. The blockchain network confirms the information through consensus and writes it into the block, forming a legally valid trusted timestamp proof.

[0036] Accordingly, the data reading and verification process includes five reverse steps (numbered 506 to 510): reading the data fragment (506), erasure coding recovery (507), verifying the ECDSA digital signature (508), verifying the SHA-256 hash digest (509), and decrypting and restoring the original data using SM4 (510). Only when both signature verification and hash verification pass is the data deemed complete and valid; if any step fails, the system will mark the data as "integrity questionable" and record the exception in the log for subsequent auditing.

[0037] The intelligent interaction module 106 is responsible for providing the analysis and judgment results to law enforcement personnel in a user-friendly manner. In one embodiment of the present invention, the module supports three interaction methods: voice feedback broadcasts key information (such as "The personnel ahead have been identified in the key personnel database, it is recommended to verify their identity") to law enforcement personnel via bone conduction speakers or Bluetooth headsets, with speech synthesis using an edge-side TTS engine and latency controlled within 200ms; visual prompts display real-time analysis results in the form of icons and text overlays on a 2.0-inch touch screen on the back of the device; and emergency upload pushes the simplified analysis results (event summary and key screenshots) of high-priority events to the comprehensive judgment platform of the command center in real time via a 4G / 5G wireless communication module or a private network communication module.

[0038] Preferably, the intelligent interaction module 106 also supports voice wake-up and continuous dialogue functions. Law enforcement officers can activate the AI ​​voice interaction function through a preset wake-up word (such as "Xiao Da Xiao Da"). After the device enters the voice interaction mode, law enforcement officers can consult the device in natural language, such as querying legal provisions, obtaining handling suggestions, or requesting the activation of specific recognition functions. The system supports multiple rounds of continuous dialogue within 5 seconds, and automatically exits the voice interaction mode after the timeout to save power. In addition, the intelligent interaction module 106 is also equipped with an SOS emergency help button. When law enforcement officers press this button, the device immediately triggers the following linkage operations: automatically starts the highest resolution video recording mode, sends the current GPS coordinates and real-time video screenshots to the command center through an encrypted channel, and starts continuous location tracking and reporting (the frequency is increased to once every 2 seconds), providing technical support for rapid response and precise reinforcement in emergency situations.

[0039] In one embodiment of the present invention, the overall power consumption management of the law enforcement recorder adopts a hierarchical strategy. In standby mode, only basic video recording and GPS positioning functions are maintained, with a total power consumption of approximately 2.5W. The built-in 6000mAh lithium polymer battery can support approximately 16 hours of continuous standby recording. In standard intelligent analysis mode, the edge computing processing module 102 runs a quantized and compressed lightweight model group, with a total power consumption of approximately 4.5W and a battery life of approximately 8 hours. In high-performance mode (e.g., simultaneously performing face recognition comparison, behavior analysis, and multimodal fusion), both the NPU and FPGA run at full load, with a total power consumption of approximately 7W and a battery life of approximately 5 hours. The CPU dynamically switches between the above power consumption modes based on the current task load and the remaining battery power, maximizing the device's battery life while ensuring the continuous availability of critical functions.

[0040] In terms of data communication architecture, the intelligent interaction module 106 supports multiple wireless communication standards to adapt to the differences in network environments under different law enforcement scenarios. In one embodiment of the present invention, the module integrates a 4G LTE Cat.4 communication module (supporting uplink speed of 50Mbps and downlink speed of 150Mbps) and a public security dedicated network PDT (Police Digital Trunking) communication module. The two communication standards can automatically switch or work simultaneously according to the network environment. Within the coverage area of ​​the public security dedicated network, the PDT communication module is used first to conduct real-time data interaction with the command center to ensure the security and high reliability of the communication link; outside the coverage area of ​​the dedicated network, it automatically switches to the 4G public network channel and establishes a secure connection with the backend system through a VPN encrypted tunnel. Preferably, the device also supports Wi-Fi 6 wireless LAN connection, and after returning to the office area, it can use the high-speed Wi-Fi channel to synchronize large-capacity law enforcement data stored locally to the backend server in batches, improving data archiving efficiency.

[0041] like Figure 6As shown, in a preferred embodiment of the present invention, the law enforcement recorder is also equipped with a federated learning module. This module enables multiple law enforcement recorder devices to collaboratively optimize a shared model without sharing the original law enforcement data, making full use of the law enforcement data resources distributed across each device to improve the model's generalization ability and scenario adaptability. Specifically, the federated learning training process is as follows: Each device uses locally accumulated law enforcement data (after anonymization, including face blurring, voice changing, and other privacy protection operations) to incrementally train its local model. During training, a differential privacy mechanism is used to add Gaussian noise protection to the gradients (the noise standard deviation is set to 1.1 times the model sensitivity, corresponding to a privacy budget of epsilon=1.0), ensuring that even if parameter updates are intercepted, the information of the original training data cannot be deduced. After training, each device only uploads the model parameter update amount (incremental parameter ΔW) to a coordination server deployed on the public security intranet. The upload process is protected by a TLS 1.3 encrypted channel. The coordination server uses a secure aggregation protocol (such as Secure Aggregation based on secret sharing) to perform weighted aggregation of parameter updates from multiple devices. The weight coefficients are dynamically calculated based on the amount of local training data and data quality scores of each device to generate new global model parameters. The global model parameters are then distributed to each device to update its local model. The above process is executed iteratively on a weekly basis to achieve continuous optimization of model capabilities. In one embodiment of the present invention, in a pilot deployment comprising 50 law enforcement recorders, after 20 rounds of federated learning iterations, the shared object detection model's mAP (memory accuracy) on the law enforcement scenario test set improved from an initial 71.2% to 79.7%, an improvement of approximately 8.5 percentage points; the behavior recognition model's Top-1 accuracy improved from 87.5% to 93.7%, an improvement of approximately 6.2 percentage points; and the face recognition model's recognition rate under large angle and low lighting conditions improved from 91.3% to 95.8%, an improvement of approximately 4.5 percentage points, verifying the effectiveness of the federated learning mechanism in the collaborative optimization of law enforcement equipment groups. Preferably, the federated learning module also supports an asynchronous aggregation mode. For devices that cannot complete parameter upload within the specified time window due to network limitations, their parameter updates can be uploaded in subsequent rounds. The coordination server performs decay weighting based on the freshness of the parameter updates to ensure the stable convergence of the global model.

[0042] From the perspective of overall system collaboration, the above six modules form a deeply coupled closed-loop processing architecture: the raw data collected by the multimodal data acquisition module 101 serves as the input to the edge computing processing module 102; the structured analysis results output by the edge computing processing module 102 serve as the input to the multimodal fusion analysis module 103; the event description vector generated by the fusion module 103 drives the local knowledge base module 104 to perform intelligent judgment; the judgment results are fed back to law enforcement personnel through the intelligent interaction module 106; and all data is encrypted and stored as evidence by the secure storage module 105. More importantly, the feedback from the back-end modules can influence the operating parameters of the front-end modules in reverse: for example, the judgment results of the local knowledge base module 104 can trigger the edge computing processing module 102 to start a higher-precision recognition model or adjust the region of interest of the detection model; the global model update of the federated learning module directly improves the inference accuracy of the edge computing processing module 102. This feedforward-feedback closed-loop architecture makes the overall performance of the system far exceed the simple summation of the modules working independently.

[0043] like Figure 4 As shown, the present invention also provides a data processing method for a multimodal edge computing law enforcement recorder, which is applied to the law enforcement recorder in the above embodiments. The complete process of the data processing method includes the following steps.

[0044] Step S1: Synchronously acquire multimodal raw data. Specifically, the multimodal data acquisition module 101 simultaneously activates each sensor to acquire data: the high-definition camera array continuously acquires video streams at 30fps and 1080p resolution (it can switch to 4K mode under sufficient lighting conditions); the directional microphone array continuously acquires audio streams at a 16kHz sampling rate; the BeiDou / GPS dual-mode positioning unit outputs latitude and longitude coordinates and elevation information at a frequency of once per second; and the environmental sensor group samples temperature, humidity, and gas concentration data every 10 seconds. All data frames acquired by the sensors are timestamped to the millisecond level by a unified clock source.

[0045] Step S2 involves preprocessing the multimodal raw data. Video preprocessing includes: calculating the pixel difference between adjacent frames using the frame difference method; marking frames with a difference exceeding a preset threshold (e.g., a change rate greater than 15%) as keyframes; reducing subsequent computational load by frame extraction during stable scene periods; and normalizing the keyframes in color space and scaling them to the model input size (640x640 pixels). Audio preprocessing includes: dividing the continuous audio stream into frames with a frame length of 25ms and a frame shift of 10ms; applying a Hamming window function to each frame to reduce spectral leakage; then performing a 1024-point Fast Fourier Transform to extract spectral features and generating 80-dimensional log-Mel spectral features on the Mel frequency scale. Location and environmental data are then converted to a standardized numerical format and written to the preprocessed data buffer queue after outlier filtering (Kalman filtering is used for smoothing location data and track prediction, and median filtering is used for noise suppression and glitch removal of sensor readings).

[0046] Step S3, Edge Intelligent Analysis. The edge computing processing module 102 performs parallel real-time analysis tasks on the preprocessed multimodal data. The CPU dynamically allocates computing tasks to the NPU and FPGA for execution based on the priority of each analysis subtask and the current computing load status through the task scheduler. This step includes several sub-processes: Sub-process S3a is target detection and tracking. The YOLOv8n model performs target detection on each frame of video image, outputting the bounding boxes and category confidence scores of detected targets such as people, vehicles, and objects. The DeepSORT tracking algorithm assigns a persistent unique identification number to each target, enabling continuous tracking across frames. In one embodiment of this invention, the DeepSORT algorithm maintains a tracking pool with a maximum capacity of 200 targets. When a target is not detected for 30 consecutive frames, its tracking number is automatically cancelled. Sub-process S3b is face recognition and comparison. First, the RetinaFace face detection model locates the face region in the video frame. Then, the detected face region is aligned, cropped, and normalized (uniformly scaled to 112x112 pixels). Next, a lightweight ArcFace model is input to extract a 512-dimensional face feature vector. Finally, cosine similarity is calculated with the feature vector in the local face database. When the matching similarity exceeds a threshold of 0.85, the identity is determined to be matched, and the matched person is output. Information; Sub-process S3c is license plate recognition. First, the vehicle area is located using an object detection model. Then, the license plate position is accurately located within the vehicle area. Finally, the CRNN model performs end-to-end sequence recognition of the license plate characters. The recognition results are then correlated with local or remote vehicle information databases for querying. Sub-process S3d is speech recognition and understanding. The Whisper model converts the audio stream into a text sequence. After punctuation restoration and sentence segmentation, the text sequence is processed by a lightweight natural language understanding model based on BERT for keyword extraction and intent classification, recognizing key law enforcement instructions, personal names, place names, and time information in the speech. Sub-process S3e is behavior analysis. The ST-GCN model identifies abnormal behavior patterns based on human skeletal key point sequences. When high-risk behavior (such as fighting or carrying weapons) is detected, a warning signal with the highest priority is immediately generated. Sub-process S3f is environmental perception. Threshold judgment and trend analysis are performed on environmental sensor data. When the CO concentration exceeds 50 ppm or the combustible gas concentration exceeds 20% of the lower explosive limit, an environmental safety warning is triggered.

[0047] Step S4, Multimodal Fusion Analysis. The multimodal fusion analysis module 103 performs deep fusion of the structured analysis results of each modality output in step S3. The detailed process of this step has been fully described in the multimodal fusion analysis module 103 section of the system embodiment, and will not be repeated here. It should be noted that, at the method level, the input of step S4 is the set of structured results output by each sub-process of step S3, including bounding boxes and category information of target detection, identity matching results of face recognition, text transcription content of speech recognition, and action classification labels of behavior analysis, etc. Each result carries a unified timestamp from step S1. The fusion process first aligns the above heterogeneous structured data to a unified timeline based on the timestamp, and then performs feature-level semantic fusion through a cross-modal attention mechanism, finally outputting a comprehensive event description vector with a fixed dimension of 512. This vector encodes the fused semantic representation of multi-dimensional information such as the identity of the person, the target behavior, the environmental state, and the speech content involved in the current law enforcement scenario. In one embodiment of the present invention, when the target detection subprocess in step S3 detects that a person is making a rapid movement, and the behavior analysis subprocess classifies the movement as "running away", and the speech recognition subprocess captures keywords such as "stop" and "don't run" in the same time period, the cross-modal fusion in step S4 will integrate the information of the above three modalities to generate a high-confidence "suspect escape" event description vector. The event type dimension encoding of this vector is significantly biased towards the "emergency pursuit" category, providing accurate and rich scene semantic input for the intelligent judgment in the subsequent step S5.

[0048] Step S5, intelligent judgment based on knowledge base. The local knowledge base module 104 receives the comprehensive event description vector generated in step S4 and executes a three-stage judgment logic: The first stage is historical case matching, which calculates the cosine similarity between the current event description vector and the case embedding vectors pre-stored in the case library, retrieves the top 5 most similar cases, and extracts their handling process and result evaluation as reference. In one embodiment of the present invention, the case retrieval adopts an approximate nearest neighbor search algorithm based on the FAISS library, and the retrieval latency on a case library of 100,000 is controlled within 10ms; The second stage is legal clause association, which classifies the event according to the event type (this classification is based on the event description vector through so... The FTMAX classifier maps to a predefined event type system (including 12 major categories and 56 subcategories such as public security disputes, traffic violations, suspected crimes, and emergency assistance). It queries applicable clauses in the legal database, generates a legal basis summary, and marks key clause numbers and applicable conditions. The third stage involves risk assessment and suggestion generation. It comprehensively considers multiple dimensions such as the number of people on site, the intensity of the behavior, whether dangerous materials are involved, time characteristics (risk coefficient increased by 0.2 at night), and geographical location characteristics (risk coefficient increased by 0.15 in remote areas) to perform a weighted risk score. Based on the score results, it generates corresponding level of handling suggestions. In one embodiment of this invention, for low-risk events (score below 30), a standard handling procedure is recommended, along with suggested law enforcement language. For medium-risk events (score between 30 and 70), it is recommended to send additional support and maintain high vigilance while automatically activating full HD video recording mode. For high-risk events (score above 70), an SOS emergency signal is immediately triggered, real-time location and on-site screenshots are automatically uploaded to the command center, and law enforcement personnel are given voice prompts to pay attention to their own safety.

[0049] Step S6: Real-time feedback and secure upload. The intelligent interaction module 106 provides real-time feedback of the analysis results from step S5 to law enforcement personnel: the voice channel broadcasts key warning information and handling suggestions through a bone conduction speaker, using concise and clear formatted statements, such as "Attention, a person 3m ahead has been identified in the fugitive database, name Zhang XX. It is recommended to verify their identity immediately and request reinforcements." The visualization channel displays augmented reality (AR) style target labeling information on the device screen, including detection boxes, identity tags, risk level color codes (green / yellow / red), and behavioral status descriptions. Simultaneously, the secure storage module 105 encrypts and signs all law enforcement data according to the aforementioned multi-layered protection process before storing it locally. In one embodiment of the present invention, the local storage uses a 256GB high-capacity flash memory chip, which is estimated to store approximately 40 hours of continuous video recording at a 1080p standard bitrate. The simplified analysis results of high-priority events (including event summary text, key screenshots, and risk assessment reports, with a data volume usually not exceeding 500KB) are uploaded to the command center's comprehensive analysis platform in real time through an encrypted channel. The upload process adopts a breakpoint resume mechanism, which automatically caches the data to be transmitted when the network is unstable and retransmits it after the network is restored, ensuring that critical information is not lost.

[0050] Preferably, the method further includes step S7, namely the federated learning collaborative training step. When the law enforcement recorder returns to its charging dock and connects to the public security intranet, the federated learning module automatically starts, using the local data accumulated that day (after differential privacy protection processing) to incrementally train the model. After training is completed, only the model parameter updates are uploaded to the coordination server according to the federated learning protocol, completing one round of collaborative training iteration, thereby achieving continuous collaborative optimization of the network-wide model capabilities while protecting the privacy of local law enforcement data of each device.

[0051] In the above data processing method, each step corresponds strictly to a functional module in the system embodiment: step S1 corresponds to the multimodal data acquisition module 101, steps S2 and S3 correspond to the edge computing processing module 102, step S4 corresponds to the multimodal fusion analysis module 103, step S5 corresponds to the local knowledge base module 104, and step S6 corresponds to the intelligent interaction module 106 and the secure storage module 105.

[0052] Comprehensive Mechanism Analysis: Significant synergistic effects exist among the various technical features of this invention. The heterogeneous computing architecture (CPU+NPU+FPGA) provides the necessary computing power foundation for real-time parallel processing of multimodal data. The multimodal fusion analysis module, based on deep semantic fusion capabilities using a cross-modal attention mechanism, elevates the system's understanding of law enforcement scenarios from fragmented perception of a single modality to a holistic multimodal cognition. This capability relies heavily on the precise spatiotemporal alignment guaranteed by the unified timestamp mechanism. The local knowledge base module further correlates the fused scenario understanding with professional knowledge in the law enforcement field, achieving a cognitive leap from "what to see" to "how to handle it." The secure storage module's four-level protection mechanism—encryption-signature-sharding-on-chain—ensures the immutability of all the aforementioned intelligent analysis results and original data as legal evidence. The federated learning mechanism enables continuous evolution of model capabilities over time, allowing the system's intelligence level to continuously improve with increased deployment scale and usage time. The organic integration of the above-mentioned technical features enables the present invention to achieve full-chain intelligence in perception, cognition, decision-making and evidence storage at the edge, achieving a qualitative breakthrough based on existing technologies.

[0053] To verify the effectiveness of the technical solution of this invention, the following experimental results were obtained in comparative tests with existing technical solutions. In a comprehensive law enforcement scenario test environment, the multimodal fusion analysis module of this invention improved the event type recognition accuracy from 82.1% to 94.6% compared to a solution using only video single-modality, an improvement of 12.5 percentage points. Particularly in noisy environments, when combined with voice modality, the recall rate for violent conflict events increased from 76.3% to 95.1%. Regarding edge-side inference performance, the heterogeneous computing architecture achieved an equivalent comprehensive inference throughput of a traditional solution at 15W power consumption under a 5W power constraint, extending battery life to more than 2.5 times that of the traditional solution. In terms of data security, the data protection link combining SM4 encryption and ECDSA signature underwent 100,000 simulated attack tests (including data tampering, replay attacks, and man-in-the-middle attacks), successfully detecting and blocking all attacks, with a zero false negative rate for data integrity verification. Regarding the federated learning effect, the model performance after 20 rounds of collaborative training on 50 devices is close to the ideal value of training on the entire dataset (the difference is less than 1.5 percentage points), while achieving the privacy protection goal of zero out-of-domain data from the original data. The above experimental results show that the collaborative work of the various technical features of this invention produces a nonlinear synergistic effect that is significantly better than the simple superposition of the technical features.

[0054] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of the claims of the present invention.

Claims

1. A multimodal edge computing law enforcement recorder, characterized in that, include: The multimodal data acquisition module is used to simultaneously acquire video streams, audio streams, satellite positioning information, and environmental sensor data from law enforcement sites, and to add a unified timestamp with millisecond-level precision to each acquired data stream based on a unified clock source, generating a time-aligned multimodal raw data stream; The edge computing processing module adopts a heterogeneous computing architecture consisting of a high-performance processor, a neural network acceleration chip, and a reconfigurable logic unit. It is used to perform real-time target detection and tracking, facial feature extraction and comparison, speech recognition and semantic understanding, and human behavior recognition on the multimodal raw data stream at the device end, and output structured analysis results for each modality. The multimodal fusion analysis module, based on a cross-modal attention mechanism, is used to perform spatiotemporal alignment and deep semantic fusion on the structured analysis results of different modalities, generating a comprehensive event description vector that encodes multidimensional scene information; The local knowledge base module stores a knowledge graph of law enforcement, a legal and regulatory database, and a historical case database. It is used to perform similar case retrieval and legal clause association based on the comprehensive event description vector, and generate hierarchical law enforcement assessment suggestions. The secure storage module employs national cryptographic algorithms and a blockchain hash anchoring mechanism to sequentially perform symmetric encryption, hash digest calculation, digital signature based on a secure chip, and erasure coding for redundant fragmented storage of law enforcement data. The intelligent interaction module is used to provide real-time feedback of the law enforcement assessment suggestions to law enforcement personnel in voice or visual form, and to interact with the command center through an encrypted channel.

2. The multimodal edge computing law enforcement recorder according to claim 1, characterized in that, The multimodal data acquisition module includes: a high-definition camera array, comprising a front-facing main camera and a wide-angle auxiliary camera, for acquiring video of law enforcement scenes; a directional microphone array, which uses beamforming technology to achieve directional sound pickup and noise suppression; a dual-mode positioning unit, providing sub-meter level positioning accuracy; and an environmental sensor group, including a temperature and humidity sensor and a gas detection sensor.

3. The multimodal edge computing law enforcement recorder according to claim 1, characterized in that, In the edge computing processing module, the high-performance processor is used for task scheduling and general computing, the neural network acceleration chip is used for parallel inference acceleration of deep learning models, and the reconfigurable logic unit is used for hardware acceleration of customized algorithms. The three components achieve data interaction and collaborative computing through a shared memory bus.

4. The multimodal edge computing law enforcement recorder according to claim 1, characterized in that, The edge computing processing module integrates a group of deep learning models that have undergone model compression processing, which includes at least one of knowledge distillation, weight quantization, and structured pruning.

5. The multimodal edge computing law enforcement recorder according to claim 1, characterized in that, The multimodal fusion analysis module employs a multimodal encoder based on the Transformer architecture. The fusion process includes: semantically encoding the video feature sequence and the audio feature sequence respectively; establishing temporal correlations between different modal features through spatiotemporal alignment; calculating the correlation weights between different modal features through a cross-modal attention layer; and performing feature weighted fusion based on the correlation weights to generate a multimodal representation vector.

6. The multimodal edge computing law enforcement recorder according to claim 5, characterized in that, The multimodal fusion analysis module also includes a dynamic weight adjustment mechanism, which is configured to dynamically adjust the weight contribution of each mode in the fusion process based on the signal quality score of each mode data. When the signal quality score of any mode is lower than a preset threshold, the fusion weight of that mode is reduced.

7. The multimodal edge computing law enforcement recorder according to claim 1, characterized in that, The secure storage module employs a multi-layered data protection link, including: encrypting the original data using the SM4 national cryptographic algorithm; calculating the SHA-256 hash digest of the encrypted data; performing ECDSA digital signature on the hash digest using the device private key stored in the secure chip; storing the signed data in fragments and generating redundant check fragments through erasure coding algorithm; and periodically synchronizing the data hash digest and signature information to the blockchain network to generate a trusted timestamp.

8. The multimodal edge computing law enforcement recorder according to claim 1, characterized in that, The local knowledge base module includes a law enforcement knowledge graph, a legal and regulatory database, and a historical case database. The process of generating law enforcement assessment suggestions includes: matching the comprehensive event description vector with the case vectors in the historical case database based on similarity; querying applicable legal and regulatory provisions according to the event type; and generating graded disposal suggestions based on on-site risk assessment.

9. The multimodal edge computing law enforcement recorder according to claim 1, characterized in that, The law enforcement recorder also includes a federated learning module, configured to collaboratively train a shared model among multiple law enforcement recorder devices via a secure aggregation protocol, with each device only uploading model parameter updates to the coordination server without sharing the original data.

10. A data processing method for a multimodal edge computing law enforcement recorder, applied to the multimodal edge computing law enforcement recorder according to any one of claims 1 to 9, characterized in that, Includes the following steps: Step S1: Simultaneously collect video streams, audio streams, location information, and environmental sensor data from the law enforcement scene through the multimodal data acquisition module, and attach a unified timestamp to each collected data stream to generate a multimodal raw data stream; Step S2: Preprocess the multimodal raw data stream, including video keyframe extraction, audio frame segmentation and windowing, and data standardization. Step S3: Input the preprocessed multimodal data into the edge computing processing module, and perform real-time intelligent analysis using a pre-deployed deep learning model to obtain structured analysis results for each modality. Step S4: Perform cross-modal spatiotemporal correlation and semantic fusion on the structured analysis results obtained in Step S3 using the multimodal fusion analysis module to generate a comprehensive event description vector. Step S5: Intelligently assess the comprehensive event description vector based on law enforcement domain knowledge in the local knowledge base module to generate law enforcement suggestions and handling plans. Step S6: Feed back the analysis results and law enforcement suggestions to law enforcement personnel in real time through the intelligent interaction module, and store the key data after encryption and signing by the secure storage module.