Campus safety analysis and early warning method and system fusing multi-modal reasoning capabilities

CN120599772BActive Publication Date: 2025-11-21SHANDONG GUOSHU DEV CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511113087.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-11-21
Estimated Expiration
2045-08-11

AI Technical Summary

Technical Problem

现有校园安全分析方法存在模态割裂、心理健康预警滞后、隐私与效率冲突及因果推理能力缺失的问题,导致检测精度低、识别延迟长、隐私泄露和事件预测不足。

Method used

采用多模态数据采集、轻量化处理和隐私化处理结合云端DeepSeek引擎分析,利用数学建模和逻辑推理加速器进行风险评估和警报触发,实现多模态数据的关联和因果推理。

Benefits of technology

提高了复杂场景下的检测精度,减少了心理危机识别延迟,保障了隐私并提供了高效的因果预测能力,支持校园安全的差异化处置。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120599772B_ABST
    Figure CN120599772B_ABST
Patent Text Reader

Abstract

The application provides a campus safety analysis and early warning method and system fusing multi-modal reasoning capability, and belongs to the technical field of intelligent campus safety; comprising multi-modal data collection in a campus; sequentially performing data processing, lightweight processing and privacy processing on the multi-modal data; analyzing the multi-modal data subjected to the lightweight processing and the privacy processing by using a cloud DeepSeek engine, and outputting a quantitative risk assessment level; triggering a three-level alarm according to the obtained risk assessment level, and pushing a differentiated disposal suggestion for the campus safety. The application can guarantee the privacy of personnel in the campus and the safety analysis efficiency, and take into account the causal reasoning capability; meanwhile, the problems of split modal data and lagging psychological health early warning are avoided, and better support and guarantee can be provided for the campus safety.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of smart campus security technology, and in particular relates to a campus security analysis and early warning method and system that integrates multimodal reasoning capabilities. Background Technology

[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.

[0003] With the rapid development of technology, campus security is undergoing a profound transformation from traditional management models to intelligent and information-based management. In this transformation, the application of a series of technologies has become crucial for improving campus security. Among them, video surveillance technology, as the foundation of campus security, has achieved a leapfrog development from analog signals to digital high-definition, and then to intelligent analysis. Intelligent video surveillance systems can not only capture images of the campus in real time, but also automatically detect and issue warnings for abnormal behavior through advanced technologies such as facial recognition and behavior recognition, effectively reducing the incidence of campus security incidents.

[0004] However, existing campus security analysis methods or systems still generally have some technical problems, such as:

[0005] (1) Modal fragmentation and shallow analysis: On the one hand, most existing technologies rely on a single visual model (such as YOLO) or independent sensors, which cannot associate the semantic relationships of multimodal data such as video, audio, and text, resulting in a high false negative rate in complex scenes and a sharp drop in detection accuracy in scenes such as low light, target occlusion, and dynamic interference; on the other hand, they can only identify "whether there is a person", but cannot understand the behavioral intention.

[0006] (2) Delayed early warning of mental health: Traditional psychological assessment relies on manual observation or questionnaire surveys, which cannot quantify physiological signals such as micro-expressions and voice frequency jitter, resulting in a delay of more than 24 hours in identifying severe psychological crises and low accuracy of teachers' interpretation of students' behavior.

[0007] (3) Privacy vs. Efficiency Conflict: Existing technologies generally require centralized processing of multimodal data, with biometric features such as faces and voiceprints being directly transmitted to the cloud. Furthermore, existing solutions struggle to balance the capabilities of large models with the low-latency requirements of campus scenarios. On one hand, centralized cloud computing suffers from high latency (pure cloud architecture results in response times exceeding 2 seconds), failing to meet emergency response requirements. Edge devices, on the other hand, have limited computing power; for example, lightweight models (with parameters less than 1B) cannot handle complex inference. On the other hand, centralized processing easily leaks biometric features (such as faces and voiceprints), while excessive desensitization can lead to the loss of feature information.

[0008] (4) Lack of causal reasoning ability: Existing technologies can only report phenomena (such as "crowd gathering"), but cannot predict the evolution path of events (such as the cause of gathering and the probability of conflict). Summary of the Invention

[0009] To overcome the shortcomings of the existing technologies, this invention provides a campus security analysis and early warning method and system that integrates multimodal reasoning capabilities. This method can ensure the privacy of personnel on campus and the efficiency of security analysis while taking into account causal reasoning capabilities. At the same time, it avoids the problems of fragmented modal data and delayed mental health early warning, and can provide better support and guarantee for campus security.

[0010] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions:

[0011] The first aspect of this invention provides a campus security analysis and early warning method that integrates multimodal reasoning capabilities.

[0012] A campus security analysis and early warning method integrating multimodal reasoning capabilities includes:

[0013] Multimodal data collection on campus;

[0014] Perform data processing on multimodal data;

[0015] The processed multimodal data is then subjected to lightweight and privacy-preserving processing.

[0016] The cloud-based DeepSeek engine is used to analyze multimodal data that has undergone lightweight and privacy processing, and outputs a quantitative risk assessment level; wherein, the cloud-based DeepSeek engine includes a mathematical modeling engine and a logical reasoning accelerator.

[0017] Based on the obtained risk assessment level, a Level 3 alarm is triggered, and differentiated handling suggestions for campus safety are pushed out.

[0018] Furthermore, the multimodal data includes video data, voice data, and text data from social media platforms within the campus monitoring area.

[0019] Furthermore, multimodal data processing is performed, including: identifying behavioral actions in the acquired video data based on spatiotemporal graph convolutional networks; separating environmental noise and quantifying speech aggression in the acquired speech data based on directional microphone arrays and voiceprint filtering algorithms; and monitoring social media platform keywords in the acquired text data in real time based on web crawling technology, and combining this with DeepSeek-NLP to identify metaphorical risk content.

[0020] Furthermore, the multimodal data is subjected to lightweight and privacy processing, including: first, the multimodal data is fed into a lightweight feature encoder, and the multimodal data is reduced in dimensionality based on the lightweight feature encoder; then, the multimodal data after dimensionality reduction is privacy-processed based on the edge of the federated learning framework.

[0021] Furthermore, when using the cloud-based DeepSeek engine to analyze multimodal data, the mathematical modeling engine is used to construct psychological resilience differential equations to quantify the psychological risk index; the logical reasoning accelerator integrates a recursive causal discovery algorithm to construct an event evolution network.

[0022] Furthermore, the three-level alarm includes blue, yellow, and red alarms, with each color representing a different alarm severity level. Specifically, red alarm > yellow alarm > blue alarm.

[0023] A second aspect of the present invention provides a campus security analysis and early warning system that integrates multimodal reasoning capabilities.

[0024] A campus security analysis and early warning system integrating multimodal reasoning capabilities includes:

[0025] The equipment layer is used for multimodal data acquisition within the campus.

[0026] The multimodal sensing layer is used for: data processing of multimodal data;

[0027] The edge computing layer is used for: lightweighting and privacy-preserving multimodal data after data processing;

[0028] The cloud-based DeepSeek engine is used to analyze multimodal data that has undergone lightweight and privacy-enhancing processing, and output a quantitative risk assessment level.

[0029] The execution and feedback layer is used to: trigger a level-three alarm based on the obtained risk assessment level, and push differentiated handling suggestions for campus safety.

[0030] Furthermore, the device layer includes cameras, directional microphones, and IoT devices;

[0031] The multimodal perception layer includes visual units, auditory units, and text units;

[0032] The edge computing layer includes a lightweight feature encoder and a privacy protection module;

[0033] The cloud-based DeepSeek engine includes a mathematical modeling engine and a logical reasoning accelerator.

[0034] The execution and feedback layer is configured with a multi-level early warning mechanism.

[0035] A third aspect of the present invention provides a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the steps of a campus security analysis and early warning method integrating multimodal reasoning capabilities as described in the first aspect of the present invention.

[0036] The fourth aspect of the present invention provides an electronic device, including a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of a campus security analysis and early warning method integrating multimodal reasoning capabilities as described in the first aspect of the present invention.

[0037] The above one or more technical solutions have the following beneficial effects:

[0038] (1) This invention utilizes the cloud-based DeepSeek engine to analyze multimodal data after lightweight and privacy processing, and outputs a quantitative risk assessment level. The multimodal data includes video data, audio data, and text data from campus social platforms within the campus monitoring area, rather than being limited to a single visual model or independent sensor. Therefore, this invention can simultaneously correlate the semantic relationships of multimodal data such as video, audio, and text, significantly reducing the false negative rate in complex scenarios.

[0039] (2) This invention utilizes a cloud-based DeepSeek engine to analyze multimodal data that has undergone lightweight and privacy processing, and outputs a quantitative risk assessment level. The cloud-based DeepSeek engine includes a mathematical modeling engine and a logical reasoning accelerator, rather than relying on traditional manual observation or questionnaires. Based on the deeply integrated cloud-based DeepSeek engine, this invention can effectively quantify physiological signals such as micro-expressions and voice fundamental frequency jitter, thereby effectively reducing the delay in identifying severe psychological crises.

[0040] (3) In this invention, multimodal data is collected and processed separately rather than centrally. Then, the cloud-based DeepSeek engine is used to analyze the lightweight and privacy-enhancing multimodal data, outputting a quantitative risk assessment level. The entire multimodal data processing process is privacy-enhancing without excessive anonymization. Federated learning combined with edge detection technology retains more feature information while ensuring biometric privacy. Therefore, this invention can simultaneously balance the privacy of personnel on campus with efficiency in security analysis.

[0041] (4) This invention combines DeepSeek’s mathematical reasoning ability with the campus scenario, and realizes the quantifiable prediction of risk through psychological resilience differential equations and group behavior game models. At the same time, it locates the root cause of the event through recursive causal discovery, so as to predict the evolution path of the event and generate intervention plans with mathematical basis, rather than being limited to reporting phenomena.

[0042] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0043] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0044] Figure 1 This is a flowchart of a campus security analysis and early warning method that integrates multimodal reasoning capabilities, as described in Embodiment 1 of the present invention.

[0045] Figure 2 This is a structural diagram of a campus security analysis and early warning system that integrates multimodal reasoning capabilities, as shown in Embodiment 2 of the present invention. Detailed Implementation

[0046] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0047] It should be noted that the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the exemplary implementations of the present invention.

[0048] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.

[0049] Example 1

[0050] This embodiment discloses a campus security analysis and early warning method that integrates multimodal reasoning capabilities.

[0051] like Figure 1 As shown, a campus security analysis and early warning method integrating multimodal reasoning capabilities includes:

[0052] Step S1: Multimodal data collection within the campus;

[0053] Step S2: Process the multimodal data;

[0054] Step S3: Perform lightweight and privacy-preserving processing on the multimodal data after data processing;

[0055] Step S4: Analyze the multimodal data after lightweight and privacy processing using the cloud-based DeepSeek engine, and output a quantitative risk assessment level; wherein, the cloud-based DeepSeek engine includes a mathematical modeling engine and a logical reasoning accelerator.

[0056] Step S5: Trigger a Level 3 alarm based on the obtained risk assessment level and push differentiated handling suggestions for campus safety.

[0057] Through the above process, this invention can ensure the privacy and security of personnel on campus while also taking into account causal reasoning capabilities; at the same time, it avoids the problems of modal data fragmentation and delayed mental health early warning, and can provide better support and guarantee for campus safety. To facilitate understanding of the technical solution of this invention, the specific implementation steps of the technical solution of this invention will be further explained and described below.

[0058] Step S1: Multimodal data collection within the campus.

[0059] In this embodiment, the multimodal data to be collected includes video data, voice data, and text data from social media platforms within the campus monitoring area.

[0060] For video data within the campus monitoring area: Multiple wide-angle cameras (supporting infrared illumination) can be deployed on campus for video data collection. As an optional implementation, the cameras can be installed in key areas such as campus entrances and exits, main roads, teaching buildings, dormitories, libraries, gymnasiums, and canteens to ensure comprehensive monitoring without blind spots; at the same time, the cameras can be connected to a terminal computer via wired, wireless, or other methods.

[0061] For voice data within the campus monitoring area: Multiple directional microphones can be deployed on campus for voice data collection. Before deploying directional microphones, the monitoring objective must be clearly defined, such as monitoring sounds related to campus safety. This helps determine the specific location and coverage area of ​​the microphones and allows for the selection of appropriate directional microphone equipment based on monitoring needs. For example, when covering a large area, a microphone with a longer pickup distance can be selected; when high-definition recording is required, a high-fidelity device can be selected. As an optional embodiment, the directional microphone can be installed in a location compatible with the camera, or it can be connected to the terminal computer via wired or wireless means. Ultimately, both video and voice data are transmitted to the terminal computer as video and voice streams, respectively. It should be noted that the voice data captured by the camera during video streaming can also be used directly as voice data without setting up a separate directional microphone. Therefore, as long as video and voice data within the required monitoring area can be effectively collected, this embodiment does not specifically limit the type of equipment or location used for collection.

[0062] For text data on campus social media platforms: web scraping technology can be used to monitor keywords on relevant campus social media platforms in real time, such as special words related to campus safety.

[0063] Step S2: Perform data processing on the multimodal data.

[0064] In this embodiment, different types of data in the multimodal data are processed separately, namely: using the Spatiotemporal Graph Convolutional Network (ST-GCN) to identify actions in the acquired video data; using a directional microphone array and voiceprint filtering algorithm to separate environmental noise in the acquired speech data and quantify speech aggression (e.g., baseband jitter, sudden changes in speech rate); and using web crawling technology to monitor social media platform keywords in the acquired text data in real time, and combining this with DeepSeek-NLP to identify metaphorical risk content.

[0065] The Spatiotemporal Graph Convolutional Network (ST-GCN) is used to identify actions in acquired video data. Specifically:

[0066] 1) The network architecture of the Spatiotemporal Graph Convolutional Network (ST-GCN) includes an input layer, a spatiotemporal graph, multi-scale convolutional blocks, and an output layer. Specifically:

[0067] Input Layer: After decoding the video stream, a 30fps frame sequence is generated. A lightweight model extracts the coordinates of 25 key points on the human body, outputting a skeleton data tensor. The spatiotemporal graph includes a spatial graph and a temporal graph. The spatial graph uses an adjacency matrix defined based on human body topology, with joint edge weights dynamically adjusted through learnable parameters. The temporal graph uses a sliding window to construct a temporal connection matrix to capture action continuity. Multi-scale convolutional blocks: Eight ST-GCN modules are stacked, each containing spatial graph convolution, temporal convolution, residual connections, and BatchNorm to accelerate convergence. Output Layer: After global average pooling, the spatiotemporal features are output through a fully connected layer, showing the probability P∈[0,1] of actions affecting campus safety.

[0068] 2) The optimization strategies of the spatiotemporal graph convolutional network ST-GCN include occlusion compensation and multimodal verification.

[0069] Specifically, occlusion compensation: By introducing an LSTM network to predict the trajectory of occluded key points, the recognition accuracy in occluded scenarios is improved. Multimodal verification: When P>0.7, the audio analysis module is triggered to verify the existence of keywords related to campus safety behaviors.

[0070] 3) The implementation steps for recognizing actions in acquired video data based on the Spatiotemporal Graph Convolutional Network (ST-GCN) include: video stream preprocessing and keyframe extraction, spatiotemporal graph construction and dynamic feature extraction, ST-GCN model forward inference, post-processing and multimodal validation, and model training and optimization. Specifically:

[0071] 3-1) Video stream preprocessing and keyframe extraction, including video slicing and multi-person scene segmentation.

[0072] Video slicing: The input video stream is sliced ​​into fixed time windows (default 1 second, 30 frames), and 1 frame (key frame) is extracted using FFmpeg to reduce redundant calculations; background subtraction algorithm is used to eliminate interference from lighting changes.

[0073] Multi-person scene segmentation: YOLOv8 is used to detect human bounding boxes in the scene, and the skeleton sequence is extracted separately for each detected target (confidence > 0.6); if the number of people detected is ≥ 3 and the bounding box overlap rate (IoU) is > 30%, the group behavior analysis mode is triggered to generate a multi-person spatiotemporal map.

[0074] 3-2) Spatiotemporal graph construction and dynamic feature extraction, including human skeleton extraction and spatiotemporal graph construction.

[0075] Human skeleton extraction: A lightweight model is used to extract the coordinates (x, y, confidence) of 25 human key points in real time and output the skeleton sequence; for key points with confidence < 0.5, their trajectories are predicted by an LSTM network.

[0076] Spatiotemporal graph construction: An initial adjacency matrix is ​​defined based on human topological connections, and a learnable weight matrix is ​​introduced to dynamically adjust edge weights; a temporal adjacency matrix is ​​constructed, and the inter-frame similarity is calculated using a Gaussian kernel function.

[0077] 3-3) Forward inference of the ST-GCN model, including spatiotemporal graph convolution calculation, multi-head spatiotemporal attention mechanism and classification output.

[0078] Spatiotemporal graph convolution computation: Each ST-GCN module contains spatial graph convolution and temporal convolution; skip connections are added to the output of each layer to alleviate gradient vanishing.

[0079] Multi-head spatiotemporal attention mechanism: An attention module is introduced in the last layer to calculate the Query, Key, and Value matrices.

[0080] Classification output: Spatiotemporal features are input into a fully connected layer after global average pooling, and the output behavior probability P is generated.

[0081] 3-4) Post-processing and multimodal verification, including logic rule enhancement and multimodal cross-validation.

[0082] Enhanced logical rules: Define a physical constraint rule base to correct the model output, such as pushing actions, detecting sudden changes in upper limb joint speed and center of gravity shift, containment behavior, and calculating the containment radius r. If r < 1 meter and lasts for more than 3 seconds, it is judged as containment.

[0083] Multimodal cross-validation: When P > 0.7, the audio analysis module is triggered to verify whether there are keywords related to campus safety within the corresponding time period; if the multimodal evidence is consistent, it is confirmed as a campus safety incident; otherwise, it enters the manual review queue.

[0084] 3-5) Model training and optimization, including loss function design and two-stage training strategy.

[0085] Loss function design: Focal Loss is used to solve the class imbalance problem.

[0086] Two-stage training strategy: In the pre-training stage, the basic ST-GCN model is trained on an open-source dataset; in the fine-tuning stage, transfer learning is performed using a self-built campus behavior dataset.

[0087] Based on directional microphone arrays and voiceprint filtering algorithms, environmental noise in the collected speech data is separated and the aggressiveness of the speech is quantified. Specifically:

[0088] 1) Noise separation and speech enhancement, including beamforming and voiceprint filtering.

[0089] 1-1) Beamforming: An algorithm is used to spatially filter the data from the 6-channel microphone array to improve the signal-to-noise ratio.

[0090] 1-2) Voiceprint filtering: Use an x-vector-based speaker recognition model to separate the speech streams of different speakers.

[0091] 2) Quantization of aggressive speech, including acoustic feature extraction and deep scoring model quantization scoring.

[0092] 2-1) Acoustic feature extraction: Calculate parameters such as fundamental frequency jitter, harmonic-to-noise ratio, and speech rate abrupt changes for each frame of speech.

[0093] 2-2) Quantitative scoring of deep scoring models: Input features into a pre-trained model and output an aggression score.

[0094] Based on web crawling technology, real-time monitoring of social media platform keywords in collected text data is conducted, and combined with DeepSeek-NLP to identify metaphorical risk content, specifically:

[0095] 1) Real-time monitoring methods for social platforms based on web crawler technology include distributed web crawler architecture design, dynamic keyword database construction, and real-time alarm mechanism.

[0096] 1-1) Distributed crawler architecture design, including multi-platform adapters and incremental crawling strategy design.

[0097] Multi-platform adapter: Deploy API crawling modules for campus social platforms such as Weibo and QQ Space, and adopt differentiated request strategies.

[0098] Incremental crawling strategy: Based on text fingerprint deduplication, only newly added content is crawled.

[0099] 1-2) Construct a dynamic keyword library based on a basic thesaurus and semantic extensions.

[0100] Basic vocabulary: includes a range of behaviors and actions related to campus safety.

[0101] Semantic expansion: Based on Word2Vec, a word vector space is generated to semantically expand sensitive words.

[0102] 1-3) Design a real-time alarm mechanism. As an optional implementation, the real-time alarm mechanism can adopt the multi-level triggering rules shown in Table 1, namely:

[0103] Table 1 Multi-level triggering mechanism

[0104]

[0105] 2) The core architecture and features of DeepSeek-NLP are based on model architecture and training strategies. The model architecture embeds attention optimization features on the basis of the basic framework. The training strategies include multi-task pre-training and adversarial training.

[0106] Regarding the model architecture: The basic framework adopts a multimodal deep Transformer structure and a dynamic vocabulary mechanism to support mixed Chinese and English text processing; the attention optimization is to add a local attention window on the basis of traditional multi-head attention to capture local metaphor features.

[0107] Regarding training strategies: Multi-task pre-training involves joint optimization of campus psychological data, with the main task being metaphor recognition and the auxiliary task being emotional polarity classification; adversarial training improves robustness by adding FGM perturbations.

[0108] 3) The implementation steps for identifying metaphorical risk content include text preprocessing and feature enhancement, deep metaphor detection, multi-granularity risk fusion, and dynamic knowledge base interaction.

[0109] 3-1) Text preprocessing and feature enhancement are achieved based on special symbol processing and dependency parsing.

[0110] For example, special character processing involves converting emoticons and internet slang into semantic tags (e.g., "T_T" is converted to "crying"). Dependency parsing can use tools to extract rhetorical structures. See Table 2 for details:

[0111] Table 2. Examples of Syntactic Analysis

[0112]

[0113] 3-2) Deep metaphor detection is achieved through context encoding and metaphor feature extraction.

[0114] For example, context encoding involves converting the input text into a vector sequence H through an embedding layer, and then generating a contextual representation through a Transformer encoder; metaphor feature extraction can use an attention head to specifically capture metaphorical patterns.

[0115] 3-3) Achieving multi-granularity risk fusion based on sentence-level risk and document-level risk; whereby sentence-level risk refers to risk fusion for each sentence. Calculate risk value Document-level risks are analyzed in conjunction with the document structure.

[0116] 3-4) Dynamic knowledge base interaction is achieved based on the rumination learning mechanism and the synchronous sensitive word database.

[0117] Rumination learning mechanism: The model is automatically updated for false positives and false negatives, as shown in Table 3. If the manual review result is not equal to the model prediction result, adversarial examples are generated, added to the training set, and the Transformer parameters are fine-tuned.

[0118] Table 3 Rumination Learning Mechanism

[0119]

[0120] Sensitive word database synchronization: High-frequency metaphorical words are extracted from the review logs every week and added to the keyword database, and the weight is updated.

[0121] Step S3: Perform lightweight and privacy processing on the multimodal data after data processing.

[0122] First, the multimodal data is fed into a lightweight feature encoder. The lightweight feature encoder then performs dimensionality reduction on the multimodal data, specifically by compressing the multimodal data into low-dimensional vectors (video). Spacetime token, audio Mel spectrograms are used to reduce transmission bandwidth requirements. Subsequently, the multimodal data after dimensionality reduction is processed for privacy at the edge of the federated learning framework. Specifically, under the federated learning framework, face blurring and voiceprint anonymization are performed at the edge to ensure that the original data does not leave the school.

[0123] Dimensionality reduction of multimodal data is performed based on a lightweight feature encoder, specifically:

[0124] 1) Video data is compressed into a spatiotemporal token.

[0125] 1-1) The input video stream is sampled at 10fps, and the coordinates of 25 key points of the human body are extracted using a lightweight model to generate a skeleton sequence.

[0126] 1-2) Spatiotemporal attention compression: A spatiotemporal separated Transformer encoder is adopted. Spatial attention is combined to calculate the correlation weight between key points within the same frame, and temporal attention is combined to calculate the motion evolution weight between consecutive frames. The weighted features are compressed into a 128-dimensional vector through a fully connected layer.

[0127] 2) Spatiotemporal tokens represent the spatiotemporal evolution pattern of actions in video data. For example, the spatial dimension represents the trajectory of the hand joints in the "waving" action; the temporal dimension represents the continuous inter-frame motion acceleration encoded in the "running" action.

[0128] 3) The speech data is compressed into a Mel spectrogram.

[0129] 3-1) Preprocessing: Pre-emphasis, applying a first-order high-pass filter; framing and windowing, with a frame length of 25ms and a frame shift of 10ms, using a Hamming window function; calculating the short-time FFT energy spectrum based on Fourier transform.

[0130] 3-2) Mel filtering: Mapping to the Mel scale through a triangular filter bank.

[0131] 3-3) Lightweight encoding: Use CNN to compress the Mel spectrum into a 64-dimensional vector.

[0132] Privacy-preserving processing is performed on multimodal data that has undergone dimensionality reduction at the edge, based on a federated learning framework. Specifically:

[0133] 1) Federated Learning Framework Architecture. This embodiment adopts a layered federated framework, specifically: a central server, responsible for global model aggregation, is deployed in the campus data center. Edge nodes are deployed in devices in various teaching buildings and dormitories, including: local model libraries, data anonymization engines, and secure communication proxies. Communication protocol: the uplink is for model parameters and anonymization features, and the downlink is for global model updates.

[0134] 2) Edge-end face blurring (differential privacy noise injection), including: detecting face region R, applying Gaussian noise matrix to the detected region, and adaptively adjusting noise intensity.

[0135] 3) Edge-end voiceprint anonymization (voiceprint feature decoupling and replacement), including: using a model to separate speech content features C from speaker features; irreversibly desensitizing speaker features, i.e.: feature replacement and adversarial generation; reconstructing anonymous speech to ensure that speaker features cannot be recovered through reverse engineering.

[0136] Step S4: Analyze the multimodal data after lightweight and privacy processing using the cloud-based DeepSeek engine, and output a quantitative risk assessment level; the cloud-based DeepSeek engine includes a mathematical modeling engine and a logical reasoning accelerator.

[0137] When using the DeepSeek engine in the cloud to analyze multimodal data, the mathematical modeling engine is used to construct psychological resilience differential equations to quantify psychological risk indices; the logical reasoning accelerator integrates a recursive causal discovery algorithm to construct event evolution networks.

[0138] DeepSeek-Math is an existing open-source mathematical reasoning model focused on solving and reasoning about complex mathematical problems. The mathematical modeling engine uses DeepSeek-Math to construct psychological resilience differential equations to quantify the psychological risk index (MPRI), i.e.:

[0139] ;

[0140] in, Indicates psychological resilience value. Indicates the strength of external support. Indicates the amount of stress stimulus. This represents the attenuation coefficient obtained by fitting campus psychological data.

[0141] The logic reasoning accelerator employs the DeepSeek-R1 recursive causal discovery algorithm to construct an event evolution network (such as "family atmosphere"). Insomnia Decreased classroom attention The self-denial method specifically includes: 1) constructing an initial causal graph with nodes representing multimodal features (such as aggressive speech); 2) pruning redundant edges through conditional independence testing; 3) recursively applying the model to solve for causal directions; and 4) outputting the event evolution path.

[0142] Furthermore, based on the psychological risk index and event evolution network, a quantitative risk assessment level is output, specifically: MPRI (Psychological Risk Index). max(100-R(t), 0); causal path length L The number of hops from the root node to the endpoint in the evolutionary network. Based on the above design, the risk assessment level is determined. The hierarchical rules can be expressed as:

[0143] ;

[0144] Step S5: Trigger a Level 3 alarm based on the obtained risk assessment level and push differentiated handling suggestions for campus safety.

[0145] The three-level alarm system includes blue, yellow, and red alerts, each color representing a different level of severity. Specifically, red alert > yellow alert > blue alert. Based on the risk assessment level, differentiated response suggestions for campus security (such as psychological intervention scripts, access control instructions, etc.) are provided, as shown in Table 4.

[0146] Table 4. Recommendations for Tiered Handling

[0147]

[0148] Based on the above process, this invention realizes a mathematically driven multimodal fusion, namely: visual data is converted into spatial coordinate parameters, audio data is modeled through sound wave differential equations, and logical predicates (such as Push(A,B)) are extracted from text data to construct a cross-modal mathematical representation space; at the same time, the strategy utility of behavioral participants is calculated based on game theory to predict the behavior diffusion path and block key nodes.

[0149] Furthermore, the multimodal fusion process of cross-modal mathematical representation space and game theory model includes:

[0150] 1) Visual data is converted into spatial coordinate parameters, and the coordinates of the joints of the human skeleton are extracted using the model to construct motion differential equations.

[0151] 2) The audio data is used to model the sound wave equation, and the sound signal s(t) is mapped to the solution of the wave equation.

[0152] 3) Logical predicates are extracted from the text data, and predicate logic is generated through AMR parsing, as shown in Table 5:

[0153] Table 5 Examples of predicates generated by AMR parsing

[0154]

[0155] Based on the above process, this invention also achieves edge-cloud collaborative optimization, namely: dynamically allocating computing resources using the DeepSeek-MoE architecture: lightweight experts at the edge (<1B parameters) handle routine detection, while deep experts in the cloud (7B-70B parameters) perform complex inference; response latency ≤300ms, which is 5 times better than a pure cloud solution.

[0156] Furthermore, the DeepSeek-MoE edge-cloud collaborative architecture adopts a dynamic routing mechanism, that is: data flow is distributed to the expert network through the controller; the architecture design of the DeepSeek-MoE edge-cloud collaborative architecture includes 1) lightweight experts at the edge: real-time face blurring, voiceprint desensitization, and basic behavior recognition; 2) deep experts in the cloud: psychological risk modeling, causal reasoning, and multimodal game simulation.

[0157] In summary, the campus safety analysis and early warning method provided by this invention integrates multimodal reasoning capabilities. It converts multimodal data into quantitative parameters using a mathematical modeling engine, constructs an event causal graph using a recursive causal discovery algorithm, and triggers tiered responses and optimizes handling strategies based on risk levels. It is particularly suitable for applications such as intelligent intervention of campus safety behaviors and dynamic monitoring of mental health. Specifically, in intelligent intervention of campus safety behaviors, it can integrate ST-GCN action recognition, voice aggression scoring, and social isolation index to calculate a behavioral risk index and trigger tiered responses (such as anonymous prompts and encrypted evidence push notifications). In dynamic monitoring of mental health, it can identify high-risk cases affecting campus safety 72 hours in advance through micro-expression Z-score standardization and psychological resilience decay curve prediction, with a false alarm rate ≤8%.

[0158] Compared to existing technologies, this invention makes many innovative improvements, such as mathematically enhanced vertical domain adaptation, privacy-efficiency balance design, and causal chain-driven proactive defense. Specifically, regarding mathematically enhanced vertical domain adaptation, this invention combines DeepSeek's mathematical reasoning capabilities with campus scenarios, achieving quantifiable risk prediction through psychological resilience differential equations and group behavior game models. Regarding the privacy-efficiency balance design, this invention combines federated learning and edge detection techniques to retain nearly 95% of feature information while ensuring biometric privacy. Regarding causal chain-driven proactive defense, this invention locates the root cause of events through recursive causal discovery to generate mathematically grounded intervention plans.

[0159] This invention, through systematic innovation in multimodal large models, mathematically enhanced reasoning, and hierarchical computing architecture, can overcome the three major technical challenges that have long existed in the field of campus security: "unclear visibility, inaccurate judgment, and slow response," and achieve a paradigm upgrade from passive monitoring to proactive defense.

[0160] Example 2

[0161] This embodiment discloses a campus security analysis and early warning system that integrates multimodal reasoning capabilities.

[0162] like Figure 2 As shown, a campus security analysis and early warning system integrating multimodal reasoning capabilities includes:

[0163] The equipment layer is used for multimodal data acquisition within the campus.

[0164] The multimodal sensing layer is used for: data processing of multimodal data;

[0165] The edge computing layer is used for: lightweighting and privacy-preserving multimodal data after data processing;

[0166] The cloud-based DeepSeek engine is used to analyze multimodal data that has undergone lightweight and privacy-enhancing processing, and output a quantitative risk assessment level.

[0167] The execution and feedback layer is used to: trigger a level-three alarm based on the obtained risk assessment level, and push differentiated handling suggestions for campus safety.

[0168] Furthermore, the device layer includes cameras, directional microphones, and IoT devices; all data collected by the device layer is sent to the terminal computer. The multimodal perception layer includes visual units, auditory units, and text units; the edge computing layer includes a lightweight feature encoder and a privacy protection module; the cloud-based DeepSeek engine includes a mathematical modeling engine and a logical reasoning accelerator; the execution and feedback layer is equipped with a multi-level early warning mechanism and supports linkage with drone tracking, access control, and psychological intervention script generation.

[0169] Furthermore, the data interaction methods between levels include:

[0170] 1) The interaction mechanism between the device layer and the multimodal perception layer is as follows: the camera is directly connected to the vision unit through the CoaXPress 2.0 interface; the directional microphone is connected to the hearing unit through the AES67 audio network protocol; and the IoT devices are aggregated to the text unit through the LoRaWAN network management system.

[0171] 2) The interaction between the multimodal perception layer and the edge computing layer is as follows: cross-modal data are interconnected through PCIe Switch; the visual unit outputs the coordinates of the skeleton joints; and the auditory unit transmits the voiceprint hash value and emotion tag.

[0172] 3) The interaction between the edge computing layer and the cloud-based DeepSeek engine is as follows: lightweight data is uploaded via an IPSec VPN tunnel.

[0173] 4) The interaction between the cloud and the execution and feedback layer is as follows: control commands are issued through 5G URLLC; priority mapping: red alerts use NS3 level dedicated channels.

[0174] Example 3

[0175] The purpose of this embodiment is to provide a computer-readable storage medium.

[0176] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of a campus security analysis and early warning method integrating multimodal reasoning capabilities as described in Embodiment 1 of this disclosure.

[0177] Example 4

[0178] The purpose of this embodiment is to provide an electronic device.

[0179] An electronic device includes a memory, a processor, and a program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps in a campus security analysis and early warning method integrating multimodal reasoning capabilities as described in Embodiment 1 of this disclosure.

[0180] The steps and methods involved in the apparatuses of Embodiments 2, 3, and 4 above correspond to those in Embodiment 1. For specific implementation details, please refer to the relevant description section of Embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood as including any medium capable of storing, encoding, or carrying an instruction set for execution by a processor and enabling the processor to perform any of the methods in this invention.

[0181] Those skilled in the art will understand that the modules or steps of the present invention described above can be implemented using general-purpose computer devices. Optionally, they can be implemented using computer-executable program code, thereby allowing them to be stored in a storage device for execution by a computer device, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. The present invention is not limited to any particular combination of hardware and software.

[0182] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.

Claims

1. A campus security analysis and early warning method integrating multimodal reasoning capabilities, characterized in that, include: Multimodal data collection on campus; Perform data processing on multimodal data; The processed multimodal data is then subjected to lightweight and privacy-preserving processing. The DeepSeek cloud engine is used to analyze multimodal data after lightweight and privacy-enhancing processing, and outputs a quantitative risk assessment level. The DeepSeek cloud engine includes a mathematical modeling engine and a logical reasoning accelerator. When analyzing multimodal data using the DeepSeek cloud engine, the mathematical modeling engine is used to construct a psychological resilience differential equation to quantify the psychological risk index. The logical reasoning accelerator integrates a recursive causal discovery algorithm for constructing an event evolution network. Specifically, the mathematical modeling engine constructs the psychological resilience differential equation based on DeepSeek-Math, i.e.: ; in, Indicates psychological resilience value. Indicates the strength of external support. Indicates the amount of stress stimulus. This represents the attenuation coefficient after fitting the campus psychological data; The logic reasoning accelerator employs the DeepSeek-R1 recursive causal discovery algorithm to construct an event evolution network. Specifically, this includes: constructing an initial causal graph with nodes featuring multimodal characteristics; pruning redundant edges through conditional independence testing; recursively applying the model to solve for causal directions; and outputting the event evolution path. Finally, based on the obtained risk assessment level, a Level 3 alarm is triggered, and differentiated handling suggestions for campus safety are pushed out.

2. The campus security analysis and early warning method integrating multimodal reasoning capabilities as described in claim 1, characterized in that, The multimodal data includes video data, voice data, and text data from social media platforms within the campus monitoring area.

3. A campus security analysis and early warning method integrating multimodal reasoning capabilities as described in any one of claims 1-2, characterized in that, Data processing for multimodal data includes: identifying behavioral actions in acquired video data based on spatiotemporal graph convolutional networks; separating environmental noise and quantifying speech aggression in acquired speech data based on directional microphone arrays and voiceprint filtering algorithms; and real-time monitoring of social media keywords in acquired text data based on web crawling technology, combined with DeepSeek-NLP to identify metaphorical risk content.

4. The campus security analysis and early warning method integrating multimodal reasoning capabilities as described in claim 1, characterized in that, The process of lightweighting and privacy processing of multimodal data includes: first, feeding the multimodal data into a lightweight feature encoder and performing dimensionality reduction on the multimodal data based on the lightweight feature encoder; then, performing privacy processing on the dimensionality-reduced multimodal data based on the edge of a federated learning framework.

5. The campus security analysis and early warning method integrating multimodal reasoning capabilities as described in claim 1, characterized in that, The three-level alarm system includes blue, yellow, and red alarms, with each color representing a different level of alarm severity. Specifically, red alarm > yellow alarm > blue alarm.

6. A campus security analysis and early warning system integrating multimodal reasoning capabilities, characterized in that, include: The equipment layer is used for multimodal data acquisition within the campus. The multimodal sensing layer is used for: data processing of multimodal data; The edge computing layer is used for: lightweighting and privacy-preserving multimodal data after data processing; The cloud-based DeepSeek engine is used to analyze multimodal data after lightweight and privacy-enhancing processing, and output a quantitative risk assessment level. The cloud-based DeepSeek engine includes a mathematical modeling engine and a logical reasoning accelerator. When analyzing multimodal data using the cloud-based DeepSeek engine, the mathematical modeling engine is used to construct psychological resilience differential equations to quantify the psychological risk index. The logical reasoning accelerator integrates a recursive causal discovery algorithm for constructing an event evolution network. Specifically, the mathematical modeling engine constructs psychological resilience differential equations based on DeepSeek-Math, i.e.: ; in, Indicates psychological resilience value. Indicates the strength of external support. Indicates the amount of stress stimulus. This represents the attenuation coefficient after fitting the campus psychological data; The logic reasoning accelerator employs the DeepSeek-R1 recursive causal discovery algorithm to construct an event evolution network. Specifically, this includes: constructing an initial causal graph with nodes featuring multimodal characteristics; pruning redundant edges through conditional independence testing; recursively applying the model to solve for causal directions; and outputting the event evolution path. The execution and feedback layer is used to: trigger a level-three alarm based on the obtained risk assessment level, and push differentiated handling suggestions for campus safety.

7. A campus security analysis and early warning system integrating multimodal reasoning capabilities as described in claim 6, characterized in that, include: The device layer includes cameras, directional microphones, and Internet of Things (IoT) devices; The multimodal perception layer includes visual units, auditory units, and text units; The edge computing layer includes a lightweight feature encoder and a privacy protection module; The cloud-based DeepSeek engine includes a mathematical modeling engine and a logical reasoning accelerator. The execution and feedback layer is configured with a multi-level early warning mechanism.

8. A computer-readable storage medium having a program stored thereon, characterized in that, When executed by the processor, the program implements the steps of the campus security analysis and early warning method that integrates multimodal reasoning capabilities as described in any one of claims 1-5.

9. An electronic device, comprising a memory, a processor, and a program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the campus security analysis and early warning method that integrates multimodal reasoning capabilities as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Campus bullying prevention system based on voice recognition and big data analysis

    CN119229593A

  • Lightweight pedestrian falling detection method and system for inspection robot

    CN120148115A