Campus safety analysis early warning method and system fused with multi-modal reasoning capability

Through the cloud-based DeepSeek engine and multimodal data processing technology, the modal split, mental health warning lag, and privacy and efficiency conflicts in campus safety analysis have been resolved, and effective correlation and timely causal reasoning of multimodal data have been achieved, thereby improving the accuracy and efficiency of campus safety analysis.

CN120599772AActive Publication Date: 2025-09-05SHANDONG GUOSHU DEV CO LTD

Patent Information

Application Number
CN202511113087.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-09-05
Estimated Expiration
2045-08-11

AI Technical Summary

Technical Problem

Existing campus safety analysis methods have problems such as modal fragmentation, delayed mental health warning, conflict between privacy and efficiency, and lack of causal reasoning ability. They are unable to effectively process multimodal data and provide timely safety support.

Method used

The cloud-based DeepSeek engine is used in combination with multimodal data collection, lightweight processing and privacy processing, and mathematical modeling and logical reasoning accelerators are used for analysis to output quantitative risk assessment levels and trigger three-level alarms and differentiated disposal recommendations.

Benefits of technology

It achieves effective correlation of multimodal data, reduces missed detection rates, reduces delays in psychological crisis identification, protects privacy, and provides timely causal reasoning capabilities, thereby improving the efficiency and accuracy of campus safety analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120599772A_ABST
    Figure CN120599772A_ABST
Patent Text Reader

Abstract

The invention provides a campus safety analysis and early warning method and system fused with a multi-modal reasoning capability, and belongs to the technical field of smart campus safety. Comprising the steps of collecting multi-modal data in a campus; sequentially carrying out data processing, lightweight processing and privacy processing on the multi-modal data; a cloud DeepSeek engine is utilized to analyze the multi-modal data subjected to the lightweight processing and the privacy processing, and a quantitative risk assessment level is output; and triggering a third-level alarm according to the obtained risk assessment level, and pushing differentiated disposal suggestions for campus security. According to the invention, the causal reasoning ability can be considered while the privacy and security analysis efficiency of people in the campus are ensured; meanwhile, the problems of modal data splitting and mental health early warning lag are avoided, and better support and guarantee can be provided for campus safety.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of smart campus safety technology, and in particular relates to a campus safety analysis and early warning method and system integrating multimodal reasoning capabilities. Background Art

[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.

[0003] With the rapid development of science and technology, campus security is undergoing a profound transformation from traditional management models to intelligent, information-based management. In this transformation, the application of a range of technologies has become crucial for improving campus safety. Among them, video surveillance technology, the foundation of campus security, has achieved a remarkable evolution from analog signals to digital high-definition and then to intelligent analysis. Intelligent video surveillance systems not only capture real-time footage on campus but also automatically detect and issue warnings for abnormal behavior through advanced technologies such as facial and behavioral recognition, effectively reducing the incidence of campus safety incidents.

[0004] However, existing campus security analysis methods or systems still have some common technical problems, such as: (1) Modal splitting and shallow analysis: On the one hand, existing technologies mostly rely on a single visual model (such as YOLO) or independent sensors, and are unable to associate the semantic relationships of multimodal data such as video, audio, and text, resulting in a high missed detection rate in complex scenes and a sharp drop in detection accuracy in scenes such as low light, target occlusion, and dynamic interference. On the other hand, they can only identify "whether there is a person" but cannot understand the behavioral intention.

[0005] (2) Delayed early warning of mental health: Traditional psychological assessments rely on manual observation or questionnaires and are unable to quantify physiological signals such as micro-expressions and voice fundamental frequency jitter, resulting in a delay of more than 24 hours in identifying severe psychological crises and low accuracy in teachers’ interpretation of student behavior.

[0006] (3) Conflict between privacy and efficiency: Most existing technologies require the centralized processing of multimodal data, with biometric features such as faces and voiceprints transmitted directly to the cloud. Furthermore, existing solutions struggle to balance the capabilities of large models with the low latency requirements of campus scenarios. On the one hand, centralized cloud computing has high latency (a pure cloud architecture results in response times greater than 2 seconds), which cannot meet emergency response requirements. Edge devices, on the other hand, have computing power limitations, such as lightweight models (with parameters less than 1B) that cannot handle complex reasoning. On the other hand, centralized processing can easily leak biometric features (such as faces and voiceprints), while excessive desensitization can lead to the loss of feature information.

[0007] (4) Lack of causal reasoning ability: Existing technologies can only report phenomena (such as “crowd gathering”) but cannot predict the evolution path of events (such as the causes of gathering and the probability of conflict). Summary of the Invention

[0008] To overcome the shortcomings of the above-mentioned existing technologies, the present invention provides a campus safety analysis and early warning method and system that integrates multimodal reasoning capabilities, which can ensure the privacy of people on campus and the efficiency of security analysis while taking into account causal reasoning capabilities; at the same time, it avoids the problems of modal data fragmentation and delayed psychological health warnings, and can provide better support and guarantees for campus safety.

[0009] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions: A first aspect of the present invention provides a campus safety analysis and early warning method integrating multimodal reasoning capabilities.

[0010] A campus safety analysis and early warning method integrating multimodal reasoning capabilities includes: Multimodal data collection on campus; Perform data processing on multimodal data; Lightweight and privacy-enhanced multimodal data after data processing; Utilize the cloud-based DeepSeek engine, which includes a mathematical modeling engine and a logic reasoning accelerator, to analyze the lightweight and privacy-protected multimodal data and output a quantitative risk assessment level. A level 3 alert is triggered based on the resulting risk assessment level, and differentiated handling recommendations for campus safety are pushed.

[0011] Furthermore, the multimodal data includes video data, voice data within the campus monitoring range, and text data on the campus social platform.

[0012] Furthermore, multimodal data is processed, including: identifying behavioral actions in the collected video data based on a spatiotemporal graph convolutional network; separating environmental noise from the collected voice data and quantifying voice aggressiveness based on a directional microphone array and voiceprint filtering algorithm; monitoring social platform keywords in the collected text data in real time based on crawler technology, and combining DeepSeek-NLP to identify metaphorical risk content.

[0013] Furthermore, multimodal data is lightweighted and privacy-enhanced, including: first, feeding the multimodal data into a lightweight feature encoder, and performing dimensionality reduction on the multimodal data based on the lightweight feature encoder; then, privacy-enhancing the multimodal data that has undergone dimensionality reduction based on the edge end under the federated learning framework.

[0014] Furthermore, when using the cloud-based DeepSeek engine to analyze multimodal data, the mathematical modeling engine is used to construct a psychological resilience differential equation to quantify the psychological risk index; the logical reasoning accelerator is integrated with a recursive causal discovery algorithm to construct an event evolution network.

[0015] Furthermore, the three-level alarm includes blue, yellow and red alarms, and each color represents a different alarm severity. Specifically, red alarm>yellow alarm>blue alarm.

[0016] A second aspect of the present invention provides a campus safety analysis and early warning system that integrates multimodal reasoning capabilities.

[0017] A campus safety analysis and early warning system that integrates multimodal reasoning capabilities includes: The equipment layer is used for multimodal data collection on campus; Multimodal perception layer, used to: process multimodal data; The edge computing layer is used to: perform lightweight processing and privacy protection on the multimodal data after data processing; The cloud-based DeepSeek engine is used to analyze lightweight and privacy-protected multimodal data and output a quantitative risk assessment level. The execution and feedback layer is used to trigger a level 3 alert based on the resulting risk assessment level and deliver differentiated response recommendations for campus safety.

[0018] Furthermore, the device layer includes cameras, directional microphones, and IoT devices; The multimodal perception layer includes a visual unit, an auditory unit, and a text unit; The edge computing layer includes a lightweight feature encoder and a privacy protection module; The cloud-based DeepSeek engine includes a mathematical modeling engine and a logical reasoning accelerator; A multi-level early warning mechanism is configured in the execution and feedback layer. The third aspect of the present invention provides a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the steps of a campus safety analysis and early warning method integrating multimodal reasoning capabilities as described in the first aspect of the present invention.

[0019] The fourth aspect of the present invention provides an electronic device, including a memory, a processor, and a program stored in the memory and runnable on the processor. When the processor executes the program, it implements the steps of a campus safety analysis and early warning method integrating multimodal reasoning capabilities as described in the first aspect of the present invention.

[0020] One or more of the above technical solutions have the following beneficial effects: (1) This invention uses the cloud-based DeepSeek engine to analyze lightweight and privacy-protected multimodal data and output a quantitative risk assessment level. Multimodal data includes video data, voice data, and text data from on-campus social platforms within the campus monitoring range, rather than being limited to a single visual model or independent sensor. Therefore, this invention can simultaneously correlate the semantic relationships among multimodal data, including video, audio, and text, significantly reducing the missed detection rate in complex scenarios.

[0021] (2) The present invention utilizes the cloud-based DeepSeek engine to analyze lightweight and privacy-protected multimodal data and output a quantitative risk assessment level. The cloud-based DeepSeek engine includes a mathematical modeling engine and a logical reasoning accelerator, rather than relying on traditional manual observation or questionnaires. Based on the deeply integrated cloud-based DeepSeek engine, the present invention can effectively quantify physiological signals such as micro-expressions and voice fundamental frequency jitter, thereby effectively reducing the delay in identifying severe psychological crises.

[0022] (3) When collecting data, the present invention collects and processes multimodal data separately, rather than centrally processing them. Subsequently, the cloud-based DeepSeek engine is used to analyze the lightweight and privacy-enhanced multimodal data and output a quantitative risk assessment level. The entire multimodal data processing process is privacy-enhanced without being overly desensitized. Federated learning is combined with edge desensitization technology to retain more feature information while ensuring biometric privacy. Therefore, the present invention can simultaneously take into account the privacy of campus personnel and the efficiency of security analysis.

[0023] (4) This invention combines DeepSeek's mathematical reasoning capabilities with campus scenarios, and achieves quantifiable risk prediction through the psychological resilience differential equation and group behavior game model. At the same time, it locates the root cause of the event through recursive causal discovery to predict the event evolution path and generate intervention plans with mathematical basis, rather than just reporting phenomena.

[0024] Advantages of additional aspects of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0026] Figure 1 This is a flowchart of a campus safety analysis and early warning method integrating multimodal reasoning capabilities in Example 1 of the present invention.

[0027] Figure 2 This is a structural diagram of a campus safety analysis and early warning system that integrates multimodal reasoning capabilities in Example 2 of the present invention. DETAILED DESCRIPTION

[0028] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.

[0029] It should be noted that the terms used herein are for describing particular embodiments only and are not intended to limit the exemplary embodiments according to the present invention.

[0030] In the absence of conflict, the embodiments of the present invention and the features thereof may be combined with each other.

[0031] Example 1 This embodiment discloses a campus safety analysis and early warning method that integrates multimodal reasoning capabilities.

[0032] like Figure 1 As shown, a campus safety analysis and early warning method integrating multimodal reasoning capabilities includes: Step S1, multimodal data collection on campus; Step S2: processing the multimodal data; Step S3: Lightweight and privacy-enhancing the multimodal data after data processing; Step S4: Analyze the lightweight and privacy-protected multimodal data using a cloud-based DeepSeek engine, and output a quantitative risk assessment level; wherein the cloud-based DeepSeek engine includes a mathematical modeling engine and a logical reasoning accelerator; Step S5: Trigger a level 3 alert based on the resulting risk assessment level, and push differentiated handling suggestions for campus safety.

[0033] Through the above process, the present invention can ensure the privacy of people on campus and the efficiency of security analysis while taking into account causal reasoning capabilities. At the same time, it avoids the problems of modal data fragmentation and delayed mental health warning, and can provide better support and guarantee for campus safety. To facilitate understanding of the technical solution of the present invention, the specific implementation steps of the technical solution of the present invention are further explained and illustrated below.

[0034] Step S1: Collect multimodal data on campus.

[0035] In this embodiment, the multimodal data that needs to be collected includes video data, voice data within the campus monitoring range, and text data on the campus social platform.

[0036] For video data within the campus's surveillance area, multiple wide-angle cameras (supporting infrared fill light) can be deployed throughout the campus to collect video data. As an optional implementation, cameras can be installed at key locations such as campus entrances and exits, main roads, teaching buildings, dormitories, libraries, gymnasiums, and cafeterias to ensure comprehensive monitoring. Furthermore, the cameras can be connected to terminal computers via various methods, including wired and wireless connections.

[0037] Regarding voice data within the campus monitoring area: Multiple directional microphones can be deployed throughout the campus to collect voice data. Before deploying directional microphones, the monitoring objective must be clearly defined, such as monitoring sounds related to campus safety. This helps determine the specific microphone location and coverage area, and allows appropriate directional microphone equipment to be selected based on the monitoring requirements. For example, if coverage of a larger area is required, a microphone with a longer pickup range can be selected; if high-definition recording is required, a device with high fidelity can be selected. As an optional embodiment, the directional microphones can be installed in a location compatible with the camera or connected to the terminal computer via any wired or wireless method. Ultimately, both video and voice data are transmitted to the terminal computer in the form of video and voice streams. It should be noted that the voice data can also be directly collected by the camera during the video stream, rather than by separately deploying directional microphones. Therefore, as long as the video and voice data within the required collection area can be effectively collected, the specific device and location used for collection are not specifically limited in this embodiment.

[0038] For text data on campus social platforms: crawler technology can be used to monitor keywords on relevant campus social platforms in real time, such as special words related to campus safety.

[0039] Step S2: Process the multimodal data.

[0040] In this embodiment, different types of data in the multimodal data are processed separately, namely: based on the spatiotemporal graph convolutional network ST-GCN, behavioral actions in the collected video data are identified; based on the directional microphone array and voiceprint filtering algorithm, environmental noise in the collected voice data is separated and the voice aggressiveness (such as fundamental frequency jitter, sudden changes in speaking rate, etc.) is quantified; based on crawler technology, social platform keywords in the collected text data are monitored in real time, and DeepSeek-NLP is combined to identify metaphorical risk content.

[0041] Based on the spatiotemporal graph convolutional network ST-GCN, we identify the behavioral actions in the collected video data. Specifically: 1) The network architecture design of the spatiotemporal graph convolutional network ST-GCN includes the input layer, spatiotemporal graph, multi-scale convolutional blocks and output layer. Specifically: Input layer: After decoding, the video stream generates a 30fps frame sequence. The coordinates of 25 key points of the human body are extracted through a lightweight model, and the skeleton data tensor is output. The spatiotemporal graph includes a spatial graph and a temporal graph. Among them, the spatial graph defines the adjacency matrix based on the topological connection of the human body, and the joint edge weights are dynamically adjusted through learnable parameters. The temporal graph uses a sliding window to construct a temporal connection matrix to capture the continuity of the action. Multi-scale convolution block: There are 8 stacked ST-GCN modules, each of which contains spatial graph convolution, temporal convolution, residual connection and BatchNorm to accelerate convergence. Output layer: After the spatiotemporal features are globally averaged and pooled, the probability of behavioral actions that affect campus safety P∈[0,1] is output through a fully connected layer.

[0042] 2) The optimization strategies of the spatiotemporal graph convolutional network ST-GCN include occlusion compensation and multimodal verification.

[0043] Specifically, occlusion compensation: By introducing an LSTM network to predict the trajectory of occluded joints, recognition accuracy in occluded scenarios is improved. Multimodal verification: When P > 0.7, the audio analysis module is triggered to verify the presence of keywords related to campus safety behavior.

[0044] 3) The implementation steps for identifying behavioral actions in collected video data based on the spatiotemporal graph convolutional network (ST-GCN) include: video stream preprocessing and keyframe extraction, spatiotemporal graph construction and dynamic feature extraction, ST-GCN model forward reasoning, post-processing and multimodal verification, and model training and optimization. Specifically: 3-1) Video stream preprocessing and key frame extraction, including video slicing and multi-person scene segmentation.

[0045] Video slicing: The input video stream is sliced ​​according to a fixed time window (default 1 second, 30 frames), and one frame (key frame) is extracted through FFmpeg to reduce redundant calculations; a background subtraction algorithm is used to eliminate interference from lighting changes.

[0046] Multi-person scene segmentation: Use YOLOv8 to detect human bounding boxes in the picture and extract a skeleton sequence for each detected target (confidence level > 0.6). If ≥ 3 people are detected and the bounding box overlap rate (IoU) is > 30%, the group behavior analysis mode is triggered to generate a multi-person spatiotemporal map.

[0047] 3-2) Spatiotemporal graph construction and dynamic feature extraction, including human skeleton extraction and spatiotemporal graph construction.

[0048] Human skeleton extraction: A lightweight model is used to extract the coordinates (x, y, confidence) of 25 key points of the human body in real time, and the skeleton sequence is output. For joint points with a confidence level less than 0.5, the trajectory is predicted using an LSTM network.

[0049] Spatiotemporal graph construction: The initial adjacency matrix is ​​defined based on the topological connections of the human body, and a learnable weight matrix is ​​introduced to dynamically adjust the edge weights; a temporal adjacency matrix is ​​constructed, and a Gaussian kernel function is used to calculate the similarity between frames.

[0050] 3-3) Forward reasoning of the ST-GCN model, including spatiotemporal graph convolution calculation, multi-head spatiotemporal attention mechanism and classification output.

[0051] Spatiotemporal graph convolution calculation: Each ST-GCN module contains spatial graph convolution and temporal convolution; skip connections are added to the output of each layer to alleviate gradient disappearance.

[0052] Multi-head spatiotemporal attention mechanism: Introduce the attention module in the last layer to calculate the Query, Key, and Value matrices.

[0053] Classification output: The spatiotemporal features are globally averaged and pooled before being input into the fully connected layer to output the behavior probability P.

[0054] 3-4) Post-processing and multimodal validation, including logical rule enhancement and multimodal cross-validation.

[0055] Logical rule enhancements: Define a physical constraint rule base to modify model outputs, such as detecting sudden changes in upper limb joint velocity and center of gravity shifts in pushing and shoving actions, and calculating the radius r of the encirclement. If r is less than 1 meter and persists for more than 3 seconds, it is considered an encirclement.

[0056] Multimodal cross-validation: When P>0.7, the audio analysis module is triggered to check whether there are keywords related to campus safety in the corresponding time period; if the multimodal evidence is consistent, it is confirmed as a campus safety incident; otherwise, it enters the manual review queue.

[0057] 3-5) Model training and optimization, including loss function design and two-stage training strategy.

[0058] Loss function design: Focal Loss is used to solve the problem of category imbalance.

[0059] Two-stage training strategy: In the pre-training stage, the basic ST-GCN model is trained on an open source dataset; in the fine-tuning stage, transfer learning is performed using a self-built campus behavior dataset.

[0060] Based on a directional microphone array and voiceprint filtering algorithm, we can separate the ambient noise from the collected voice data and quantify the voice aggressiveness. Specifically: 1) Noise separation and speech enhancement, including beamforming and voiceprint filtering.

[0061] 1-1) Beamforming: An algorithm is used to spatially filter the 6-channel microphone array data to improve the signal-to-noise ratio.

[0062] 1-2) Voiceprint filtering: Use an x-vector-based speaker recognition model to separate the voice streams of different speakers.

[0063] 2) Aggressive speech quantification, including acoustic feature extraction and deep scoring model quantification scoring.

[0064] 2-1) Acoustic feature extraction: Calculate parameters such as fundamental frequency jitter, harmonic-to-noise ratio, and speech rate mutation for each frame of speech.

[0065] 2-2) Deep Scoring Model Quantification: Input features into the pre-trained model and output the aggressiveness score.

[0066] We use crawler technology to monitor social platform keywords in collected text data in real time, and combine it with DeepSeek-NLP to identify metaphorical risk content. Specifically: 1) The real-time monitoring method of social platforms based on crawler technology includes the design of distributed crawler architecture, the construction of dynamic keyword library and real-time alarm mechanism.

[0067] 1-1) Distributed crawler architecture design, including multi-platform adapters and incremental crawling strategy design.

[0068] Multi-platform adapter: Deploy API crawling modules for campus social platforms such as Weibo and QQ Space, and adopt differentiated request strategies.

[0069] Incremental crawling strategy: based on text fingerprint deduplication, only crawl new content, 1-2) Build a dynamic keyword library based on the basic vocabulary and semantic extensions.

[0070] Basic vocabulary: includes a series of behaviors and actions that affect campus safety.

[0071] Semantic expansion: Generate word vector space based on Word2Vec and perform semantic expansion on sensitive words.

[0072] 1-3) Design a real-time alarm mechanism. As an optional embodiment, the real-time alarm mechanism can adopt the multi-level triggering rules shown in Table 1, namely: Table 1 Multi-level trigger mechanism

[0073] 2) DeepSeek-NLP's core architecture and features are implemented based on the model architecture and training strategy. The model architecture embeds attention optimization features on the basis of the basic framework; the training strategy includes multi-task pre-training and adversarial training.

[0074] Regarding the model architecture: The basic framework adopts a multimodal deep Transformer structure and a dynamic vocabulary mechanism to support mixed Chinese and English text processing; attention optimization adds a local attention window on the basis of traditional multi-head attention to capture local metaphor features.

[0075] Regarding the training strategy: Multi-task pre-training is a joint optimization of campus psychological data. Its main task is metaphor recognition, and its auxiliary task is sentiment polarity classification. Adversarial training improves robustness by adding FGM perturbations.

[0076] 3) Implementation steps for metaphorical risk content identification, including text preprocessing and feature enhancement, deep metaphor detection, multi-granularity risk fusion, and dynamic knowledge base interaction.

[0077] 3-1) Text preprocessing and feature enhancement are achieved based on special symbol processing and dependency syntax parsing.

[0078] For example, special symbol processing involves converting emoticons and Martian text into semantic labels (e.g., "T_T" is converted to "crying"). Dependency parsing tools can be used to extract rhetorical structures. See Table 2 for details: Table 2 Syntactic analysis examples

[0079] 3-2) Deep metaphor detection is achieved through context encoding and metaphor feature extraction.

[0080] For example, context encoding is the process of converting the input text into a vector sequence H through the Embedding layer, and then generating a context representation through the Transformer encoder; metaphor feature extraction can use the attention head to specifically capture metaphor patterns.

[0081] 3-3) Implement multi-granularity risk fusion based on sentence-level risk and document-level risk; sentence-level risk is the risk of each sentence. Calculating Value at Risk ; Document-level risks are analyzed in combination with the chapter structure.

[0082] 3-4) Dynamic knowledge base interaction is achieved based on the ruminative learning mechanism and the synchronization of sensitive vocabulary.

[0083] Rumination learning mechanism: False positives and false negatives automatically trigger model updates, as shown in Table 3. If the manual review result is not equal to the model prediction result, adversarial examples are generated and added to the training set, and the Transformer parameters are fine-tuned.

[0084] Table 3 Rumination learning mechanism

[0085] Sensitive word library synchronization: extract high-frequency metaphor words from the audit log every week and add them to the keyword library, and update the weights.

[0086] Step S3: Lightweight and privacy-enhanced processing is performed on the multimodal data after data processing.

[0087] First, the multimodal data is fed into a lightweight feature encoder, and the dimensionality reduction operation is performed on the multimodal data based on the lightweight feature encoder, that is, by compressing the multimodal data into a low-dimensional vector (video Space-time token, audio Mel-spectrograms) to reduce transmission bandwidth requirements. Subsequently, privacy processing is performed on the multimodal data after dimensionality reduction at the edge within the federated learning framework. Specifically, face blurring and voiceprint anonymization are performed at the edge within the federated learning framework to ensure that the original data remains calibrated.

[0088] Perform dimensionality reduction on multimodal data based on lightweight feature encoders. Specifically: 1) Video data is compressed into spatiotemporal tokens.

[0089] 1-1) The input video stream is sampled at 10fps, and a lightweight model is used to extract the coordinates of 25 key points of the human body to generate a skeleton sequence.

[0090] 1-2) Spatiotemporal Attention Compression: A spatiotemporal separation Transformer encoder is used. Spatial attention is combined to calculate the association weights between joints within the same frame, and temporal attention is combined to calculate the action evolution weights between consecutive frames. The weighted features are compressed into a 128-dimensional vector through a fully connected layer.

[0091] 2) Spatiotemporal Tokens represent the spatiotemporal evolution of actions in video data. For example, the spatial dimension captures the trajectory of hand joints in a “waving” action, while the temporal dimension encodes the acceleration of motion between consecutive frames in a “running” action.

[0092] 3) Speech data is compressed into Mel-spectrogram.

[0093] 3-1) Preprocessing: Pre-emphasis, applying a first-order high-pass filter; frame division and windowing, with a 25ms frame length, a 10ms frame shift, and a Hamming window function; calculation of the short-time FFT energy spectrum based on the Fourier transform.

[0094] 3-2) Mel filtering: Mapping to Mel scale through triangular filter bank.

[0095] 3-3) Lightweight encoding: Use CNN to compress the Mel spectrum into a 64-dimensional vector.

[0096] Based on the edge of the federated learning framework, privacy processing is performed on multimodal data that has undergone dimensionality reduction operations. Specifically: 1) Federated Learning Framework Architecture. This embodiment adopts a hierarchical federated framework. Specifically, the central server, responsible for global model aggregation, is deployed in the campus data center. Edge nodes, deployed in equipment in each teaching building and dormitory, contain a local model library, a data desensitization engine, and a secure communication agent. The communication protocol: upstream transmits model parameters and desensitization features, and downstream transmits global model updates.

[0097] 2) Edge-side face blurring (differential privacy noise injection), including: detecting the face region R, applying a Gaussian noise matrix to the detected region, and adaptively adjusting the noise intensity.

[0098] 3) Edge voiceprint anonymization (voiceprint feature decoupling and replacement), including: using a model to separate the speech content feature C from the speaker feature; irreversibly desensitizing the speaker feature, namely: feature replacement and adversarial generation; and reconstructing the anonymous speech to ensure that the speaker feature cannot be recovered through reverse engineering.

[0099] Step S4: Use the cloud-based DeepSeek engine to analyze the lightweight and privacy-protected multimodal data and output a quantitative risk assessment level; wherein the cloud-based DeepSeek engine includes a mathematical modeling engine and a logical reasoning accelerator.

[0100] When using the cloud-based DeepSeek engine to analyze multimodal data, the mathematical modeling engine is used to construct a psychological resilience differential equation to quantify the psychological risk index; the logical reasoning accelerator is integrated with a recursive causal discovery algorithm to construct an event evolution network.

[0101] DeepSeek-Math is an existing open-source mathematical reasoning model that focuses on solving and reasoning complex mathematical problems. The mathematical modeling engine builds a psychological resilience differential equation based on DeepSeek-Math to quantify the psychological risk index (MPRI), which is: ; in, Represents the psychological resilience value, Indicates the strength of external support, Indicates the amount of pressure stimulation, Represents the attenuation coefficient fitted by campus psychological data.

[0102] The logical reasoning accelerator uses the DeepSeek-R1 recursive causal discovery algorithm to build event evolution networks (such as "family atmosphere Insomnia Decreased attention span in class The method involves: 1) constructing an initial causal graph with nodes being multimodal features (e.g., aggressive speech); 2) pruning redundant edges through conditional independence testing; 3) recursively applying the model to solve for causal directions; and 4) outputting the event evolution path.

[0103] Furthermore, based on the psychological risk index and event evolution network, a quantitative risk assessment level is output, specifically: MPRI (Psychological Risk Index) max(100-R(t),0); causal path length L The number of hops from the root node to the end point in the evolving network. Based on the above design, the risk assessment level The classification rules can be expressed as: ; Step S5: Trigger a level 3 alert based on the resulting risk assessment level, and push differentiated handling suggestions for campus safety.

[0104] The three-level alert includes blue, yellow, and red. Each color represents a different level of severity. Specifically, red alert > yellow alert > blue alert. Differentiated campus safety measures (such as psychological intervention and access control instructions) are pushed based on the risk assessment level. The details are shown in Table 4: Table 4: Push of graded disposal suggestions

[0105] Based on the above process, the present invention realizes multimodal fusion driven by mathematics, namely: visual data is converted into spatial coordinate parameters, audio data is modeled through acoustic wave differential equations, and logical predicates (such as Push(A,B)) are extracted from text data to construct a cross-modal mathematical representation space; at the same time, the strategic utility of the behavior participants is calculated based on game theory to predict the behavior diffusion path and block key nodes.

[0106] Furthermore, the multimodal fusion process of the cross-modal mathematical representation space and the game theory model includes: 1) Convert visual data into spatial coordinate parameters, use the model to extract the coordinates of the human skeleton joints, and construct the motion differential equation.

[0107] 2) Audio data is modeled using the wave equation, mapping the sound signal s(t) into a solution to the wave equation.

[0108] 3) Logical predicates are extracted from text data and predicate logic is generated through AMR parsing, as shown in Table 5: Table 5 Examples of predicates generated by AMR parsing

[0109] Based on the above process, the present invention also realizes edge-cloud collaborative optimization, namely: using the DeepSeek-MoE architecture to dynamically allocate computing resources: lightweight experts on the edge (<1B parameters) handle routine detection, and deep experts on the cloud (7B-70B parameters) perform complex reasoning; the response delay is ≤300ms, which is 5 times higher than the pure cloud solution.

[0110] Furthermore, the DeepSeek-MoE edge-cloud collaborative architecture adopts a dynamic routing mechanism, namely: data flows are distributed to the expert network through the controller; the architectural design of the DeepSeek-MoE edge-cloud collaborative architecture includes 1) lightweight experts at the edge: real-time face blurring, voiceprint desensitization, and basic behavior recognition; 2) deep experts in the cloud: psychological risk modeling, causal reasoning, and multimodal game simulation.

[0111] In summary, the present invention provides a campus safety analysis and early warning method that integrates multimodal reasoning capabilities. It converts multimodal data into quantitative parameters through a mathematical modeling engine, constructs an event causal graph using a recursive causal discovery algorithm, and triggers graded responses and optimizes disposal strategies based on risk levels. It is particularly suitable for application scenarios of intelligent intervention in campus safety behavior and dynamic monitoring of mental health. Among them, when intelligently intervening in campus safety behavior, ST-GCN action recognition, voice aggressiveness score, and social isolation index can be integrated to calculate the behavioral risk index and trigger graded responses (such as anonymous prompts and encrypted evidence push); when dynamically monitoring mental health, micro-expression Z-score standardization and psychological resilience attenuation curve prediction can be used to identify high-risk cases that affect campus safety 72 hours in advance, with a false alarm rate of ≤8%.

[0112] Compared with existing technologies, the present invention has made many innovative improvements, such as mathematically enhanced vertical field adaptation, privacy-efficiency balance design, and causal chain-driven active defense. Among them, with regard to mathematically enhanced vertical field adaptation, the present invention combines DeepSeek's mathematical reasoning capabilities with campus scenarios, and realizes quantifiable risk prediction through psychological resilience differential equations and group behavior game models; with regard to privacy-efficiency balance design, the present invention combines federated learning and edge desensitization technology to retain nearly 95% of feature information while ensuring biometric privacy; with regard to causal chain-driven active defense, the present invention locates the root cause of the event through recursive causal discovery to generate a mathematically based intervention plan.

[0113] Through systematic innovations in multimodal large models, mathematically enhanced reasoning, and hierarchical computing architecture, this invention can overcome the three long-standing technical challenges in the field of campus security: unclear vision, inaccurate judgment, and slow response, and achieve a paradigm upgrade from passive monitoring to active defense.

[0114] Example 2 This embodiment discloses a campus safety analysis and early warning system that integrates multimodal reasoning capabilities.

[0115] like Figure 2 As shown, a campus safety analysis and early warning system integrating multimodal reasoning capabilities includes: The equipment layer is used for multimodal data collection on campus; Multimodal perception layer, used to: process multimodal data; The edge computing layer is used to: perform lightweight processing and privacy protection on the multimodal data after data processing; The cloud-based DeepSeek engine is used to analyze lightweight and privacy-protected multimodal data and output a quantitative risk assessment level. The execution and feedback layer is used to trigger a level 3 alert based on the resulting risk assessment level and deliver differentiated response recommendations for campus safety.

[0116] Furthermore, the device layer includes cameras, directional microphones, and IoT devices; all data collected by the device layer is sent to the terminal computer. The multimodal perception layer includes visual units, auditory units, and text units; the edge computing layer includes a lightweight feature encoder and a privacy protection module; the cloud-based DeepSeek engine includes a mathematical modeling engine and a logical reasoning accelerator; and the execution and feedback layer is equipped with a multi-level early warning mechanism and supports linked drone tracking, access control, and the generation of psychological intervention scripts.

[0117] Furthermore, the data interaction between layers includes: 1) The interaction mechanism between the device layer and the multimodal perception layer is as follows: the camera is directly connected to the vision unit through the CoaXPress 2.0 interface; the directional microphone is connected to the hearing unit through the AES67 audio network protocol; and the IoT devices are aggregated to the text unit through the LoRaWAN network management.

[0118] 2) The interaction between the multimodal perception layer and the edge computing layer is as follows: cross-modal data is interconnected through PCle Switch; the visual unit outputs the coordinates of the skeleton joint points; and the auditory unit transmits the voiceprint hash value and emotion label.

[0119] 3) The interaction between the edge computing layer and the cloud-based DeepSeek engine is as follows: lightweight data is uploaded via the IPSec VPN tunnel.

[0120] 4) The interaction between the cloud and the execution and feedback layer is as follows: control instructions are issued through 5G URLLC; priority mapping: red alerts use NS3-level dedicated channels.

[0121] Example 3 The purpose of this embodiment is to provide a computer-readable storage medium.

[0122] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of a campus safety analysis and early warning method integrating multimodal reasoning capabilities as described in the first embodiment of the present disclosure.

[0123] Example 4 The purpose of this embodiment is to provide an electronic device.

[0124] An electronic device includes a memory, a processor, and a program stored in the memory and executable on the processor. When the processor executes the program, the steps of a campus safety analysis and early warning method integrating multimodal reasoning capabilities as described in the first embodiment of the present disclosure are implemented.

[0125] The steps involved in the apparatuses of Examples 2, 3, and 4 above correspond to those of Method Example 1. For detailed implementations, please refer to the relevant description of Example 1. The term "computer-readable storage medium" should be understood to mean a single medium or multiple media containing one or more instruction sets; it should also be understood to include any medium capable of storing, encoding, or carrying an instruction set for execution by a processor and causing the processor to perform any method of the present invention.

[0126] Those skilled in the art will appreciate that the modules or steps of the present invention described above can be implemented using a general-purpose computer device. Alternatively, they can be implemented using program code executable by a computing device, which can then be stored in a storage device and executed by the computing device. Alternatively, they can be fabricated into separate integrated circuit modules, or multiple modules or steps can be fabricated into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.

[0127] Although the above describes the specific embodiments of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without any creative work are still within the scope of protection of the present invention.

Claims

1. A campus safety analysis and early warning method integrating multimodal reasoning capabilities, characterized by: include: Multimodal data collection on campus; Perform data processing on multimodal data; Lightweight and privacy-enhanced multimodal data after data processing; Utilize the cloud-based DeepSeek engine, which includes a mathematical modeling engine and a logic reasoning accelerator, to analyze the lightweight and privacy-protected multimodal data and output a quantitative risk assessment level. A level 3 alert is triggered based on the resulting risk assessment level, and differentiated handling recommendations for campus safety are pushed.

2. A campus safety analysis and early warning method integrating multimodal reasoning capabilities as claimed in claim 1, characterized in that: The multimodal data includes video data, voice data within the campus monitoring range, and text data on the campus social platform.

3. A campus safety analysis and early warning method integrating multimodal reasoning capabilities according to any one of claims 1-2, characterized in that: Data processing of multimodal data includes: identifying behavioral actions in collected video data based on spatiotemporal graph convolutional networks; separating environmental noise from collected voice data and quantifying voice aggressiveness based on directional microphone arrays and voiceprint filtering algorithms; real-time monitoring of social platform keywords in collected text data based on crawler technology, and combining DeepSeek-NLP to identify metaphorical risk content.

4. The campus safety analysis and early warning method integrating multimodal reasoning capabilities according to claim 1 is characterized in that: Lightweight processing and privacy protection are performed on multimodal data, including: first, feeding the multimodal data into a lightweight feature encoder, and performing dimensionality reduction on the multimodal data based on the lightweight feature encoder; then, privacy protection is performed on the multimodal data after the dimensionality reduction operation based on the edge end under the federated learning framework.

5. The campus safety analysis and early warning method integrating multimodal reasoning capabilities as claimed in claim 1 is characterized in that: When using the cloud-based DeepSeek engine to analyze multimodal data, the mathematical modeling engine is used to construct a psychological resilience differential equation to quantify the psychological risk index; the logical reasoning accelerator is integrated with a recursive causal discovery algorithm to construct an event evolution network.

6. The campus safety analysis and early warning method integrating multimodal reasoning capabilities according to claim 1, characterized in that: The three-level alarm includes blue, yellow and red alarms, and each color represents a different alarm severity. Specifically, red alarm>yellow alarm>blue alarm.

7. A campus safety analysis and early warning system integrating multimodal reasoning capabilities, characterized by: include: The equipment layer is used for multimodal data collection on campus; Multimodal perception layer, used to: process multimodal data; The edge computing layer is used to: perform lightweight processing and privacy protection on the multimodal data after data processing; The cloud-based DeepSeek engine is used to analyze lightweight and privacy-protected multimodal data and output a quantitative risk assessment level. The execution and feedback layer is used to trigger a level 3 alert based on the resulting risk assessment level and deliver differentiated response recommendations for campus safety.

8. The campus safety analysis and early warning system integrating multimodal reasoning capabilities as claimed in claim 7, characterized in that: include: The device layer includes cameras, directional microphones, and IoT devices; The multimodal perception layer includes a visual unit, an auditory unit, and a text unit; The edge computing layer includes a lightweight feature encoder and a privacy protection module; The cloud-based DeepSeek engine includes a mathematical modeling engine and a logical reasoning accelerator; A multi-level early warning mechanism is configured in the execution and feedback layer.

9. A computer-readable storage medium having a program stored thereon, characterized in that: When the program is executed by a processor, the steps of the campus safety analysis and early warning method integrating multimodal reasoning capabilities as described in any one of claims 1 to 6 are implemented.

10. An electronic device comprising a memory, a processor, and a program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps in the campus safety analysis and early warning method integrating multimodal reasoning capabilities as described in any one of claims 1-6 are implemented.

Citation Information

Patent Citations

  • Campus bullying prevention system based on voice recognition and big data analysis

    CN119229593A

  • Lightweight pedestrian falling detection method and system for inspection robot

    CN120148115A

  • Chinese text classification method based on metaphor association and label constraint comparative learning

    CN120256638A

  • Campus safety monitoring method and system based on multi-mode perception

    CN120299169A

  • Artificial intelligence-based system for analyzing student behavior

    DE202024107622U1

Cited By

  • Campus security risk prevention and control cloud service platform

    CN121012835A

  • A campus safety risk prevention and control cloud service platform

    CN121012835B