Campus safety monitoring method and system based on multi-mode perception

Through multimodal perception technology, multimodal data on campus is collected and analyzed in real time, early warning signals are generated and integrated, and alarm thresholds are dynamically adjusted, which solves the privacy protection and scenario adaptability problems in campus bullying monitoring and improves the accuracy and efficiency of monitoring.

CN120299169APending Publication Date: 2025-07-11LANGFANG XIHANG TECHNOLOGY CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510686016.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

When monitoring campus bullying behaviors, the existing technology has privacy protection problems and the difficulty of a single data source to capture complex features in a comprehensive way, resulting in the risk of false alarms and missed detection, and the fixed threshold determination mechanism leads to poor scenario adaptability.

Method used

The front-end perception unit collects multimodal data in real time, including human movements, environmental sounds and personnel physiological data, independently analyzes and generates early warning signals, and conducts signal strength, superimposed weights and space-time dynamic thresholds to determine the comprehensive alarm level on the cloud platform, and dynamically adjusts the alarm threshold to achieve adaptive scene monitoring.

Benefits of technology

Significantly reduce the false alarm rate, realize adaptive scenario monitoring, optimize resource allocation and response efficiency, and support intelligent system upgrades and function expansion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299169A_ABST
    Figure CN120299169A_ABST
Patent Text Reader

Abstract

The invention provides a campus safety monitoring method and system based on multi-modal perception, and relates to the technical field of campus safety, and the method comprises the steps: collecting multi-modal data in real time through a front-end perception unit, the multi-modal data comprising human body motion data, environmental sound data and personnel physiological data; independently analyzing the multi-modal data, and respectively generating a first early warning signal, a second early warning signal and a third early warning signal; the three early warning signals are uploaded to a cloud analysis platform for fusion analysis, so that the cloud analysis platform determines a comprehensive alarm level based on the signal intensity, the superposition weight and the time-space dynamic threshold of the three early warning signals; the space-time dynamic threshold value is dynamically adjusted based on the incident time and the incident place; early warning information is directionally distributed to a preset terminal according to the comprehensive alarm level and the incident site, and emergency response operation is triggered. According to the method provided by the invention, the campus spoofing event can be detected under the condition that the privacy of students is not influenced, and the detection result is more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to the technical field of campus security, and particularly relates to a campus security monitoring method and system based on multi-modal perception. Background Art

[0002] Campus bullying, as an important social problem endangering the physical and mental health of teenagers, has characteristics such as strong concealment and diverse scenarios. It not only causes physical and mental trauma to the bullied, but also may have a profound negative impact on bystanders and the entire campus environment. Therefore, it is necessary to effectively monitor campus bullying behavior.

[0003] Traditional monitoring methods use cameras to record. Due to privacy protection, they cannot be deployed in sensitive areas such as dormitories and toilets, resulting in these places becoming blind spots for bullying monitoring. In addition, a single data source is difficult to comprehensively capture complex features, and existing monitoring methods based on multi-source data fusion generally adopt a fixed threshold determination mechanism, resulting in poor scene adaptability and prone to false alarms and missed detections. Summary of the Invention

[0004] In view of the above-mentioned defects or deficiencies in the prior art, it is desirable to provide a campus security monitoring method and system based on multi-modal perception to solve the above problems.

[0005] The first aspect of the present invention provides a campus security monitoring method based on multi-modal perception, including: Real-time collection of multi-modal data in the target area through a front-end perception unit, where the multi-modal data includes human motion data, environmental sound data, and personnel physiological data; Independently analyze the multi-modal data to generate a first warning signal, a second warning signal, and a third warning signal respectively; Upload the three warning signals to a cloud analysis platform for fusion analysis, so that the cloud analysis platform determines the comprehensive alarm level based on the signal strength, superposition weight, and spatio-temporal dynamic threshold of the three warning signals; the spatio-temporal dynamic threshold is dynamically adjusted based on the incident time and incident location; Directly distribute warning information to a preset terminal according to the comprehensive alarm level and incident location, and trigger an emergency response operation.

[0006] According to the technical solution provided by the present invention, the independently analyzing the multi-modal data to generate a first warning signal, a second warning signal, and a third warning signal respectively includes: Call a first preset database according to the human motion data to match dangerous action labels, and output a first warning signal when the match is successful; Call a second preset database according to the environmental sound data to match dangerous vocabulary labels, and output a second warning signal when the match is successful; According to the physiological data of the person, call the third preset database to compare the breathing frequency threshold and the heart rate threshold, and output a third warning signal when the threshold is exceeded.

[0007] According to the technical solution provided by the present invention, uploading the three warning signals to the cloud analysis platform for fusion analysis, so that the cloud analysis platform determines the comprehensive alarm level based on the signal strength, superimposed weight and spatio-temporal dynamic threshold of the three warning signals, includes: Determine the signal strength of the three warning signals, and the signal strength includes action matching degree, vocabulary matching degree and physiological abnormality index; Based on the incident time and incident location, assign weight values to the signal strength of the three warning signals respectively, and calculate the weighted sum to obtain a risk assessment parameter; Call the risk level database according to the incident time and incident location to obtain the spatio-temporal dynamic threshold, and determine the comprehensive alarm level according to the matching result of the risk assessment parameter and the spatio-temporal dynamic threshold range.

[0008] According to the technical solution provided by the present invention, the acquisition of the human body movement data includes: Obtain multiple infrared images through an infrared imaging device, perform heat source segmentation and time series analysis on the multiple infrared images, and identify the torso, limbs and head contours of the target person; According to the area change rate and temperature gradient distribution of the heat source contact area between adjacent frames in the infrared image, determine the type of contact event between the target persons; the type of contact event includes a dynamic contact event and an instantaneous contact event; Determine a pre-trained action classification model according to the type of contact event, and input the area change rate and temperature gradient distribution into the action classification model to output a preliminary action label, and generate human body movement data.

[0009] According to the technical solution provided by the present invention, determining the type of contact event between the target persons according to the area change rate and temperature gradient distribution of the heat source contact area between adjacent frames in the infrared image includes: Calculate the area change rate of the heat source contact area between consecutive frames in the infrared image. If the area change rate of the contact area between consecutive frames exceeds the first set value, it is determined as an effective contact event; Extract the temperature gradient distribution of the contact area to construct a three-dimensional thermodynamic feature matrix, and divide the type of contact event in combination with the contact duration.

[0010] According to the technical solution provided by the present invention, the method further includes: Based on the heat source contour, reversely track the movement trajectory for a preset duration before contact, and extract the acceleration change rate, direction deflection angle and relative speed; Calculate the action rationality score through the motion intention analysis model to correct the confidence of the preliminary action label; the confidence of the preliminary action label is used to adjust the action intensity.

[0011] According to the technical solution provided by the present invention, the method further includes: Calculate the dynamic energy density of the contact area according to the three-dimensional thermodynamic feature matrix; Call the confidence correction database, match the confidence correction value based on the dynamic energy density and the preliminary action label, and correct the confidence of the preliminary action label.

[0012] According to the technical solution provided by the present invention, the method further includes: Obtain the event type and injury level according to the intermediate data of the comprehensive alarm level; generate a bullying event report including the event type and injury level.

[0013] According to the technical solution provided by the present invention, the method further includes: Monitor the power supply status of the front-end sensing unit, enable the backup power supply when power is off, and generate a power-off alarm signal.

[0014] According to the technical solution provided by the present invention, for implementing the above-mentioned campus security monitoring method based on multi-modal perception, the system includes: A data acquisition module, which is configured to collect multi-modal data in the target area in real time through a front-end sensing unit, and the multi-modal data includes human body action data, environmental sound data and personnel physiological data; A data analysis module, which is configured to independently analyze the multi-modal data and generate a first warning signal, a second warning signal and a third warning signal respectively; A data sending module, which is configured to upload the three warning signals to a cloud analysis platform for fusion analysis, so that the cloud analysis platform determines the comprehensive alarm level based on the signal strength, superposition weight and spatio-temporal dynamic threshold of the three warning signals; the spatio-temporal dynamic threshold is dynamically adjusted based on the incident time and incident location; A cloud response module, which is configured to direct and distribute warning information to a preset terminal according to the comprehensive alarm level and the incident location, and trigger an emergency response operation Compared with the prior art, the beneficial effects of the present invention are as follows: By synchronously collecting human actions, environmental sounds, and personnel physiological data through the front-end perception unit, and independently analyzing and then fusing and processing them, multi-dimensional cross-verification is achieved, significantly reducing the false alarm rate; By dynamically adjusting the alarm threshold based on the incident time and incident location, scene-adaptive monitoring is realized, avoiding the rigidity problem caused by a fixed threshold; By fusing and analyzing the signal strength, weight, and dynamic threshold to judge the comprehensive alarm level, and directing and distributing early warning information to preset terminals, hierarchical emergency response is achieved, optimizing resource allocation and response efficiency; Through the data fusion architecture of the cloud analysis platform, it supports model training and multi-terminal expansion, realizing the intelligent upgrade and function expansion of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] By reading the detailed description of the non-restrictive embodiments with reference to the following drawings, other features, objectives, and advantages of the present invention will become more apparent: Figure 1 It is a flowchart of the steps of a campus security monitoring method based on multi-modal perception provided in Embodiment 1 of the present invention; Figure 2 It is a schematic structural diagram of a campus security monitoring system based on multi-modal perception provided in Embodiment 2 of the present application.

[0016] Reference numerals in the drawings: 10, data acquisition module; 20, data analysis module; 30, data sending module; 40, cloud response module. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] The present invention will be further described in detail below with reference to the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related invention, rather than limiting the invention. Additionally, it should be noted that for the sake of description, only parts related to the invention are shown in the drawings.

[0018] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the drawings and embodiments.

[0019] Please refer to the figure. This embodiment provides a campus security monitoring method based on multi-modal perception, including: S100: Real-time collect multi-modal data within the target area through the front-end perception unit, and the multi-modal data includes human action data, environmental sound data, and personnel physiological data.

[0020] In step S100, the front-end sensing unit includes an action pickup device, a sound pickup device, and a physiological data pickup device. The action pickup device is only used to pick up the human actions within the target area and does not collect the optical images within the area, so as to avoid infringing on the privacy of students. Therefore, the action pickup device can be deployed in privacy places such as toilets and dormitories. The sound pickup device is used to collect the ambient sound, especially the voices of students within the area. The physiological data pickup device can be optionally a millimeter-wave radar, which is used to detect the breathing and heart rate of students within the area. The front-end sensing unit is deployed at multiple locations within the school. The front-end sensing unit is signal-connected to the back-end monitoring system through the Internet of Things or the Internet. The human action data, ambient sound data, and personnel physiological data collected by the front-end sensing unit are sent to the monitoring system for processing in real time. When each front-end sensing unit uploads data, it also sends the time and location of data collection to the monitoring system for subsequent tracking of the incident time and incident location.

[0021] The human action data includes multiple preliminary action tags. The preliminary action tags indicate that a certain part of the body of person A acts on a certain part of the body of person B, which is used to judge whether a physical conflict occurs. The ambient sound data includes multiple words recognized from the speech, which is used to judge whether a language conflict occurs. The personnel physiological data includes the breathing rate and heart rate, which is used to monitor whether there is an abnormal physiological characteristic phenomenon.

[0022] Furthermore, since the action pickup device only picks up the human actions within the target area, an optical camera cannot be used. Therefore, in this embodiment, the human action data is collected through an infrared imaging device. The collection of the human action data includes: S110: Obtain multiple infrared images through the infrared imaging device, perform heat source segmentation and time series analysis on the multiple infrared images, and identify the torso, limbs, and head contours of the target person.

[0023] In step S110, the infrared imaging device is arranged at multiple corners within the campus. The infrared imaging device collects infrared images of the environment within its field of view at regular time intervals to obtain multiple frames of infrared images. In this embodiment, the infrared imaging device is integrated with a personnel recognition function. When the infrared imaging device detects the presence of a person in the area, it collects infrared images at regular time intervals, enabling the infrared imaging device to be activated only under set conditions, thereby reducing the power consumption of the infrared imaging device. There is at least one target person in the infrared image, and only the heat source contour of the target person is shown in the infrared image. The torso, limbs, and head of the target person can be identified from the infrared image. When there is more than one target person in the infrared image, heat source segmentation is performed on the multiple target persons, and thus the heat sources in the infrared image are segmented into multiple individual target persons. By performing temporal analysis on multiple frames of infrared images, the action states of each target person in the infrared image over time can be represented, and thus the actions of the target persons can be recognized.

[0024] S120: Determine the type of contact event between target persons according to the area change rate and temperature gradient distribution of the heat source contact area between adjacent frames in the infrared image; the contact event types include dynamic contact events and instantaneous contact events.

[0025] In step S120, this step is used to recognize the actions of the target persons in the infrared image. After obtaining multiple frames of infrared images, by comparing the area change rate of the contact area of different heat source contours between two adjacent frames of infrared images, it can be determined whether two target persons in the infrared image are in contact. Combining the temperature gradient distribution within the contact area between two adjacent frames of infrared images and the temperature change rate within the contact area, the type of contact event between target persons can be classified. By classifying the contact time types, it is convenient to perform targeted processing according to different contact event types subsequently.

[0026] Further, step S120 specifically includes: S121: Calculate the area change rate of the heat source contact area between consecutive frames in the infrared image. If the area change rate of the contact area between consecutive frames exceeds a first set value, it is determined as an effective contact event; In step S121, after obtaining multiple frames of infrared images, calculate the change rate of the contact area between consecutive frames. When it is determined that the change rate of the contact area between consecutive frames does not exceed the first set value, it is determined as a non-effective contact event at this time. For example, if the first set value is 10% and the change rate of the contact area of five consecutive frames of infrared images is 8%, it means that the two target persons may just be walking hand in hand or with one person's arm on the other person's shoulder while walking, etc., and step S122 is not continued; when it is determined that the contact area between consecutive frames exceeds the first set value, it is determined as an effective contact event at this time. For example, if the contact area between the palm and the face suddenly increases by 18% in five consecutive frames of infrared images, it is possible that one of the target persons slaps the other target person, so step S122 is continued.

[0027] S122: Extract the temperature gradient distribution of the contact area to construct a three-dimensional thermodynamic feature matrix, and classify the contact event types in combination with the contact duration.

[0028] In step S122, first, based on the result of heat source segmentation, obtain the contact area in the infrared image, and smooth the boundary and eliminate noise through the morphological operation of OpenCV; then divide the contact area into 1mm×1mm grids, and record the temperature value at each grid point , to obtain the sampling matrix of the contact area. For example, when the contact area is 10cm×10cm, a 100×100 sampling matrix is generated; then continue to calculate the horizontal and vertical temperature gradients using the central difference method:

[0029]

[0030] where and are the spatial resolutions, with a default value of 1mm; Then continue to construct the three-dimensional thermodynamic feature matrix. Define the first dimension as the spatial coordinates to record the position of each grid point , position the second dimension as the temperature parameters, including the temperature at the contact center point , the edge temperature attenuation coefficient and the temperature change rate , define the third dimension as the time series, and record the consecutive frame data according to the timestamp ; the single-frame matrix form is:

[0031] After multiple frames are stacked, a three-dimensional structure is formed; where through exponential fitting where is the distance from the center; the temperature change rate Calculation based on the temperature difference between adjacent frames; Next, the contact duration is determined. First, the starting point is detected. When the contact area exceeds the second set value within the first set number of consecutive frames, it is marked as the contact start time. Then, the ending point is detected. When the contact area is lower than the second set value for the second set number of consecutive frames, it is marked as the contact end time. In this embodiment, when the contact area exceeds 10 cm within 3 consecutive frames (time window 0.1 s) 2 , it is marked as the base start time , when the contact area is lower than 10 cm for 5 consecutive frames (0.17 s) 2 , it is marked as the contact end time ; The calculated duration is ; Then, according to the duration and the temperature change rate, the contact event type is judged. When seconds and , it is judged as a contact event, such as a quick violent act like a slap. At this time, the edge temperature decay coefficient is relatively high; when seconds and , it is judged as a dynamic contact event, such as continuous behaviors like dragging and pressing. At this time, the edge temperature decay coefficient is relatively low.

[0032] By using the contact time and the temperature change rate of the contact area to judge the contact event type, the misjudgment probability of non-violent actions is reduced. For example, high-fiving belongs to a quick non-violent action. Among them, the edge temperature decay coefficient also has a certain guiding role in the contact event type. Therefore, in other embodiments, the influence of the edge temperature decay coefficient on the contact time type judgment can also be considered.

[0033] S130: Determine the pre-trained action classification model according to the contact event type, input the area change rate and the temperature gradient distribution into the action classification model to output a preliminary action label, and generate human action data.

[0034] In step S130, the action classification model is pre-trained. In this embodiment, there are two action classification models, which are respectively used to identify specific actions for dynamic contact events and instantaneous contact events. Here, two action classification models are used for action recognition. The purpose is to reduce the number of parameters of a single model, reduce the inference latency of a high-complexity model on edge devices, and ensure the real-time performance of action recognition. Specifically, while judging the type of contact event, according to the contour position of the target person in the contact area, the contact area is determined as the specific contact part of the target person, and then the contact part combination is obtained; after obtaining the three-dimensional thermodynamic feature matrix according to steps S121 - S122, it is input into the corresponding action classification model and combined with the contact part combination. Then, the preliminary action label is output by the action classification model, and the human action data is generated according to the preliminary action label. Thus, the human action data is obtained.

[0035] Specifically, for the convenience of understanding steps S110 - S130, an example of the preliminary action label is given here. In this embodiment, the preliminary action label defines four types of violent contact labels and two types of normal contact labels. Among them, the violent contact labels include slapping, pushing, dragging, and suppressing, and the normal contact labels include high-fiving and patting on the shoulder; the purpose of judging the type of contact event is to determine the corresponding action recognition model, and the process of judging the type of contact event is actually to determine the contact duration and the temperature change rate in the contact area. Therefore, steps S110 - S130 can be simply represented by Table 1 below: Table 1

[0036] S200: Independently analyze the multi-modal data to generate the first warning signal, the second warning signal, and the third warning signal respectively.

[0037] In step S200, according to the multi-modal data obtained in step S100, each type of data is independently analyzed. When it is judged that there is a physical conflict based on the human action data, the first warning signal is generated for the human action data; when it is judged that there is a verbal conflict based on the environmental sound data, the second warning signal is generated for the environmental sound data; when it is judged that there is an abnormal physiological feature based on the personnel physiological data, the third warning signal is generated for the personnel physiological data. Independent analysis based on different types of data can simplify the processing process and reduce the resource occupation in the calculation process.

[0038] Further, step S200 specifically includes: S201: Call the first preset database to match the dangerous action label according to the human action data, and output the first warning signal when the match is successful.

[0039] In step S201, multiple dangerous action tags are pre-stored in the first preset database. After calling the first preset database with the human motion data, the preliminary action tags in the human motion data are matched with the dangerous action tags. It should be noted that there may be more than one type of preliminary action tag. When matching the preliminary action tags and the dangerous action tags, when it is determined that at least one preliminary action tag and a dangerous action tag match, the first warning signal is output. According to the different numbers of matching preliminary action tags and dangerous action tags, the signal strength of the first warning signal output is also different. The more the number of matches, the stronger the signal strength.

[0040] S202: Call the second preset database according to the environmental sound data to match the dangerous vocabulary tags, and output the second warning signal when it is determined that the match is successful.

[0041] In step S202, multiple dangerous vocabulary tags are pre-stored in the second preset database. After calling the second preset database with the environmental sound data, the vocabulary in the environmental sound data is matched with the dangerous vocabulary tags. There is also more than one vocabulary in the environmental sound data. When matching the vocabulary in the environmental sound data and the dangerous vocabulary tags, when it is determined that at least one vocabulary in the environmental sound data and a dangerous vocabulary tag match, the second warning signal is output. According to the different numbers of vocabulary matches, the second warning signal with different signal strengths is output. The more the number of vocabulary matches, the stronger the signal strength.

[0042] S203: Call the third preset database according to the personnel physiological data to compare the breathing frequency threshold and the heart rate frequency threshold, and output the third warning signal when it is determined that the threshold is exceeded.

[0043] In step S203, multiple breathing frequency ranges and multiple heart rate frequency ranges are pre-stored in the third preset database. After calling the third preset database with the personnel physiological data, the breathing frequency is compared with the breathing frequency range. When the breathing frequency exceeds the threshold of the breathing frequency range, it indicates that the personnel in the area have abnormal breathing. At the same time, the heart rate frequency is compared with the heart rate frequency range. When the heart rate frequency exceeds the threshold of the line frequency range, it indicates that the personnel in the area have abnormal heart rate. When it is recognized that both the heart rate and breathing are abnormal, the third warning signal is output. According to how much the heart rate frequency and the breathing frequency exceed the corresponding thresholds, the third warning signal with different signal strengths is output. The more the threshold is exceeded, the stronger the signal strength.

[0044] S300: Upload the three warning signals to the cloud analysis platform for fusion analysis, so that the cloud analysis platform determines the comprehensive alarm level based on the signal strength, superposition weight and spatio-temporal dynamic threshold of the three warning signals; the spatio-temporal dynamic threshold is dynamically adjusted based on the incident time and incident location.

[0045] In step S300, the cloud analysis platform serves as the core processor. The cloud analysis platform is deployed in a school or an operator, and communicates via the Internet of Things or the Internet. After the three warning signals are output, they are processed by the cloud analysis platform. The cloud analysis platform performs fusion analysis on the three warning signals based on the two dimensions of time and space when the event occurs, and then determines the comprehensive alarm level. The comprehensive alarm level is used to characterize the severity of the bullying event.

[0046] Furthermore, step S300 specifically includes: S310: Determine the signal strengths of the three warning signals. The signal strengths include action matching degree, vocabulary matching degree, and physiological abnormality index.

[0047] In step S310, after the cloud analysis platform receives the signal strengths of the three warning signals, it respectively obtains the signal strengths of the three warning signals and normalizes the signal strengths of the three warning signals. It should be noted that the higher the action matching degree, the more the number of preliminary action labels matching the dangerous action labels, that is, the higher the signal strength of the first warning signal, and the more serious the situation of physical conflict at this time; the higher the vocabulary matching degree, the more the number of words in the environmental sound data matching the dangerous vocabulary labels, that is, the higher the intensity of the second warning signal, and the more serious the situation of language conflict at this time; the higher the physiological abnormality index, the more the breathing frequency and heart rate of the people in the area deviate from the normal values, that is, the higher the intensity of the third warning signal, and it can be verified that there is indeed a physical conflict or a language conflict at this time.

[0048] S320: Based on the time and location of the incident, assign weight values to the signal strengths of the three warning signals, and calculate the weighted sum to obtain a risk assessment parameter.

[0049] In step S320, the weight values assigned to the signal strengths of the three warning signals are determined based on the time and location of the incident. Specifically, in this embodiment, a weight value database is established in advance based on time and location. The weight value database includes multiple incident times, multiple incident locations corresponding to each incident time, and the weight values of the three signal strengths corresponding to each incident location. The weight value database is shown in Table 1 below: Table 2

[0050] In Table 2, α, β, and γ are the signal strength weight values of the first warning signal, the second warning signal, and the third warning signal respectively. Only some examples are given in Table 2. Specifically, the settings of the incident time, incident location, and weight values are determined according to the actual situation of the school.

[0051] Refer to Table 2 and explain the situations therein. Taking the playground during the time period from 08:00 to 17:00 as an example, since this time period is usually during class or break activities, there may be physical education classes or students having break activities on the playground. Due to the frequent activities on the playground, the weight value of human motion data is set to the highest. The environmental sound data is easily interfered by environmental noise, so the weight value of environmental sound data is set lower; taking the dormitory during the time period from 18:00 to 22:00 as an example, since it is the night dormitory environment, the weight value of personnel physiological data is set higher, and the human motion data is the second.

[0052] In step S320, by matching the incident time and incident location corresponding to the three warning signals with the weight value database, a unique set of weight values can be obtained, realizing the dynamic adjustment of the weight values based on the incident time and incident location. After obtaining the signal strength weight values of the three warning signals, calculate the weighted sum of the three warning signal strengths to obtain the risk assessment parameter. The calculation of the weighted sum enables multi-dimensional cross-verification of the occurrence of bullying incidents, improves the true judgment of the occurrence of bullying incidents, and significantly reduces the false alarm rate.

[0053] S330: Call the risk level database according to the incident time and incident location to obtain the spatio-temporal dynamic threshold, and determine the comprehensive alarm level according to the matching result of the risk assessment parameter and the spatio-temporal dynamic threshold range.

[0054] In step S330, the risk level database includes multiple incident times, multiple incident locations corresponding to each incident time, multiple spatio-temporal dynamic ranges corresponding to each incident location, and comprehensive alarm levels corresponding to each spatio-temporal dynamic range. Among them, the threshold of the spatio-temporal dynamic range is the spatio-temporal dynamic threshold. The risk level database is shown in Table 3 below: Table 3

[0055] Only examples in some cases are given in Table 3. Specifically, the settings of the incident time, incident location, spatio-temporal dynamic range, and comprehensive alarm level are determined according to the actual situation of the school. Referring to Table 3, taking the classroom and playground during the time period from 08:00 to 17:00 as an example, in the classroom scenario, due to limited activities, when the risk assessment parameter satisfies the spatio-temporal dynamic range of 0 - 50, it belongs to the low-risk situation. While in the playground scenario, due to engaging in sports activities, when the risk assessment parameter satisfies the spatio-temporal dynamic range of 0 - 60, it belongs to the low-risk situation; taking the dormitory from 22:00 to 06:00 as an example, since it is a late-night scenario, the risk assessment parameter only belongs to the low-risk situation when it satisfies the spatio-temporal dynamic range of 0 - 20.

[0056] In step S330, by matching the incident time, incident location, and risk assessment parameters with the risk level database, a unique comprehensive alarm level can be obtained, realizing dynamic adjustment of the spatio-temporal dynamic range based on the incident time and incident location. By dynamically adjusting the alarm threshold based on the incident time and incident location, scene adaptive monitoring is achieved, avoiding the rigidity problem caused by fixed thresholds, being more in line with the actual situation, and thus enabling more accurate judgment of bullying incidents.

[0057] S400: Directly distribute the warning information to the preset terminals according to the comprehensive alarm level and the incident location, and trigger the emergency response operation.

[0058] In step S400, the terminal can be the mobile phones of teachers, security personnel, dormitory administrators, or even the police. Corresponding terminals are set in advance according to different combinations of comprehensive alarm levels and incident locations. For example, in the classroom during the period from 08:00 to 17:00, if the comprehensive alarm level is medium risk, a minor conflict may occur, and the comprehensive alarm level is sent to the class teacher's mobile phone to notify the class teacher to check. If the integrated alarm level is high risk, a serious conflict may occur, and the comprehensive alarm level is sent to the security mobile phone. While sending the comprehensive alarm level and the incident location, the emergency response operation is also triggered. The emergency response operation includes triggering the campus broadcast alarm for warning or directly calling the police. By directly distributing the warning information to the preset terminals, hierarchical emergency response is achieved, optimizing resource allocation and response efficiency.

[0059] Furthermore, the method further includes: Obtain the event type and the injury level according to the intermediate data of the comprehensive alarm level; Generate a bullying incident report including the event type and the injury level.

[0060] Specifically, the event type includes physical bullying incidents and verbal bullying incidents, and the injury levels respectively include the levels of physical bullying incidents and verbal bullying incidents. During the process of obtaining the comprehensive alarm level, the intermediate data is obtained, and the intermediate data is input into a pre-trained bullying classification model. The bullying classification model identifies the human action data and environmental sound data in the intermediate data and outputs the event type and the corresponding injury level. Finally, a bullying incident report is generated and stored according to the event type and the corresponding injury level for the teacher to trace the bullying incident.

[0061] Furthermore, the method further includes: Monitor the power supply status of the front-end sensing unit, enable the backup power supply when power is off, and generate a power-off alarm signal.

[0062] Specifically, by monitoring the power supply status of all front-end sensing units, it is possible to determine whether the front-end sensing units are in a normal working state. When it is determined that a certain front-end sensing unit is powered off, the backup power supply is activated to supply power to the powered-off front-end sensing unit, avoiding the inability to obtain multi-modal data in a timely manner, thereby affecting the monitoring of campus security.

[0063] Furthermore, a protective cover is sleeved outside each front-end sensing unit, and the protective cover is electrically connected to the monitoring system. When it is detected that the protective cover is damaged, an alarm signal is automatically generated.

[0064] In addition, the infrared imaging device is also integrated with an anti-occlusion function. When the infrared imaging device detects that the heat source distribution at each position in the infrared image within the field of view is consistent, it is determined that the infrared imaging device is occluded, and an alarm signal is generated at this time.

[0065] Embodiment 2 Based on the content of the above Embodiment 1, this embodiment provides another campus security monitoring method based on multi-modal perception. The same content as in Embodiment 1 will not be elaborated here. The differences are as follows: The method further includes: Based on the heat source contour, reverse-track the motion trajectory for a preset duration before contact, and extract the acceleration change rate, direction deflection angle, and relative velocity; Calculate the action rationality score through the motion intention analysis model, and correct the confidence of the preliminary action label; the confidence of the preliminary action label is used to adjust the action intensity.

[0066] Specifically, when the preliminary action label is output through the action classification model in Embodiment 1, the corresponding confidence is also generated for each preliminary action label. Therefore, when determining the signal intensity of the first warning signal in step S300, not only the number of label matches needs to be considered, but also the confidence of each preliminary action label needs to be considered. The signal intensity of the first warning signal is jointly determined by the number of label matches and the confidence, thereby making the signal intensity more accurate.

[0067] In Embodiment 1, effective action recognition can be performed through the current contact behavior. However, in many cases, the "motivation" of the action can better explain the problem. This embodiment judges the action motivation of the target person, thereby improving the accuracy of action recognition.

[0068] Specifically, through multiple frames of infrared images collected by the infrared imaging device, based on the temporal change of the heat source contour coordinates, reverse-track the motion trajectories of two target persons within a preset duration before contact. The preset duration can be selected as 1 second, and kinematic parameters such as the acceleration change rate, motion direction deflection angle, and relative velocity within the preset duration are extracted. Among them, the acceleration change rate is the absolute value change of the acceleration per unit time, and the calculation formula is:

[0069] Among them, is the acceleration of the current frame, is the acceleration of the previous frame, is the frame interval time; The deflection angle of the movement direction is calculated by the direction angle of the displacement vector of the heat source contour of adjacent frames. The calculation formula is:

[0070] Among them, and are the velocity vectors of the current frame and the previous frame respectively; The relative velocity is the relative motion velocity between target persons. The calculation formula is:

[0071] Among them, and are the velocity vectors of target persons A and B respectively; After obtaining the acceleration change rate, the movement direction deflection angle, and the relative velocity, according to the preset movement intention analysis model, input the above kinematic parameters, and output the action rationality score. The score range is 0 - 100 points. When training the movement intention analysis model, the violent contact label and the normal contact label in the historical data are used as positive and negative samples. The following is an example to illustrate the scoring rules of the movement intention analysis model. When it is judged that the acceleration change rate > 1 m / s 2 , 20 points are added for each item. When it is judged that the movement direction deflection angle > 45°, 15 points are added. When it is judged that the relative velocity > 2 m / s 2 , 25 points are added.

[0072] Then map the action rationality score to the confidence correction coefficient. For example: When the score > 50 points, the confidence is increased by k1×score / 100 (k1 is the correction factor, and 0.3 can be selected); When the score ≤ 50 points, the confidence is decreased by k2×(50 - score) / 50 (k2 is the correction factor, and 0.2 can be selected).

[0073] After correcting the confidence of the preliminary action label by the method provided in this embodiment, the judgment of the signal strength of the first warning signal is more accurate, and the accuracy of detecting bullying behavior is further improved.

[0074] Example 3 Based on the content of the above-mentioned Embodiment 2, this embodiment provides another campus security monitoring method based on multi-modal perception. The same content as that in Embodiment 2 will not be elaborated here. The differences are as follows: The method further includes: Calculating the dynamic energy density of the contact area according to the three-dimensional thermodynamic feature matrix; Invoking the confidence correction database, and based on the dynamic energy density and the preliminary action label, matching the confidence correction value to correct the confidence of the preliminary action label.

[0075] Based on the above Embodiment 2, this embodiment further corrects the confidence of the preliminary action label by analyzing the thermodynamic characteristics of the contact area to distinguish contact behaviors with different intensities.

[0076] Specifically, first calculate the energy density within the contact area according to the three-dimensional thermodynamic feature matrix obtained in step S122. The calculation formula for the dynamic energy density is as follows:

[0077] Wherein, represents the temperature change of the contact area at time represents the contact area; After obtaining the dynamic energy density, call the confidence correction database according to the dynamic energy density and the preliminary action label. The confidence correction database includes multiple preliminary action labels, multiple different energy density ranges corresponding to each preliminary action label, and confidence correction values corresponding to each energy density range. The confidence correction database is shown in Table 4 below: Table 4

[0078] According to the dynamic energy density and the preliminary action label, a unique confidence correction value can be obtained in Table 4. The confidence is further corrected by the confidence correction value, so as to improve the accuracy of human action data recognition through a multi-level verification method, and further significantly improve the accuracy and robustness of campus bullying monitoring. It should be noted that Table 4 is only an example, and the specific settings of the preliminary action labels and the energy density ranges are determined according to the actual situations of different schools.

[0079] Embodiment 4 Please refer to Figure 2 , this embodiment provides a campus security monitoring system based on multi-modal perception for performing the campus security monitoring method based on multi-modal perception as described in Embodiments 1-3. The system includes: A data acquisition module, wherein the data acquisition module is configured to collect multimodal data in the target area in real time through a front-end sensing unit, wherein the multimodal data includes human motion data, environmental sound data, and personnel physiological data; A data analysis module, wherein the data analysis module is configured to independently analyze the multimodal data to generate a first warning signal, a second warning signal, and a third warning signal respectively; A data sending module, the data sending module is configured to upload the three warning signals to the cloud analysis platform for fusion analysis, so that the cloud analysis platform determines the comprehensive alarm level based on the signal strength, superposition weight and spatiotemporal dynamic threshold of the three warning signals; the spatiotemporal dynamic threshold is dynamically adjusted based on the time and location of the incident; A cloud response module is configured to distribute warning information to a preset terminal according to the comprehensive alarm level and the location of the incident, and trigger an emergency response operation.

[0080] The above description is only a preferred embodiment of the present invention and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the present invention is not limited to the technical solution formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the inventive concept. For example, the above features are replaced with the technical features with similar functions disclosed in the present invention (but not limited to) to form a technical solution.

Claims

1. A campus security monitoring method based on multi-modal perception, characterized in that, Including: Collecting multi-modal data in the target area in real time through a front-end perception unit, where the multi-modal data includes human action data, environmental sound data, and personnel physiological data; Independently analyzing the multi-modal data to generate a first warning signal, a second warning signal, and a third warning signal respectively; Uploading the three warning signals to a cloud analysis platform for fusion analysis, so that the cloud analysis platform determines the comprehensive alarm level based on the signal strength, superposition weight, and spatio-temporal dynamic threshold of the three warning signals; the spatio-temporal dynamic threshold is dynamically adjusted based on the incident time and incident location; Directly distributing warning information to a preset terminal according to the comprehensive alarm level and incident location, and triggering an emergency response operation.

2. The campus security monitoring method based on multimodal perception according to claim 1, wherein The independently analyzing the multi-modal data to generate a first warning signal, a second warning signal, and a third warning signal respectively includes: Invoking a first preset database according to the human action data to match dangerous action labels, and outputting a first warning signal when the match is successful; Invoking a second preset database according to the environmental sound data to match dangerous vocabulary labels, and outputting a second warning signal when the match is successful; Invoking a third preset database according to the personnel physiological data to compare the respiration frequency threshold and the heart rate frequency threshold, and outputting a third warning signal when it exceeds the threshold.

3. The campus security monitoring method based on multi-modal perception according to claim 2, characterized in that, The uploading the three warning signals to a cloud analysis platform for fusion analysis, so that the cloud analysis platform determines the comprehensive alarm level based on the signal strength, superposition weight, and spatio-temporal dynamic threshold of the three warning signals includes: Determining the signal strength of the three warning signals, where the signal strength includes action matching degree, vocabulary matching degree, and physiological abnormality index; Assigning weight values to the signal strength of the three warning signals respectively based on the incident time and incident location, and calculating the weighted sum to obtain a risk assessment parameter; Invoking a risk level database according to the incident time and incident location to obtain the spatio-temporal dynamic threshold, and determining the comprehensive alarm level according to the matching result between the risk assessment parameter and the spatio-temporal dynamic threshold range.

4. The campus security monitoring method based on multi-modal perception according to claim 3, characterized in that, The collection of the human action data includes: Obtaining multiple infrared images through an infrared imaging device, performing heat source segmentation and time series analysis on the multiple infrared images to identify the torso, limbs, and head contours of the target person; Determining the type of contact event between target persons according to the area change rate and temperature gradient distribution of the heat source contact area between adjacent frames in the infrared image; the type of contact event includes a dynamic contact event and an instantaneous contact event; Determining a pre-trained action classification model according to the type of contact event, and inputting the area change rate and temperature gradient distribution into the action classification model to output a preliminary action label, and generating human action data.

5. The campus security monitoring method based on multi-modal perception according to claim 4, characterized in that, The determining the type of contact event between target persons according to the area change rate and temperature gradient distribution of the heat source contact area between adjacent frames in the infrared image includes: Calculating the area change rate of the heat source contact area between consecutive frames in the infrared image, and if the area change rate of the contact area between consecutive frames exceeds a first set value, determining it as an effective contact event; The temperature gradient distribution of the contact area is extracted to construct a three-dimensional thermodynamic characteristic matrix, and the contact event type is divided according to the contact duration.

6. The campus security monitoring method based on multimodal perception according to claim 5, wherein The method further comprises: Based on the heat source contour, reversely track the motion trajectory of the preset time before contact, and extract the acceleration change rate, direction deflection angle and relative speed; The action rationality score is calculated through the motion intention analysis model to correct the confidence of the preliminary action label; the confidence of the preliminary action label is used to adjust the action intensity.

7. The campus security monitoring method based on multi-modal perception according to claim 6, wherein, The method further comprises: Calculating the dynamic energy density of the contact area according to the three-dimensional thermodynamic characteristic matrix; The confidence correction database is called, and the confidence of the preliminary action label is corrected based on the dynamic energy density and the confidence correction value of the preliminary action label matching.

8. The campus security monitoring method based on multi-modal perception according to claim 7, characterized in that, The method further comprises: The event type and injury level are obtained according to the intermediate data of the comprehensive alarm level; and a bullying event report including the event type and injury level is generated.

9. The campus security monitoring method based on multi-modal perception according to claim 8, characterized in that The method further comprises: Monitor the power supply status of the front-end sensing unit, enable the backup power supply and generate a power failure alarm signal when the power is off.

10. A campus security monitoring system based on multimodal perception, characterized in that, The system is used to execute the campus safety monitoring method based on multimodal perception as described in any one of claims 1 to 9, the system comprising: A data acquisition module, wherein the data acquisition module is configured to collect multimodal data in the target area in real time through a front-end sensing unit, wherein the multimodal data includes human motion data, environmental sound data, and personnel physiological data; A data analysis module, wherein the data analysis module is configured to independently analyze the multimodal data to generate a first warning signal, a second warning signal, and a third warning signal respectively; A data sending module, the data sending module is configured to upload the three warning signals to the cloud analysis platform for fusion analysis, so that the cloud analysis platform determines the comprehensive alarm level based on the signal strength, superposition weight and spatiotemporal dynamic threshold of the three warning signals; the spatiotemporal dynamic threshold is dynamically adjusted based on the time and location of the incident; A cloud response module is configured to distribute warning information to a preset terminal according to the comprehensive alarm level and the location of the incident, and trigger an emergency response operation.

Citation Information

Cited By

  • Campus safety analysis early warning method and system fused with multi-modal reasoning capability

    CN120599772A

  • Automatic buzzing alarm control communication system

    CN120766405A

  • An automated buzzer alarm control communication system

    CN120766405B