A campus intelligent monitoring method for campus anti-bullying
By performing audio and video preprocessing at the edge and combining dynamic weighting of interaction anomaly index and semantic bullying index, the problems of high false alarm rate and high false alarm rate of existing school bullying monitoring system are solved, and rapid and accurate identification and response to school bullying incidents are achieved.
Patent Information
- Application Number
- CN202511642970.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-11
- Publication Date
- 2026-03-10
- Estimated Expiration
- 2045-11-11
AI Technical Summary
Existing school bullying monitoring systems suffer from high false alarm and false negative rates when identifying bullying behavior, especially in distinguishing between playful behavior and malicious physical attacks and in identifying low-volume verbal threats.
By performing audio and video preprocessing at the edge, suspected bullying incident segments are screened, and dynamic weighting is performed by combining interaction anomaly index and semantic bullying index. Using lightweight target detection and speech recognition technology, physical conflicts and verbal threats can be identified, enabling rapid and accurate bullying location and early warning.
It enables rapid and accurate identification and response to school bullying incidents, reduces data transmission and analysis latency, improves identification robustness and accuracy, and can accurately locate bullies and provide timely warnings.
Smart Images

Figure CN121121656B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image processing. More particularly, the present application relates to a campus intelligent monitoring method for campus anti-bullying. BACKGROUND
[0002] Campus bullying behavior poses a serious challenge to campus safety due to its suddenness, concealment and diversity of forms. Traditional security monitoring mainly relies on manual real-time monitoring or post-video tracking, which not only consumes a lot of manpower, but also is difficult to achieve real-time discovery and timely intervention of bullying incidents. Therefore, existing technologies begin to introduce artificial intelligence to automatically analyze monitoring audio and video in order to achieve proactive early warning, but still face technical bottlenecks.
[0003] In the field of video analysis, conventional action recognition algorithms have difficulty in distinguishing between normal play and malicious physical attacks because the two may be highly similar in apparent characteristics such as action amplitude and speed, resulting in high false positive rates of existing systems. In the field of audio analysis, existing methods are mostly focused on detecting high-decibel sounds such as screaming and quarreling, but lack the ability to recognize low-volume verbal bullying containing key information such as threats and insults, which can easily cause false negatives.
[0004] To improve accuracy, some technical solutions attempt to combine human pose estimation algorithms to analyze the details of body interaction and analyze speech content through natural language processing technology. However, simple pose analysis usually treats all interactors equally and cannot effectively determine the relationship between the attacker and the victim in an attack. Basic keyword matching technology also cannot evaluate the real level of verbal threats according to the severity of the words and the context, which comprehensively leads to poor monitoring effect of campus bullying. SUMMARY
[0005] To solve the above technical problem of poor monitoring effect of existing campus bullying, the present application provides a campus intelligent monitoring method for campus anti-bullying, comprising:
[0006] Acquire the audio and video data stream of the campus monitoring scene, and perform edge preprocessing to screen out a suspected bullying event segment, wherein the suspected bullying event segment comprises a suspected bullying video segment and a suspected bullying audio segment; analyze the spatiotemporal position changes of the human body key nodes of different human body targets in each video frame in the suspected bullying video segment, acquire an interaction abnormality index of the suspected bullying video segment, and the interaction abnormality index is used to represent the cumulative situation of the difference between the instantaneous interaction indexes of any two human body targets; perform vocabulary extraction on the suspected bullying audio segment, acquire a semantic bullying index of the suspected bullying audio segment according to the relevance of the vocabulary and bullying-related vocabulary; dynamically weight the interaction abnormality index and the semantic bullying index according to the environmental noise of the video scene, and acquire a comprehensive bullying index of the suspected bullying event segment; and based on the comprehensive bullying index of the suspected bullying event segment, locate the bullying event to complete the campus intelligent monitoring for campus anti-bullying.
[0007] The present application effectively reduces the data transmission and analysis delay by preprocessing at the edge end, solves the problem of high delay and high bandwidth occupation of the existing cloud analysis scheme; the present application not only evaluates the inequality of the attack behavior from the video through the interaction abnormality index to distinguish bullying from playing, but also combines the semantic bullying index in the audio to identify verbal threats, furthermore, the present application dynamically weights the audio and video features according to the environmental signal-to-noise ratio, improves the recognition robustness and accuracy in different scenes such as playgrounds and corridors, and finally realizes the rapid and accurate positioning and early warning of the campus bullying event.
[0008] Preferably, the edge preprocessing and screening to obtain the suspected bullying event segment comprise:
[0009] A lightweight target detection algorithm is used to detect the video data stream frame by frame to acquire the bounding box of each human body target; based on the bounding box, a video segment meeting any one of the following conditions is preliminarily screened out, which is recorded as a suspected bullying video segment: there is a behavior of multiple bounding boxes overlapping for more than a preset first threshold; in the video frame where no human body target is detected, there are continuous video frames with an information entropy lower than a preset second threshold and a duration longer than the preset first threshold.
[0010] Preferably, the instantaneous interaction index between any two human body targets satisfies the expression:
[0011] ;
[0012] In the formula, represents the instantaneous interaction index of the i-th human body target to the c-th human body target in the target frame; represents the hand point set of the i-th human body target in the target frame; represents the core point set of the c-th human body target in the target frame; , express The Data points The The instantaneous velocity vector of each data point; express The Data points and The Euclidean distance between data points; Indicates the minimum value; The symbol indicates the calculation of the magnitude of a vector.
[0013] This invention combines the relative speed and distance between the attacker's hand and the victim's core area for modeling, which can more accurately reflect the physical nature of aggressive actions such as pushing and hitting, that is, the hand approaching the opponent's core body part at high speed. This allows the system to more accurately quantify the attack intensity from a physical perspective and improve the accuracy of distinguishing between aggressive actions and non-aggressive limb contact.
[0014] Preferably, the interaction anomaly index satisfies the expression:
[0015] ;
[0016] In the formula, F represents the interaction anomaly index of the suspected bullying incident segment; I represents the set of human targets detected in the suspected bullying incident segment; This indicates the degree of abnormality in the interaction between the i-th individual and the c-th individual in a suspected bullying incident segment. The method of obtaining the instantaneous interaction index between the i-th human target and the c-th human target in all video frames is to accumulate the difference between the instantaneous interaction index between the c-th human target and the i-th human target in the target frame. This represents the maximum value function.
[0017] This invention calculates and accumulates the differences in instantaneous interaction indices between different individuals, and identifies the individual who experiences the greatest interaction energy, thereby determining the directionality and asymmetry of aggressive behavior throughout the entire event. A larger index indicates the presence of a concentrated target, a typical characteristic of bullying incidents. Therefore, this invention can not only determine the severity of an event but also accurately identify the bullied individual, solving the problem of difficulty in determining the aggressor relationship.
[0018] Preferably, the step of extracting words from suspected bullying audio segments includes: converting suspected bullying audio segments into text using automatic speech recognition technology; segmenting the text into words using a word segmentation algorithm to obtain multiple words from the suspected bullying audio segments; constructing a bullying keyword lexicon and obtaining a preset threat level for each bullying keyword; and converting the words and bullying keywords into word vectors using a word vector model.
[0019] Preferably, the semantic bullying index of the suspected bullying audio segment satisfies the expression:
[0020] ;
[0021] In the formula, The semantic bullying index indicates the likelihood of audio clips being used to represent bullying. The number of words indicating suspected bullying audio clips; The bullying index contribution of the m-th word in the suspected bullying audio clip; express The maximum value corresponds to the preset threat level of the bullying keyword. This represents the word vector of the m-th word in a suspected bullying audio clip, and the set of cosine similarities between this vector and the set of word vectors in a bullying keyword lexicon.
[0022] This invention multiplies the semantic similarity of keywords by their preset threat levels and then sums them up. It comprehensively considers the degree of closeness of the words used to the core bullying words and the severity of the words themselves, so that the final index can more accurately reflect the strength of aggression in the entire conversation. Compared with simple keyword matching, its evaluation results are more objective and accurate.
[0023] Preferably, the bullying index contribution of the m-th word in the suspected bullying audio segment. Satisfying the expression:
[0024] ;
[0025] In the formula, The word vector representing the m-th word in a suspected bullying audio clip; A set of word vectors representing a keyword database for bullying; The word vector of the m-th word in the suspected bullying audio segment is represented by the set of cosine similarities between the word vectors of the bullying keyword lexicon and the set of word vectors of the bullying keyword lexicon. Represents the maximum value function; Preset condition: The m-th word in the suspected bullying audio clip has the largest contribution to the bullying index among all words. One word; This indicates that the condition is met. ; It is a negation symbol. This indicates that the condition is not met. ; This is the default value.
[0026] This invention selects and sums the N words that contribute the most to the bullying index, which is equivalent to grasping the most critical words in the dialogue for analysis. It can eliminate noise interference from irrelevant words, making the semantic bullying index more focused on core threat information, thus improving the accuracy and robustness of the calculation.
[0027] Preferably, the step of obtaining the comprehensive bullying index of suspected bullying event segments includes: obtaining the signal-to-noise ratio of suspected bullying audio segments; the dynamic weighting is to perform a weighted summation of the interaction anomaly index and the semantic bullying index based on the signal-to-noise ratio.
[0028] Preferably, the step of locating bullying incidents includes: obtaining the index value corresponding to the interaction anomaly index and the corresponding human target, which is designated as the core target of attack; collecting and constructing a verification dataset of labeled audio and video clips of school bullying incidents, and calculating the comprehensive bullying index of all clips in the verification dataset; determining the warning threshold based on the ROC curve; when the comprehensive bullying index reaches the warning threshold, the system immediately pushes an alarm to designated personnel such as school security and class teachers through preset communication channels such as SMS and mobile applications. The alarm includes the location of the camera where the incident occurred, information on the core target of attack, and real-time video footage.
[0029] Preferably, the warning threshold includes a primary warning threshold and a secondary warning threshold, wherein the secondary warning threshold is greater than the primary warning threshold; when the comprehensive bullying index reaches the secondary warning threshold, the alarm pushed by the system to the designated person in charge will be marked as high priority to prompt priority handling.
[0030] The beneficial effects of this invention are as follows:
[0031] (1) This invention establishes an interaction anomaly index and a semantic bullying index to reflect the inequality of physical conflict and the severity of verbal threats, respectively. It dynamically adjusts the weight of audio and video information according to the on-site audio signal-to-noise ratio, adapts to various campus environments, and combines the low latency advantage of edge computing with a hierarchical early warning mechanism to finally achieve rapid and accurate identification and response to bullying incidents.
[0032] (2) The interaction anomaly index of the present invention not only identifies bullying, but also locates specific individuals, making the response to bullying incidents more accurate and efficient. Attached Figure Description
[0033] Figure 1 This is a flowchart illustrating an intelligent campus monitoring method for preventing bullying in schools, as described in this invention.
[0034] Figure 2 It is a schematic diagram showing 17 key nodes of the human body. Detailed Implementation
[0035] S1: Acquire audio and video data streams from campus surveillance scenes, perform edge preprocessing, and filter out suspected bullying event segments, which include suspected bullying video segments and suspected bullying audio segments.
[0036] It should be noted that school bullying is characterized by its suddenness and concealment, and often occurs in diverse settings such as corridors, playgrounds, and around restrooms. Uploading all raw audio and video data from surveillance cameras to the cloud for analysis would not only consume enormous network bandwidth and cloud storage resources but also lead to high analysis latency, failing to meet the needs of real-time early warning. Considering that most surveillance footage shows normal activity, this invention first performs real-time preprocessing at edge computing nodes close to the data source. A lightweight model quickly filters out most normal activity data, uploading only audio and video clips with potentially abnormal behavior, thereby reducing the load on data transmission and analysis and laying the foundation for subsequent accurate and real-time identification.
[0037] Specifically, the system acquires audio and video data streams from campus surveillance scenes, performs edge preprocessing, and filters out segments suspected of bullying incidents, including:
[0038] Surveillance devices are installed in multiple areas on campus to obtain real-time audio and video data streams. These areas include high-incidence areas of bullying, such as corridors, playgrounds, and areas around toilets. The surveillance devices include audio and video recording equipment such as fixed cameras, pan-tilt cameras, and microphones.
[0039] It should be noted that in order for computers to perform preliminary analysis of student activities in audio and video data streams, it is first necessary to accurately identify and locate student targets from complex backgrounds.
[0040] The video data stream is processed in real time, and a lightweight object detection algorithm is used to detect each frame of the video data stream to complete the detection of human targets, obtaining the bounding box of each human target. For example, the lightweight object detection algorithm can adopt the YOLOv8n model.
[0041] It should be noted that students' specific behaviors and interaction patterns are reflected through their body posture and relative position. Considering that the key points of the human skeleton can accurately describe human posture and movement with a relatively small amount of data, this invention further extracts human skeletal information and uses this as a basis for preliminary screening of abnormal behaviors.
[0042] For each detected human target's bounding box, the coordinates of key human nodes are obtained frame-by-frame using a skeleton extraction module. For example, the skeleton extraction module can employ a lightweight OpenPose model to extract the coordinates of 17 key human nodes, such as... Figure 2 This diagram illustrates 17 key points of the human body. 0 represents the nose, 1 and 2 represent the eyes, 3 and 4 represent the ears, 5 and 6 represent the shoulders, 7 and 8 represent the elbows, 9 and 10 represent the wrists, 11 and 12 represent the hips, 13 and 14 represent the knees, and 15 and 16 represent the ankles.
[0043] Based on the coordinate sequence, video clips meeting any of the following conditions are initially selected as suspected bullying video clips: 1) Multiple-person physical contact behavior exists, defined as the overlap time of the bounding boxes of any one or more human targets exceeding a preset first threshold (e.g., 2 seconds); 2) Continuous abnormal video frames exist, defined as consecutive video frames in which no human targets are detected, with an information entropy lower than a preset second threshold and a duration exceeding the preset first threshold. The preset second threshold is obtained by measuring the information entropy of video frames from the corresponding camera in a normal, unoccupied scene (e.g., 0.5 times the average information entropy of all video frames in which no human targets are detected). It should be noted that detecting continuous abnormal video frames is to identify situations where bullying behavior cannot be detected in a timely manner due to the bully obstructing the camera.
[0044] Suspected bullying video clips and corresponding audio data from the same time period are recorded as suspected bullying event audio and video clips, and uploaded to the cloud analysis layer via low-latency transmission links such as 5G or WiFi 6.
[0045] At this point, audio and video clips of what appeared to be a bullying incident were obtained.
[0046] S2: Analyze the spatiotemporal position changes of key human nodes of different human targets in each video frame of the suspected bullying video clip to obtain the interaction anomaly index of the suspected bullying video clip; extract words from the suspected bullying audio clip, and obtain the semantic bullying index of the suspected bullying audio clip based on the correlation between the words and bullying-related words.
[0047] It should be noted that audio and video clips of suspected bullying incidents only indicate the possibility of abnormal gatherings or video anomalies, but cannot distinguish between normal playful roughhousing and malicious physical conflict, nor can they distinguish between camera malfunctions and deliberate obstruction. Considering that bullying behavior visually manifests as specific aggressive or oppressive spatiotemporal dynamics, and is accompanied by strong negative emotions and threatening language auditorily, this invention conducts in-depth analysis of audio and video clips of suspected bullying incidents, extracting refined features highly correlated with bullying behavior from three dimensions: visual, auditory, and contextual.
[0048] It should be further explained that in order to identify specific behaviors such as pushing and punching from the skeletal sequence, it is necessary to model the spatiotemporal relationships between key nodes. This invention quantifies the intensity of aggressive actions by calculating the speed and proximity of key nodes such as hands and feet relative to the core areas of other individuals' bodies.
[0049] Specifically, the spatiotemporal position changes of key human nodes of different human targets in each video frame of a suspected bullying video clip are analyzed to obtain the interaction anomaly index of the suspected bullying video clip, including:
[0050] Record any video frame of a suspected bullying incident as the target frame, and obtain the set of key human body nodes of any human target in the target frame.
[0051] For the i-th and c-th human targets in the target frame, obtain the hand-related nodes from the human body key node set of the i-th human target in the target frame, denoted as the hand point set of the i-th human target in the target frame; obtain the core-related nodes from the human body key node set of the i-th human target in the target frame, denoted as the core point set of the i-th human target in the target frame. It should be noted that, with Figure 2 Taking the 17 key nodes of the human body as an example, the hand point set includes 7, 8, 9, and 10, and the core point set includes 5, 6, 11, 12, 13, 14, 15, and 16.
[0052] The instantaneous velocity vector of the data points of the hand point set of the i-th human target in the target frame is obtained by taking the coordinates of any data point of the hand point set of the i-th human target in the target frame and the coordinates of the data points of the hand point set of the i-th human target in the previous frame, and dividing the vector by the interval between consecutive frames. Similarly, the instantaneous velocity vector of the data points of the core point set of the c-th human target in the target frame is obtained.
[0053] The instantaneous interaction index between any two human targets in the target frame satisfies the expression:
[0054] ;
[0055] In the formula, It represents the instantaneous interaction index between the i-th human target and the c-th human target in the target frame; This represents the set of hand points of the i-th human target in the target frame; This represents the set of core points of the c-th human target in the target frame; , express The Data points The The instantaneous velocity vector of each data point; express The Data points and The Euclidean distance between data points; This represents a minimum value, used to avoid a denominator of 0. For example, ; The symbol indicates the calculation of the magnitude of a vector.
[0056] In the formula, express The Data points The The magnitude of the difference between the instantaneous velocity vectors of the data points; the larger this value, the greater the magnitude. The Data points The The more inconsistent the directions of the data points are and the greater the instantaneous speed, the more it indicates that the hand of the i-th human target in the target frame is approaching the core area of the c-th human target at high speed, thus representing the instantaneous interaction index of the i-th and c-th human targets in the target frame. express The Data points The The smaller the distance between the data points, the closer the two human targets are, which means that the instantaneous speed of the data points has a greater impact on the offensiveness and the greater the damage that may be caused. Therefore, the instantaneous interaction index of the i-th human target and the c-th human target in the target frame is larger.
[0057] It should be noted that normal interaction between two human targets should be mutual. Therefore, if the instantaneous interaction index differs significantly between the two, and if the instantaneous interaction indices of multiple human targets in a suspected bullying video clip show significant differences and the interaction is directional, then the abnormal interaction of the suspected bullying event clip is more intense. Thus, this invention obtains the abnormal interaction index of a suspected bullying event clip based on the cumulative difference in the instantaneous interaction indices of each human target in the suspected bullying video clip.
[0058] The interaction anomaly index of the suspected bullying incident clips satisfies the expression:
[0059] ;
[0060] ;
[0061] In the formula, An abnormal interaction index indicating suspected bullying incident clips; This represents the set of human targets detected from suspected bullying incident footage; This indicates the degree of abnormality in the interaction between the i-th human target and the c-th human target in a suspected bullying incident segment; This indicates the number of video frames in a segment suspected of being a bullying incident; , Let i represent the instantaneous interaction index between the i-th human target and the c-th human target in the t-th frame, and let c represent the instantaneous interaction index between the c-th human target and the i-th human target in the target frame. This represents the maximum value function. It should be noted that t starts accumulating from 2 because the calculation of the instantaneous interaction index requires data from the frame preceding the t-th frame; if t=1, the instantaneous interaction index cannot be calculated.
[0062] In the formula, The difference between the instantaneous interaction index of the i-th human target to the c-th human target in the t-th frame and the instantaneous interaction index of the c-th human target to the i-th human target in the target frame represents the degree of interaction abnormality between the i-th human target and the c-th human target in the t-th frame. This represents the cumulative degree of abnormal interaction between the i-th human target and the c-th human target across all video frames. This represents the sum of the cumulative abnormality of interactions between all human targets and the c-th human target, reflecting the degree to which the c-th human target is being targeted. This represents the maximum degree to which all human targets are targeted. The larger the value, the greater the interaction anomaly index of the suspected bullying incident segment. This feature can not only reflect the severity of the suspected bullying incident segment, but also clearly point to the bullied object.
[0063] It should be noted that verbal bullying is an important form of bullying, and its identification relies on understanding the content of spoken language. Considering that speech in bullying scenarios often contains negative emotions and threatening words, this invention combines speech emotion and keywords for dual detection to establish an audio emotional semantic feature vector.
[0064] Preferably, the suspected bullying audio clips are subjected to word extraction, and a semantic bullying index is obtained based on the relevance of the extracted words to bullying-related words, including:
[0065] Suspected bullying audio clips are converted into text using automatic speech recognition (ASR). A word segmentation algorithm is then used to segment the text, yielding multiple words related to the suspected bullying audio clips. A bullying keyword lexicon is constructed, and a preset threat level is assigned to each bullying keyword. These preset threat levels are manually determined. Finally, a word vector model is used to convert the words and bullying keywords into word vectors. It should be noted that Automatic Speech Recognition (ASR), word segmentation algorithms, and word vector models are existing technologies and will not be elaborated upon here.
[0066] The semantic bullying index of the suspected bullying audio clip satisfies the expression:
[0067] ;
[0068] ;
[0069] In the formula, The semantic bullying index indicates the likelihood of audio clips being used to represent bullying. The number of words indicating suspected bullying audio clips; The bullying index contribution of the m-th word in the suspected bullying audio clip; The word vector representing the m-th word in a suspected bullying audio clip; A set of word vectors representing a keyword database for bullying; The word vector of the m-th word in the suspected bullying audio segment is represented by the set of cosine similarities between the word vectors of the bullying keyword lexicon and the set of word vectors of the bullying keyword lexicon. Represents the maximum value function; express The maximum value corresponds to the preset threat level of the bullying keyword; A represents the preset condition: the m-th word in the suspected bullying audio segment has the largest bullying index contribution among all words. One word; This indicates that condition A is satisfied; It is a negation symbol. This indicates that condition A is not met; For example, the default value is... It is 10.
[0070] In the formula, This reflects the maximum relevance between the m-th word and the bullying keyword lexicon. The larger the value, the closer the meaning of the m-th word is to bullying. Indicates to Filter by values, because The semantic bullying index is obtained by accumulating the bullying index contribution of all words. If the vocabulary is large, the overall semantic bullying index may be high. Therefore, the words are filtered by condition A, and only the N words with the largest bullying index contribution are accumulated. This makes the semantic bullying index focus more on the words with the highest degree of bullying and improves the accuracy of the semantic bullying index calculation.
[0071] At this point, the interaction anomaly index of the suspected bullying incident segment and the semantic bullying index of the suspected bullying audio segment were obtained.
[0072] S3: Based on the environmental noise of the video scene, dynamically weight the interaction anomaly index and semantic bullying index to obtain the comprehensive bullying index of suspected bullying event segments.
[0073] It should be noted that the Interaction Anomaly Index and the Semantic Bullying Index assess the likelihood of bullying behavior from visual and auditory dimensions, respectively. However, in a real school environment, the reliability and importance of these two modalities vary dynamically across different events. For example, on a noisy playground, physical conflict may be key evidence, while audio signals may be filled with noise; and in a quiet corridor, clear threatening language may reveal the nature of bullying more clearly than a slight shove. Therefore, this invention dynamically weights the Interaction Anomaly Index and the Semantic Bullying Index.
[0074] Specifically, based on the environmental noise level of the video scene, the interaction anomaly index and semantic bullying index are dynamically weighted to obtain a comprehensive bullying index for suspected bullying incident segments, including:
[0075] ;
[0076] In the formula, The overall bullying index indicates the amount of data captured in footage that appears to be a bullying incident; Indicates the signal-to-noise ratio of the audio clip suspected of being bullying; An abnormal interaction index indicating suspected bullying incident clips; The semantic bullying index indicates the likelihood of audio clips being used to represent bullying. This represents the normalization function.
[0077] In the formula, This indicates that the interaction anomaly index and semantic bullying index are weighted by the signal-to-noise ratio (SNR) of suspected bullying audio clips. A low SNR indicates poor audio quality. Larger The interaction anomaly index has a smaller weight, while the semantic bullying index has a smaller weight.
[0078] At this point, a comprehensive bullying index was obtained from the suspected bullying incident footage.
[0079] S4: Based on the comprehensive bullying index of suspected bullying incident fragments, the bullying incident is located, and intelligent monitoring, early warning and handling of bullying prevention on campus are completed.
[0080] Specifically, the identification, early warning, and handling of bullying incidents include:
[0081] Obtain the index value corresponding to the interaction anomaly index and the corresponding human target, and denote it as the core target to be attacked.
[0082] A verification dataset of labeled audio and video clips of school bullying incidents is collected and constructed, with labels including "normal activity" and "bullying incident". The comprehensive bullying index of all clips in the verification dataset is calculated. Thresholds are determined based on the ROC curve: the Receiver Operational Characteristics (ROC) curve of the system is plotted, showing the relationship between the proportion of bullying identified and the proportion of false positives for normal activities at all possible thresholds. A primary warning threshold and a secondary warning threshold are generated. For example, the primary warning threshold is the comprehensive bullying index corresponding to the point closest to the upper left corner (0,1) on the ROC curve, and the secondary warning threshold is the comprehensive bullying index corresponding to a true positive rate of 99%. When the comprehensive bullying index reaches the primary warning index, the system immediately pushes the location of the corresponding camera, information about the core attacked target, and real-time footage to school security personnel, homeroom teachers, and other responsible persons via SMS, APP, etc. When the comprehensive bullying index reaches the secondary warning index, the system adds a priority alarm to the push notifications to school security personnel, homeroom teachers, and other responsible persons.
[0083] This completes the intelligent monitoring, early warning, and response system for preventing bullying on campus.
[0084] While various embodiments of the invention have been shown and described in this specification, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Many modifications, alterations, and alternatives will occur to those skilled in the art without departing from the spirit and essence of the invention.
Claims
1. A campus intelligent monitoring method for anti-bullying in a campus, characterized in that, The method comprises the following steps: Acquire the audio and video data stream of the campus monitoring scene and perform edge preprocessing to screen out a suspected bullying event segment, wherein the suspected bullying event segment comprises a suspected bullying video segment and a suspected bullying audio segment; and record any video frame of the suspected bullying event segment as a target frame; Analyze the spatiotemporal position changes of the human body key nodes of different human body targets in each video frame of the suspected bullying video segment to obtain an interaction abnormality index of the suspected bullying video segment, wherein the interaction abnormality index is used to represent the cumulative situation of the difference between the instantaneous interaction indexes of any two human body targets; and perform vocabulary extraction on the suspected bullying audio segment to obtain a semantic bullying index of the suspected bullying audio segment according to the relevance of the vocabulary and bullying-related vocabulary; Dynamically weight the interaction abnormality index and the semantic bullying index according to the environmental noise of the video scene to obtain a comprehensive bullying index of the suspected bullying event segment; and position the bullying event based on the comprehensive bullying index of the suspected bullying event segment to complete the campus intelligent monitoring for campus anti-bullying. The instantaneous interaction exponent between any two human targets satisfies the expression: In the formula, It represents the instantaneous interaction index between the i-th human target and the c-th human target in the target frame; This represents the set of hand points of the i-th human target in the target frame; This represents the set of core points of the c-th human target in the target frame; , express The Data points The The instantaneous velocity vector of each data point; express The Data points and The Euclidean distance between data points; Indicates the minimum value; Indicates the sign for finding the magnitude of a vector; The interaction anomaly index satisfies the expression: ; in the formula, F represents the interaction anomaly index of the suspected bullying event segment; I represents a set of human body targets detected in the suspected bullying event segment; represents the interaction anomaly degree of the ith human body target to the cth human body target in the suspected bullying event segment, The acquisition manner of the interaction anomaly index is to accumulate the difference between the instantaneous interaction index of the ith human body target to the cth human body target in all video frames and the instantaneous interaction index of the cth human body target to the ith human body target in the target frame; represents the maximum function; The semantic bullying index of the suspected bullying audio clip satisfies the expression: In the formula, The semantic bullying index indicates the likelihood of audio clips being used to represent bullying. The number of words indicating suspected bullying audio clips; The bullying index contribution of the m-th word in the suspected bullying audio clip; express The maximum value corresponds to the preset threat level of the bullying keyword. The word vector of the m-th word in the suspected bullying audio segment is represented by the set of cosine similarities between the word vectors of the bullying keyword lexicon and the set of word vectors of the bullying keyword lexicon. a bullying index contribution degree of an mth vocabulary of the suspected bullying audio clip satisfies an expression: ; in the expression, represents a word vector of an mth vocabulary of the suspected bullying audio clip; represents a word vector set of the bullying keyword library; represents a preset condition: the mth vocabulary of the suspected bullying audio clip belongs to the maximum one of the bullying index contribution degrees of all the vocabularies; represents that the condition is satisfied; is a negative symbol, represents that the condition is not satisfied; is a preset value. 2.The campus intelligent monitoring method for anti-bullying in a campus according to claim 1, wherein, The edge preprocessing and screening to obtain the suspected bullying event segment comprises the following steps: Detect each human body target by using a lightweight target detection algorithm to obtain the bounding box of each human body target; and preliminarily screen out a video segment that meets any one of the following conditions based on the bounding box, and record the video segment as a suspected bullying video segment: there are multiple bounding boxes that overlap for more than a preset first threshold; and in the video frame in which no human body target is detected, there are continuous video frames with an information entropy lower than a preset second threshold and a duration longer than the preset first threshold. 3.The campus intelligent monitoring method for anti-bullying in a campus according to claim 1, wherein, The vocabulary extraction on the suspected bullying audio segment comprises the following steps: convert the suspected bullying audio segment into text by using an automatic speech recognition technology; perform word segmentation on the text by using a word segmentation algorithm to obtain multiple vocabularies of the suspected bullying audio segment; construct a bullying keyword library to obtain a preset threat level of each bullying keyword; and convert the vocabularies and the bullying keywords into word vectors by using a word vector model. 4.The campus intelligent monitoring method for anti-bullying in a campus according to claim 1, wherein, The comprehensive bullying index of the suspected bullying event segment is obtained by the following steps: obtain the signal-to-noise ratio of the suspected bullying audio segment; and perform weighted summation on the interaction abnormality index and the semantic bullying index according to the signal-to-noise ratio. 5.The campus intelligent monitoring method for anti-bullying in a campus according to claim 1, wherein, The positioning of the bullying event comprises the following steps: obtain an index value and a corresponding human body target of the interaction abnormality index, and record the human body target as a core attacked target; collect and construct a verification data set of the labeled campus bullying event audio and video segments to calculate the comprehensive bullying index of all segments in the verification data set; determine a warning threshold based on a ROC curve; and when the comprehensive bullying index reaches the warning threshold, the system pushes an alarm to the campus security or the class teacher in real time through a short message or a mobile application, wherein the alarm comprises the location of the camera, the information of the core attacked target and a real-time video picture.
6. The campus intelligent monitoring method for anti-bullying in a campus according to claim 5, characterized in that, The warning threshold comprises a first warning threshold and a second warning threshold, and the second warning threshold is greater than the first warning threshold; and when the comprehensive bullying index reaches the second warning threshold, the alarm pushed by the system to the campus security or the class teacher is marked as high priority to prompt priority processing.
Citation Information
Patent Citations
Campus bullying behavior detection method and system based on video phone voice recognition
CN119479692A
Campus bullying event detection method and device, electronic equipment and storage medium
CN119992443A