Campus bullying event detection method and device, electronic equipment and storage medium
By extracting and integrating the video and audio features in campus monitoring data and performing classification processing, the problem of insufficient detection accuracy of campus bullying incidents in the existing technology is solved, and more efficient early warning and prevention effects are achieved.
Patent Information
- Application Number
- CN202411976795.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-13
AI Technical Summary
When preventing campus bullying, it is difficult to accurately identify and detect conflicts between students, especially emerging bullying methods or highly concealed bullying behaviors, and the prediction accuracy is limited.
By obtaining monitoring data on campus, extracting video features and audio features, and weighted summing in the pre-constructed feature sharing subspace, fused features are obtained, and these features are used for classification processing to improve the detection accuracy of campus bullying events.
Accurate identification and detection of campus bullying incidents has been achieved, early warning capabilities have been improved, and campus bullying has been effectively prevented.
Smart Images

Figure CN119992443A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and in particular to a method, device, electronic device and storage medium for detecting campus bullying incidents. Background Art
[0002] In the campus environment, preventing school bullying has always been a key issue of concern to educational institutions and all sectors of society. Most of the existing technical solutions for preventing school bullying use knowledge graphs to predict possible bullying incidents by analyzing student relationships, student information, and establishing a bullying knowledge base. Although this method can improve the accuracy of predictions to a certain extent, there are still some problems. First, the prediction model based on historical records and knowledge bases often finds it difficult to capture the contradictions between students, so the accuracy of the prediction is limited. Secondly, this method can usually only predict known bullying behaviors, and it is often difficult to effectively identify and detect newly emerging bullying methods or more concealed bullying behaviors. Therefore, there is an urgent need for a new technical solution that can accurately identify, detect, and effectively prevent school bullying. Summary of the invention
[0003] The embodiment of the present invention provides a method for detecting campus bullying incidents, aiming to provide a solution that can accurately identify and detect campus bullying incidents, and can effectively prevent campus bullying from occurring. The present invention obtains monitoring data to be detected in the campus, extracts a first video feature and a first audio feature from the monitoring data, performs weighted summation of the first video feature and the first audio feature in a pre-constructed feature sharing subspace, obtains a fusion feature, and uses the fusion feature for classification processing, thereby improving the detection accuracy of campus bullying incidents, and can timely warn campus bullying incidents, effectively preventing the occurrence of campus bullying.
[0004] In a first aspect, an embodiment of the present invention provides a method for detecting campus bullying incidents, the method comprising the following steps:
[0005] Acquire monitoring data to be detected on campus, wherein the monitoring data to be detected includes multiple target persons;
[0006] Respectively extracting a first video feature and a first audio feature of the monitoring data to be detected;
[0007] Based on the first video feature and the first audio feature, determining a plurality of pairs of target features related to the school bullying incident and a weight combination corresponding to each pair of the target features in a pre-constructed feature sharing subspace, each pair of the target features comprising a video target feature and an audio target feature aligned in time, and the weight combination comprising a first weight corresponding to the video target feature and a second weight corresponding to the audio target feature;
[0008] Performing weighted sum processing on the target features based on the weight combination to obtain fused features, each fused feature corresponding to a pair of the target features;
[0009] Classify and process the plurality of fusion features to determine whether the campus bullying incident exists in the monitoring data to be detected.
[0010] Optionally, based on the first video feature and the first audio feature, determining multiple pairs of target features related to the school bullying incident and weight combinations corresponding to each pair of the target features in a pre-constructed feature sharing subspace includes:
[0011] Construct feature sharing subspace;
[0012] Based on linear discriminant analysis, mapping the first video feature and the first audio feature to the feature shared subspace to obtain a second video feature and a second audio feature, wherein the second video feature corresponds to the first video feature, and the second audio feature corresponds to the first audio feature;
[0013] Based on the second video feature and the second audio feature, multiple pairs of target features related to the school bullying incident and weight combinations corresponding to each pair of the target features are determined.
[0014] Optionally, determining multiple pairs of target features related to the school bullying incident and weight combinations corresponding to each pair of the target features based on the second video feature and the second audio feature includes:
[0015] According to the feature operator corresponding to the campus bullying incident, the second video feature and the second audio feature are respectively calculated and processed to obtain a plurality of video target features related to the campus bullying incident and a plurality of audio target features related to the campus bullying incident;
[0016] Determine one of the video target features and one of the audio target features aligned in the time dimension as a pair of target features related to the campus bullying incident;
[0017] For each pair of the target features, a weight combination corresponding to the target features is determined based on the video target feature and the audio target feature.
[0018] Optionally, for each pair of the target features, determining a weight combination corresponding to the target features based on the video target feature and the audio target feature includes:
[0019] For each pair of the target features, the video target feature and the audio target feature are respectively decomposed to obtain a first feature vector corresponding to the video target feature and a second feature vector corresponding to the audio target feature;
[0020] The first feature vector and the second feature vector are respectively normalized to obtain a first weight corresponding to the video target feature and a second weight corresponding to the audio target feature.
[0021] Optionally, performing weighted sum processing on the target features based on the weight combination to obtain fused features includes:
[0022] For each pair of the target features, weighting the video target features by a first weight, and weighting the audio target features by a second weight;
[0023] The weighted video target feature and the audio target feature are added together to obtain a fusion feature.
[0024] Optionally, the classifying and processing the plurality of fusion features to determine whether the campus bullying incident exists in the monitoring data to be detected includes:
[0025] Based on the behavioral characteristics related to the campus bullying incident, a plurality of the fusion features are classified and processed to obtain a campus bullying incident score;
[0026] Based on the school bullying incident score, determine whether the school bullying incident exists in the monitoring data to be detected.
[0027] Optionally, the determining whether the school bullying incident exists in the monitoring data to be detected based on the school bullying incident score includes:
[0028] If the school bullying incident score is within the first score range, determining that a school bullying incident of the first emergency type exists in the monitoring data to be detected;
[0029] If the school bullying incident score is within a second score range, it is determined that a school bullying incident of a second emergency type exists in the monitoring data to be detected, and the maximum score boundary of the second score range is less than the minimum score boundary of the first classification range;
[0030] If the school bullying incident score is within a third score range, it is determined that a school bullying incident of a third emergency type exists in the monitoring data to be detected, and the maximum score boundary of the third score range is less than the minimum score boundary of the second classification range;
[0031] If the school bullying incident score is less than the minimum score boundary of the third score range, it is determined that the school bullying incident does not exist in the monitoring data to be detected.
[0032] In a second aspect, an embodiment of the present invention further provides a campus bullying incident detection device, the campus bullying incident detection device comprising:
[0033] An acquisition module is used to acquire monitoring data to be detected on campus, wherein the monitoring data to be detected includes multiple target persons;
[0034] A first processing module, used for respectively extracting a first video feature and a first audio feature of the monitoring data to be detected;
[0035] A second processing module is used to determine, based on the first video feature and the first audio feature, a plurality of pairs of target features related to the school bullying incident and a weight combination corresponding to each pair of the target features in a pre-constructed feature sharing subspace, each pair of the target features includes a video target feature and an audio target feature that are aligned in time, and the weight combination includes a first weight corresponding to the video target feature and a second weight corresponding to the audio target feature;
[0036] A third processing module is used to perform weighted sum processing on the target features based on the weight combination to obtain fused features, each fused feature corresponding to a pair of the target features;
[0037] The fourth processing module is used to classify the multiple fusion features to determine whether the school bullying incident exists in the monitoring data to be detected.
[0038] In a third aspect, an embodiment of the present invention provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of the campus bullying incident detection method provided in the embodiment of the present invention are implemented.
[0039] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the campus bullying incident detection method provided in the embodiment of the invention are implemented.
[0040] In an embodiment of the present invention, monitoring data to be detected in a campus is obtained, and the monitoring data to be detected includes multiple target persons; the first video feature and the first audio feature of the monitoring data to be detected are respectively extracted; based on the first video feature and the first audio feature, multiple pairs of target features related to campus bullying incidents and weight combinations corresponding to each pair of target features are determined in a pre-constructed feature sharing subspace, each pair of target features includes a video target feature and an audio target feature aligned in time, and the weight combination includes a first weight corresponding to the video target feature and a second weight corresponding to the audio target feature; based on the weight combination, the target features are weighted and summed to obtain fusion features, each fusion feature corresponds to a pair of the target features; multiple fusion features are classified and processed to determine whether there is a campus bullying incident in the monitoring data to be detected. The present invention obtains monitoring data to be detected in a campus, extracts the first video feature and the first audio feature from the monitoring data, performs weighted summation on the first video feature and the first audio feature in a pre-constructed feature sharing subspace, obtains the fusion feature, and uses the fusion feature for classification processing, thereby improving the detection accuracy of campus bullying incidents, and can timely warn campus bullying incidents, effectively preventing the occurrence of campus bullying. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0042] Figure 1 is a flow chart of a method for detecting campus bullying incidents provided by an embodiment of the present invention;
[0043] Figure 2 It is a structural schematic diagram of a campus bullying incident detection device provided by an embodiment of the present invention;
[0044] Figure 3 It is a structural schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0045] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0046] like Figure 1As shown, Figure 1 : is a flow chart of a campus bullying incident detection method provided by an embodiment of the present invention, the campus bullying incident detection method comprises the steps of:
[0047] 101. Obtain the monitoring data to be tested on campus.
[0048] In an embodiment of the present invention, the campus bullying incident detection method can be applied to an event detection platform, which can be constructed based on a server or a distributed server, and the event detection platform includes a data interface (available for user upload or monitoring device upload), an event detection model, a lightweight classification model, and a campus bullying incident detection program. The data interface can be used to obtain data to be processed or training data, and the data to be processed can be data of multiple modes, and the campus bullying incident detection program can be used to implement each step of the campus bullying incident detection method.
[0049] The monitoring data to be detected can be collected by monitoring equipment installed on campus. The monitoring equipment can be installed at a location that is allowed to be installed on campus. The embodiment of the present invention does not limit the location and method of installing the monitoring equipment. The monitoring equipment can collect image data and audio data as monitoring data, and upload the collected monitoring data to the event detection platform through a data interface.
[0050] The above-mentioned monitoring data to be detected includes multiple target persons, and the above-mentioned target persons are students. Personnel detection can be performed on the monitoring data. When two or more student persons are detected in the monitoring data, the monitoring data when two or more student persons are detected can be used as the monitoring data to be detected.
[0051] The above-mentioned campus bullying incident detection method can be used in campus management scenarios to detect campus bullying incidents.
[0052] 102. Extract first video features and first audio features of the monitoring data to be detected respectively.
[0053] In the embodiment of the present invention, the feature extraction module within the multimodal large model can be used to perform video feature extraction and audio feature extraction on the monitoring data to be detected, so as to obtain the first video feature and the first audio feature of the monitoring data to be detected.
[0054] The monitoring data to be detected can be split according to data modality to obtain video monitoring data of image modality and audio monitoring data of audio modality. Feature extraction is performed on the video monitoring data through the feature extraction module inside the multimodal large model to obtain the first video feature of the monitoring data to be detected. Feature extraction is performed on the audio monitoring data through the feature extraction module inside the multimodal large model to obtain the first audio feature of the monitoring data to be detected.
[0055] Of course, in a possible embodiment, the feature extraction module can also be trained by itself, rather than being a feature extraction module inside a multimodal large model. The role of using a feature extraction module inside a multimodal large model is that through the self-learning ability of the multimodal large model, it can continuously learn the features required for campus bullying incidents, and then extract more accurate video features and audio features. The role of using a self-trained feature extraction module is that it has stronger pertinence and stronger interpretability.
[0056] In a possible embodiment, feature extraction of video surveillance data may include spatiotemporal feature extraction and behavioral feature extraction. Further, the feature extraction module may include a 3D convolutional neural network (3D CNN) and a deep learning model P3D+LSTM network. The 3D convolutional neural network may be used to extract spatiotemporal features and a deep learning model P3D+LSTM network to extract action features such as waving, punching, collision, pushing, quarreling, etc., thereby obtaining a first video feature including spatiotemporal features and action features.
[0057] In a possible embodiment, feature extraction of audio monitoring data may include audio preprocessing and emotional feature extraction. Further, the audio monitoring data is converted from an audio signal into a digital signal, noise and interference in the audio are removed, and the audio signal is divided into multiple short segments for local feature extraction. Mel spectrum analysis MFCC is used to extract emotional features in the audio. Emotional features may include: insults, anger, sadness, anxiety, etc. Specifically, a deep learning model LSTM (Long Short-Term Memory Network) can be used to capture changes in the psychological state of the target person in the audio monitoring data, thereby obtaining a first audio feature including language features and emotional features.
[0058] 103. Based on the first video feature and the first audio feature, multiple pairs of target features related to campus bullying incidents and weight combinations corresponding to each pair of target features are determined in a pre-constructed feature sharing subspace.
[0059] In an embodiment of the present invention, after obtaining the first video feature and the first audio feature, since the first video feature and the first audio feature are features extracted from different modal data and their dimensions are greatly different, the first video feature and the first audio feature can be mapped to a pre-constructed feature shared subspace, so that the first video feature and the first audio feature can be aligned in dimension, which facilitates the fusion of the video feature map and the audio feature.
[0060] The above-mentioned target features related to campus bullying incidents can be the pre-behavior features of bullying incidents, such as waving, punching, collision, pushing, quarreling and other action behavior features, as well as insults, anger, sadness, anxiety and other emotional behavior features. Action behavior features such as waving, punching, collision, pushing, quarreling and other action behavior features can be further extracted from the first video feature, and emotional behavior features such as insults, anger, sadness, anxiety and other emotional behavior features can be extracted from the first audio feature. The action behavior features and emotional behavior features can be matched in time first, and the successfully matched action behavior features and emotional behavior features can be combined into a pair of target features, and the condition for the above successful matching is the same time. For each action behavior feature that fails to match, a segment of audio features of the same time can be matched in the first audio feature to extract emotional behavior features, and the extracted emotional behavior features can be used as audio target features to form a pair of target features. At this time, the audio target feature can be a low-confidence feature, and the low-confidence feature can be used to indicate that the reliability of the audio target feature as an emotional behavior feature related to campus bullying incidents is low, and it can be considered as an emotional behavior feature that is not related to campus bullying incidents. For each emotional behavior feature that fails to match, a video feature of the same time can also be matched in the first video feature to extract the action behavior feature, and the extracted action behavior feature can be used as the video target feature. At this time, the above-mentioned video target feature can be a low-confidence feature. The low-confidence feature can be used to indicate that the reliability of the video target feature as an action behavior feature related to campus bullying incidents is low, and it can be considered as an action behavior feature that is not related to campus bullying incidents.
[0061] It should be noted that each pair of target features includes a video target feature and an audio target feature that are aligned in time, and the weight combination includes a first weight corresponding to the video target feature and a second weight corresponding to the audio target feature.
[0062] The first weight and the second weight can be set manually or dynamically adjusted according to the video target feature and the audio target feature. For example, if the video target feature and the audio target feature in the target feature are both features related to campus bullying incidents, the first weight is a0, the second weight is b0, and a0 and b0 belong to the interval [0, 1]; if the video target feature in the target feature is a feature related to campus bullying incidents, and the audio target feature is a feature not related to campus bullying incidents, the first weight is a1, the second weight is b1, b1 is less than b0, and a1 and b1 belong to the interval [0, 1]; if the video target feature in the target feature is a feature not related to campus bullying incidents, and the audio target feature is a feature related to campus bullying incidents, the first weight is a2, the second weight is b2, a2 is less than a0, and a2 and b2 belong to the interval [0, 1].
[0063] 104. Perform weighted summation on the target features based on the weight combination to obtain fused features, each fused feature corresponding to a pair of target features.
[0064] In an embodiment of the present invention, after obtaining the target features and the weight combination, the first weight can be used to perform weighted calculation on the video target features in the target features to obtain the weighted video features, and the second weight can be used to perform weighted calculation on the audio target features in the target features to obtain the weighted audio features. The weighted video features and the weighted audio features are added and fused to form fused features. Each pair of target features can be processed by weighted summation to obtain a fused feature. The fused feature contains action features in a visual sense and emotional features in a linguistic sense, which can better express whether there is a possibility of school bullying between the target persons. If there are multiple pairs of target features, there are multiple fused features, that is, one monitoring data to be detected can correspond to multiple fused features.
[0065] 105. Classify and process multiple fusion features to determine whether there are campus bullying incidents in the monitoring data to be detected.
[0066] In the embodiment of the present invention, multiple fused features can be classified independently, and finally the multiple classification results are integrated to determine whether there is a school bullying incident in the monitoring data to be detected. Multiple fused features can also be spliced into a feature to be classified, and the feature to be classified is classified, and then the classification result is used to determine whether there is a school bullying incident in the monitoring data to be detected.
[0067] When multiple fusion features are classified and processed independently, it can be achieved through the classification module of the large language model or through a self-trained classification model. The classification model is used to output the event scores corresponding to each fusion feature. Each fusion feature corresponds to an event score. The average event score of all event scores can be calculated. The higher the average event score, the higher the probability that there is a school bullying incident in the monitoring data to be detected. When the average event score is higher than a preset average score threshold, it can be determined that there is a school bullying incident in the monitoring data to be detected. All scores can also be integrated. The higher the integral, the higher the probability that there is a school bullying incident in the monitoring data to be detected. When the integral is higher than a preset integral, it can be determined that there is a school bullying incident in the monitoring data to be detected.
[0068] When it is determined that there is a campus bullying incident in the monitoring data to be detected, the monitoring data to be detected and the detection results of the campus bullying incident can be used as push data and pushed to relevant managers, such as the target person’s teacher or the person in charge of campus bullying incidents on campus. The managers can timely find the students involved in the campus bullying incident based on the monitoring data to be detected and the detection results of the campus bullying incident, and then take corresponding measures.
[0069] In an embodiment of the present invention, monitoring data to be detected in a campus is obtained, and the monitoring data to be detected includes multiple target persons; the first video feature and the first audio feature of the monitoring data to be detected are respectively extracted; based on the first video feature and the first audio feature, multiple pairs of target features related to campus bullying incidents and weight combinations corresponding to each pair of target features are determined in a pre-constructed feature sharing subspace, each pair of target features includes a video target feature and an audio target feature aligned in time, and the weight combination includes a first weight corresponding to the video target feature and a second weight corresponding to the audio target feature; based on the weight combination, the target features are weighted and summed to obtain fusion features, each fusion feature corresponds to a pair of the target features; multiple fusion features are classified and processed to determine whether there is a campus bullying incident in the monitoring data to be detected. The present invention obtains monitoring data to be detected in a campus, extracts the first video feature and the first audio feature from the monitoring data, performs weighted summation on the first video feature and the first audio feature in a pre-constructed feature sharing subspace, obtains the fusion feature, and uses the fusion feature for classification processing, thereby improving the detection accuracy of campus bullying incidents, and can timely warn campus bullying incidents, effectively preventing the occurrence of campus bullying.
[0070] It is understandable that in the specific implementation of this application, it involves monitoring data, student data, teacher data and other related data. When the embodiments in this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data, as well as the training, deployment and calling of algorithm models, must comply with the relevant laws, regulations and standards of relevant countries and regions.
[0071] Optionally, in the step of determining, in a pre-constructed feature sharing subspace based on the first video feature and the first audio feature, multiple pairs of target features related to school bullying incidents and the weight combination corresponding to each pair of target features, a feature sharing subspace can be constructed; based on linear discriminant analysis, the first video feature and the first audio feature are mapped to the feature sharing subspace to obtain a second video feature and a second audio feature, the second video feature corresponds to the first video feature, and the second audio feature corresponds to the first audio feature; based on the second video feature and the second audio feature, multiple pairs of target features related to school bullying incidents and the weight combination corresponding to each pair of target features are determined.
[0072] In the embodiment of the present invention, it can be understood that the above-mentioned feature sharing subspace is a data representation space, in which different types of features (such as video features and audio features) can be converted and represented into a unified dimension so that feature data of different modalities can be directly compared and combined in this common space.
[0073] Linear discriminant analysis (LDA) is a dimensionality reduction technique that finds a projection direction so that sample projection points of the same type are as close as possible, while sample projection points of different types are as far apart as possible. In an embodiment of the present invention, linear discriminant analysis can be used to project the first video feature and the first audio feature into a feature sharing subspace, thereby obtaining the corresponding second video feature and the second audio feature.
[0074] After obtaining the second video feature and the second audio feature, action behavior features such as waving, punching, collision, pushing, quarreling, etc. can be further extracted from the second video feature, and emotional behavior features such as insults, anger, sadness, anxiety, etc. can be extracted from the second audio feature. The action behavior features and emotional behavior features can be matched in time first, and the successfully matched action behavior features and emotional behavior features are combined into a pair of target features. The condition for the above successful matching is that the time is the same or the intersecting time accounts for more than 1 / n of the total time, and n is greater than 1. For each action behavior feature that fails to match, an audio feature of the same time can be matched in the second audio feature to extract the emotional behavior feature, and the extracted emotional behavior feature is used as the audio target feature to form a pair of target features. At this time, the audio target feature can be a low-confidence feature. The low-confidence feature can be used to indicate that the reliability of the audio target feature as an emotional behavior feature related to campus bullying incidents is low, and it can be considered as an emotional behavior feature that is not related to campus bullying incidents. For each emotional behavior feature that fails to match, a video feature of the same time can also be matched in the second video feature to extract the action behavior feature, and the extracted action behavior feature can be used as the video target feature. At this time, the above-mentioned video target feature can be a low-confidence feature. The low-confidence feature can be used to indicate that the reliability of the video target feature as an action behavior feature related to campus bullying incidents is low, and it can be considered as an action behavior feature that is not related to campus bullying incidents.
[0075] The above-mentioned first weight and the above-mentioned second weight can be set manually, or can be obtained by dynamically adjusting according to the video target features and the audio target features. After obtaining the target features, the video target features and the audio target features in the target features can be weighted. Specifically, the first weight value of the video target features can be determined according to the correlation between the video target features and the campus bullying incident features. The higher the correlation between the video target features and the campus bullying incident features, the greater the first weight value of the video target features. The second weight value of the audio target features can be determined according to the correlation between the audio target features and the campus bullying incident features. The higher the correlation between the audio target features and the campus bullying incident features, the greater the second weight value of the audio target features.
[0076] Optionally, in the step of determining multiple pairs of target features related to the campus bullying incidents and the weight combination corresponding to each pair of target features based on the second video feature and the second audio feature, the second video feature and the second audio feature can be respectively calculated and processed according to the feature operator corresponding to the campus bullying incident to obtain multiple video target features related to the campus bullying incident and multiple audio target features related to the campus bullying incident; a video target feature and an audio target feature aligned in the time dimension are determined as a pair of target features related to the campus bullying incident; for each pair of target features, the weight combination corresponding to the target features is determined based on the video target feature and the audio target feature.
[0077] In an embodiment of the present invention, the above-mentioned feature operator may correspond to behaviors related to campus bullying incidents, such as waving, punching, colliding, pushing, quarreling and other action behavior characteristics, as well as insults, anger, sadness, anxiety and other emotional behavior characteristics, and each behavior corresponds to an operator. Therefore, there may be multiple feature operators related to campus bullying incidents in the feature sharing subspace. The second video feature is calculated and processed by the feature operator to determine whether the second video feature is a behavioral feature related to the campus bullying incident. If so, the second video feature will be determined as a video target feature, if not, it will no longer be concerned. The second audio feature is calculated and processed by the feature operator to determine whether the second audio feature is a behavioral feature related to the campus bullying incident. If so, the second audio feature will be determined as an audio target feature, if not, it will no longer be concerned.
[0078] For video target features and audio target features, their corresponding appearance times in the monitoring data to be detected can be obtained, and the video target features and audio target features with the same appearance time can be determined as a successful match and a pair of target features, or the video target features and audio target features with a high degree of intersection in appearance time can be determined as a successful match, such as the intersection time of the appearance time of the video target feature and the audio target feature accounts for 1 / n or more of the sum of the appearance time of the video target feature and the audio target feature, and n is greater than 1. In this case, the video target feature and the audio target feature can be regarded as a pair of target features. The video target features and audio target features with a low degree of intersection in appearance time or no intersection in appearance time can be determined as a failed match, such as the intersection time of the appearance time of the video target feature and the audio target feature accounts for less than 1 / n of the sum of the appearance time of the video target feature and the audio target feature. In this case, the video target features that failed to match can be matched with a segment of audio features with the same appearance time in the second audio feature as the audio target feature corresponding to the video target feature, and the audio target features that failed to match can be matched with a segment of video features with the same appearance time in the second video feature as the video target feature corresponding to the audio target feature.
[0079] The above-mentioned first weight and the above-mentioned second weight can be set manually, or can be obtained by dynamically adjusting according to the video target features and the audio target features. After obtaining the target features, the video target features and the audio target features in the target features can be weighted. Specifically, the first weight value of the video target features can be determined according to the correlation between the video target features and the campus bullying incident features. The higher the correlation between the video target features and the campus bullying incident features, the greater the first weight value of the video target features. The second weight value of the audio target features can be determined according to the correlation between the audio target features and the campus bullying incident features. The higher the correlation between the audio target features and the campus bullying incident features, the greater the second weight value of the audio target features.
[0080] Optionally, in the step of determining the weight combination corresponding to the target features based on the video target features and the audio target features for each pair of target features, the video target features and the audio target features can be decomposed separately for each pair of target features to obtain a first feature vector corresponding to the video target feature and a second feature vector corresponding to the audio target feature; the first feature vector and the second feature vector can be normalized separately to obtain a first weight corresponding to the video target feature and a second weight corresponding to the audio target feature.
[0081] In an embodiment of the present invention, for each pair of target features, there is a video target feature and an audio target feature, which can be calculated by the characteristic equation A(x)=λx, wherein A(x) is the feature to be decomposed, λ is the eigenvalue, and x is the eigenvector. When solving the characteristic equation, the eigenvalue λ is solved with the goal of maximizing the eigenvalue λ, and the corresponding eigenvector when the eigenvalue λ reaches the maximum value is used as the target feature vector. The video target feature pair can be decomposed by the above characteristic equation, and the video target feature is decomposed into an eigenvalue and an eigenvector. When the eigenvalue reaches the maximum value, the eigenvector is used as the first eigenvector corresponding to the video target feature. Similarly, the audio target feature pair can be decomposed by the above characteristic equation, and the audio target feature is decomposed into an eigenvalue and an eigenvector. When the eigenvalue reaches the maximum value, the eigenvector is used as the second eigenvector corresponding to the audio target feature.
[0082] After obtaining the first eigenvector and the second eigenvector, normalization is performed on the first eigenvector and the second eigenvector respectively. Through normalization, a value between [0, 1] is obtained as a weight value, that is, the first eigenvector is normalized to obtain a first weight, and the second eigenvector is normalized to obtain a second weight. Through normalization, the vector can be converted into a scalar, thereby measuring the importance between the first eigenvector and the second eigenvector. If the first weight is greater than the second weight, it means that the video target feature is more important than the audio target feature. If the second weight is greater than the first weight, it means that the audio target feature is more important than the video target feature.
[0083] In a possible embodiment, the normalization process can be to calculate the modulus of the feature vector itself, that is, to determine the modulus of the first feature vector as the first weight, and the modulus of the second feature vector as the second weight, and then obtain the weight combination corresponding to the target feature. It should be noted that the larger the modulus, the more relevant it is to campus bullying incidents, and the smaller the modulus, the less relevant it is to campus bullying incidents.
[0084] In a possible embodiment, the modulus of the video target feature may be directly calculated and determined as the first weight, and the modulus of the audio target feature may be calculated and determined as the second weight.
[0085] Optionally, in the step of performing weighted summation processing on the target features based on the weight combination to obtain the fused features, for each pair of target features, the video target features can be weighted by the first weight, and the audio target features can be weighted by the second weight; the weighted video target features and audio target features are added to obtain the fused features.
[0086] In the embodiment of the present invention, after obtaining the target features and the weight combination, the video target features in the target features can be weighted by the first weight, and the audio target features in the target features can be weighted by the second weight, and the weighted video target features and audio target features are added. The weighted processing is a multiplication calculation. Specifically, it can be expressed as:
[0087] A mix =w1A1+w2A2
[0088] Among them, A mix is the fusion feature, w1 is the first weight, A1 is the video target feature, w2 is the first weight, and A2 is the audio target feature.
[0089] The video target features and the audio target features are weighted and fused through the first weight and the second weight, so that the features related to campus bullying incidents are more prominent, thereby improving the detection accuracy of campus bullying incidents.
[0090] Optionally, in the step of classifying and processing multiple fused features to determine whether there is a school bullying incident in the monitoring data to be detected, the multiple fused features can be classified based on behavioral characteristics related to the school bullying incident to obtain a school bullying incident score; based on the school bullying incident score, it is determined whether there is a school bullying incident in the monitoring data to be detected.
[0091] In the embodiment of the present invention, when multiple fusion features are independently classified, it can be implemented through the classification module of the large language model, and can also be implemented through a self-trained classification model.
[0092] The above classification module is provided with multiple feature classifiers, each of which has a reference feature obtained through training. One feature classifier corresponds to one reference feature, and one feature classifier corresponds to one behavior category. The classification model will calculate the similarity between each fusion feature and each feature classifier. If the similarity between a fusion feature and a feature classifier is greater than or equal to a threshold, it means that the fusion feature belongs to the behavior category corresponding to the feature classifier, such as punching, pushing, collision, insulting, emotional anger, quarreling, sadness and other behavior categories. If the similarity between a fusion feature and a feature classifier is less than the threshold, it means that the fusion feature does not belong to the behavior category corresponding to the feature classifier. Finally, the score of campus bullying incidents is calculated according to the behavior categories to which all fusion features belong. For example, if the number of behavior categories corresponding to the feature classifier in the fusion feature is greater, the score of campus bullying incidents is higher. If the number of behavior categories corresponding to the feature classifier in the fusion feature is smaller, the score of campus bullying incidents is lower.
[0093] In a possible embodiment, the above-mentioned behavior categories can also be divided into emergency type, general type and minor type, wherein the emergency type indicates that the nature of the school bullying incident is serious, the general type indicates that the nature of the school bullying incident is general, and the minor type indicates that the nature of the school bullying incident is relatively minor. Features such as punching, pushing, collision, insulting, and emotional anger can be classified as emergency types, features such as quarrels and sadness can be classified as general types, and other features can be classified as minor types.
[0094] Optionally, in the step of determining whether there is a school bullying incident in the monitoring data to be detected based on the school bullying incident score, if the school bullying incident score is in a first score range, it is determined that there is a first emergency type of school bullying incident in the monitoring data to be detected; if the school bullying incident score is in a second score range, it is determined that there is a second emergency type of school bullying incident in the monitoring data to be detected, and the maximum score boundary of the second score range is less than the minimum score boundary of the first classification range; if the school bullying incident score is in a third score range, it is determined that there is a third emergency type of school bullying incident in the monitoring data to be detected, and the maximum score boundary of the third score range is less than the minimum score boundary of the second classification range; if the school bullying incident score is less than the minimum score boundary of the third score range, it is determined that there is no school bullying incident in the monitoring data to be detected.
[0095] In an embodiment of the present invention, school bullying incidents can be divided according to emergency types, and different emergency types correspond to different score ranges. The score range corresponding to the first emergency type is the highest, indicating that a school bullying incident of high severity is detected, the score range corresponding to the second emergency type is below the score range corresponding to the first emergency type, indicating that a school bullying incident of average severity is detected, and the score range corresponding to the third emergency type is below the score range corresponding to the second emergency type, indicating that a school bullying incident of low severity is detected. If the score is lower than the third score range, it can be determined that there is no school bullying incident in the monitoring data to be detected.
[0096] Different early warning strategies can be designed according to different emergency types. For example, for the first emergency type of school bullying incidents, on-site prevention measures can be taken immediately, for the second emergency type, inquiry measures can be taken, and for the third emergency type, visiting measures can be taken.
[0097] Different emergency types of campus bullying incidents and monitoring data to be detected are pushed to the event task. The first emergency type of tasks are filtered out through the event task, and the first emergency type of tasks are issued device warnings and pushed to the teachers of the target students.
[0098] like Figure 2 As shown, an embodiment of the present invention provides a campus bullying incident detection device, the campus bullying incident detection device comprising:
[0099] An acquisition module 201 is used to acquire monitoring data to be detected on campus, wherein the monitoring data to be detected includes multiple target persons;
[0100] A first processing module 202 is used to extract a first video feature and a first audio feature of the monitoring data to be detected respectively;
[0101] A second processing module 203 is used to determine, based on the first video feature and the first audio feature, multiple pairs of target features related to the school bullying incident and a weight combination corresponding to each pair of the target features in a pre-constructed feature sharing subspace, each pair of the target features includes a video target feature and an audio target feature that are aligned in time, and the weight combination includes a first weight corresponding to the video target feature and a second weight corresponding to the audio target feature;
[0102] A third processing module 204 is used to perform weighted sum processing on the target features based on the weight combination to obtain fused features, each fused feature corresponding to a pair of the target features;
[0103] The fourth processing module 205 is used to classify the multiple fusion features to determine whether the school bullying incident exists in the monitoring data to be detected.
[0104] Optionally, the second processing module 203 is also used to construct a feature sharing subspace; based on linear discriminant analysis, the first video feature and the first audio feature are mapped to the feature sharing subspace to obtain a second video feature and a second audio feature, the second video feature corresponds to the first video feature, and the second audio feature corresponds to the first audio feature; based on the second video feature and the second audio feature, multiple pairs of target features related to the school bullying incident and the weight combination corresponding to each pair of the target features are determined.
[0105] Optionally, the second processing module 203 is also used to calculate and process the second video feature and the second audio feature respectively according to the feature operator corresponding to the school bullying incident, so as to obtain multiple video target features related to the school bullying incident and multiple audio target features related to the school bullying incident; determine one of the video target features and one of the audio target features aligned in the time dimension as a pair of target features related to the school bullying incident; for each pair of the target features, determine the weight combination corresponding to the target features based on the video target feature and the audio target feature.
[0106] Optionally, the second processing module 203 is also used to decompose the video target feature and the audio target feature for each pair of target features, respectively, to obtain a first feature vector corresponding to the video target feature, and a second feature vector corresponding to the audio target feature; and to normalize the first feature vector and the second feature vector, respectively, to obtain a first weight corresponding to the video target feature and a second weight corresponding to the audio target feature.
[0107] Optionally, the third processing module 204 is also used to weight the video target features by a first weight and weight the audio target features by a second weight for each pair of target features; and add the weighted video target features and the audio target features to obtain a fusion feature.
[0108] Optionally, the fourth processing module 205 is also used to classify the multiple fusion features based on behavioral characteristics related to the school bullying incident to obtain a school bullying incident score; based on the school bullying incident score, determine whether the school bullying incident exists in the monitoring data to be detected.
[0109] Optionally, the fourth processing module 205 is also used to determine that a first emergency type of school bullying incident exists in the monitoring data to be detected if the school bullying incident score is in a first score range; if the school bullying incident score is in a second score range, it is determined that a second emergency type of school bullying incident exists in the monitoring data to be detected, and the maximum score boundary of the second score range is less than the minimum score boundary of the first classification range; if the school bullying incident score is in a third score range, it is determined that a third emergency type of school bullying incident exists in the monitoring data to be detected, and the maximum score boundary of the third score range is less than the minimum score boundary of the second classification range; if the school bullying incident score is less than the minimum score boundary of the third score range, it is determined that the school bullying incident does not exist in the monitoring data to be detected.
[0110] like Figure 3 As shown, an embodiment of the present invention further provides an electronic device, including a processor, and the processor can execute any one of the above-mentioned campus bullying incident detection methods.
[0111] Specifically, it includes a processor 301 and a memory 302, and a computer program for executing a campus bullying incident detection method stored in the memory 302 and capable of running on the processor 301, wherein:
[0112] The processor 301 runs the computer program of the campus bullying incident detection method stored in the memory 302 and performs the following steps:
[0113] Acquire monitoring data to be detected on campus, wherein the monitoring data to be detected includes multiple target persons;
[0114] Respectively extracting a first video feature and a first audio feature of the monitoring data to be detected;
[0115] Based on the first video feature and the first audio feature, determining a plurality of pairs of target features related to the school bullying incident and a weight combination corresponding to each pair of the target features in a pre-constructed feature sharing subspace, each pair of the target features comprising a video target feature and an audio target feature aligned in time, and the weight combination comprising a first weight corresponding to the video target feature and a second weight corresponding to the audio target feature;
[0116] Performing weighted sum processing on the target features based on the weight combination to obtain fused features, each fused feature corresponding to a pair of the target features;
[0117] Classify and process the plurality of fusion features to determine whether the campus bullying incident exists in the monitoring data to be detected.
[0118] Optionally, the processor 301 determines, based on the first video feature and the first audio feature, a plurality of pairs of target features related to the school bullying incident and a weight combination corresponding to each pair of the target features in a pre-constructed feature sharing subspace, including:
[0119] Construct feature sharing subspace;
[0120] Based on linear discriminant analysis, mapping the first video feature and the first audio feature to the feature shared subspace to obtain a second video feature and a second audio feature, wherein the second video feature corresponds to the first video feature, and the second audio feature corresponds to the first audio feature;
[0121] Based on the second video feature and the second audio feature, multiple pairs of target features related to the school bullying incident and weight combinations corresponding to each pair of the target features are determined.
[0122] Optionally, the processor 301 determines, based on the second video feature and the second audio feature, a plurality of pairs of target features related to the school bullying incident and a weight combination corresponding to each pair of the target features, including:
[0123] According to the feature operator corresponding to the campus bullying incident, the second video feature and the second audio feature are respectively calculated and processed to obtain a plurality of video target features related to the campus bullying incident and a plurality of audio target features related to the campus bullying incident;
[0124] Determine one of the video target features and one of the audio target features aligned in the time dimension as a pair of target features related to the campus bullying incident;
[0125] For each pair of the target features, a weight combination corresponding to the target features is determined based on the video target feature and the audio target feature.
[0126] Optionally, the processor 301 performs, for each pair of the target features, determining, based on the video target feature and the audio target feature, a weight combination corresponding to the target features, including:
[0127] For each pair of the target features, the video target feature and the audio target feature are respectively decomposed to obtain a first feature vector corresponding to the video target feature and a second feature vector corresponding to the audio target feature;
[0128] The first feature vector and the second feature vector are respectively normalized to obtain a first weight corresponding to the video target feature and a second weight corresponding to the audio target feature.
[0129] Optionally, the processor 301 performs weighted summation processing on the target features based on the weight combination to obtain a fusion feature, including:
[0130] For each pair of the target features, weighting the video target features by a first weight, and weighting the audio target features by a second weight;
[0131] The weighted video target feature and the audio target feature are added together to obtain a fusion feature.
[0132] Optionally, the processor 301 performs classification processing on the plurality of fusion features to determine whether the campus bullying incident exists in the monitoring data to be detected, including:
[0133] Based on the behavioral characteristics related to the campus bullying incident, a plurality of the fusion features are classified and processed to obtain a campus bullying incident score;
[0134] Based on the school bullying incident score, determine whether the school bullying incident exists in the monitoring data to be detected.
[0135] Optionally, the processor 301 determines whether the campus bullying incident exists in the monitoring data to be detected based on the campus bullying incident score, including:
[0136] If the school bullying incident score is within the first score range, determining that a school bullying incident of the first emergency type exists in the monitoring data to be detected;
[0137] If the school bullying incident score is within a second score range, it is determined that a school bullying incident of a second emergency type exists in the monitoring data to be detected, and the maximum score boundary of the second score range is less than the minimum score boundary of the first classification range;
[0138] If the school bullying incident score is within a third score range, it is determined that a school bullying incident of a third emergency type exists in the monitoring data to be detected, and the maximum score boundary of the third score range is less than the minimum score boundary of the second classification range;
[0139] If the school bullying incident score is less than the minimum score boundary of the third score range, it is determined that the school bullying incident does not exist in the monitoring data to be detected.
[0140] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the various processes of the campus bullying incident detection method provided by the embodiment of the present invention are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0141] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, the storage medium can be a disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc.
[0142] The above disclosure is only the preferred embodiment of the present invention, which certainly cannot be used to limit the scope of the present invention. Therefore, equivalent changes made according to the claims of the present invention are still within the scope of the present invention.
Claims
1. A method for detecting campus bullying incidents, characterized in that: The method comprises the following steps: Acquire monitoring data to be detected on campus, wherein the monitoring data to be detected includes multiple target persons; Respectively extracting a first video feature and a first audio feature of the monitoring data to be detected; Based on the first video feature and the first audio feature, determining a plurality of pairs of target features related to the school bullying incident and a weight combination corresponding to each pair of the target features in a pre-constructed feature sharing subspace, each pair of the target features comprising a video target feature and an audio target feature aligned in time, and the weight combination comprising a first weight corresponding to the video target feature and a second weight corresponding to the audio target feature; Performing weighted sum processing on the target features based on the weight combination to obtain fused features, each fused feature corresponding to a pair of the target features; Classify and process the plurality of fusion features to determine whether the campus bullying incident exists in the monitoring data to be detected.
2. The campus bullying incident detection method according to claim 1, characterized in that: The method of determining, based on the first video feature and the first audio feature, multiple pairs of target features related to the campus bullying incident and weight combinations corresponding to each pair of the target features in a pre-constructed feature sharing subspace includes: Construct feature sharing subspace; Based on linear discriminant analysis, mapping the first video feature and the first audio feature to the feature shared subspace to obtain a second video feature and a second audio feature, wherein the second video feature corresponds to the first video feature, and the second audio feature corresponds to the first audio feature; Based on the second video feature and the second audio feature, multiple pairs of target features related to the school bullying incident and weight combinations corresponding to each pair of the target features are determined.
3. The campus bullying incident detection method according to claim 2, characterized in that: The determining, based on the second video feature and the second audio feature, multiple pairs of target features related to the campus bullying incident and weight combinations corresponding to each pair of the target features includes: According to the feature operator corresponding to the campus bullying incident, the second video feature and the second audio feature are respectively calculated and processed to obtain a plurality of video target features related to the campus bullying incident and a plurality of audio target features related to the campus bullying incident; Determine one of the video target features and one of the audio target features aligned in the time dimension as a pair of target features related to the campus bullying incident; For each pair of the target features, a weight combination corresponding to the target features is determined based on the video target feature and the audio target feature.
4. The campus bullying incident detection method according to claim 1, characterized in that: For each pair of the target features, determining a weight combination corresponding to the target features based on the video target feature and the audio target feature includes: For each pair of the target features, the video target feature and the audio target feature are respectively decomposed to obtain a first feature vector corresponding to the video target feature and a second feature vector corresponding to the audio target feature; The first feature vector and the second feature vector are respectively normalized to obtain a first weight corresponding to the video target feature and a second weight corresponding to the audio target feature.
5. The campus bullying incident detection method according to claim 1, characterized in that: The step of performing weighted summation processing on the target features based on the weight combination to obtain fused features includes: For each pair of the target features, weighting the video target features by a first weight, and weighting the audio target features by a second weight; The weighted video target feature and the audio target feature are added together to obtain a fusion feature.
6. The campus bullying incident detection method according to any one of claims 1 to 5, characterized in that: The classifying and processing the plurality of fusion features to determine whether the campus bullying incident exists in the monitoring data to be detected includes: Based on the behavioral characteristics related to the campus bullying incident, a plurality of the fusion features are classified and processed to obtain a campus bullying incident score; Based on the school bullying incident score, determine whether the school bullying incident exists in the monitoring data to be detected.
7. The campus bullying incident detection method according to claim 6, characterized in that: The determining, based on the school bullying incident score, whether the school bullying incident exists in the monitoring data to be detected includes: If the school bullying incident score is within the first score range, determining that a school bullying incident of the first emergency type exists in the monitoring data to be detected; If the school bullying incident score is within a second score range, it is determined that a school bullying incident of a second emergency type exists in the monitoring data to be detected, and the maximum score boundary of the second score range is less than the minimum score boundary of the first classification range; If the school bullying incident score is within a third score range, it is determined that a school bullying incident of a third emergency type exists in the monitoring data to be detected, and the maximum score boundary of the third score range is less than the minimum score boundary of the second classification range; If the school bullying incident score is less than the minimum score boundary of the third score range, it is determined that the school bullying incident does not exist in the monitoring data to be detected.
8. A campus bullying incident detection device, characterized in that: The campus bullying incident detection device comprises: An acquisition module is used to acquire monitoring data to be detected on campus, wherein the monitoring data to be detected includes multiple target persons; A first processing module, used for respectively extracting a first video feature and a first audio feature of the monitoring data to be detected; A second processing module is used to determine, based on the first video feature and the first audio feature, a plurality of pairs of target features related to the school bullying incident and a weight combination corresponding to each pair of the target features in a pre-constructed feature sharing subspace, each pair of the target features includes a video target feature and an audio target feature that are aligned in time, and the weight combination includes a first weight corresponding to the video target feature and a second weight corresponding to the audio target feature; A third processing module is used to perform weighted sum processing on the target features based on the weight combination to obtain fused features, each fused feature corresponding to a pair of the target features; The fourth processing module is used to classify the multiple fusion features to determine whether the school bullying incident exists in the monitoring data to be detected.
9. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps in the campus bullying incident detection method as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the campus bullying incident detection method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Campus intelligent monitoring method for preventing campus bullying
CN121121656A
A campus intelligent monitoring method for campus anti-bullying
CN121121656B