A campus anti-bullying system supporting AI audio and video processing by POE or POF transmission
AI audio and video processing systems supported by POE or POF transmission utilize skeletal topological geometric features to analyze interactive behavior, solving the problem of bullying recognition in complex multi-person interaction scenarios in campus security systems, and achieving high-precision, low-latency bullying behavior recognition and early warning.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ANKANG YOUJIAO ENLIGHTENMENT TECHNOLOGY CO LTD
- Filing Date
- 2026-01-14
- Publication Date
- 2026-06-02
AI Technical Summary
Existing campus security systems struggle to accurately distinguish between playful and bullying behavior when dealing with complex multi-person interaction scenarios. Traditional algorithms lack in-depth analysis of the relative spatial relationship and posture evolution logic between the interacting parties, leading to false alarms and missed alarms. Furthermore, edge computing devices are limited by power and cannot meet high computing power requirements.
The AI audio and video processing system, which supports POE or POF transmission, collects audio-visual data streams in real time through a multimodal acquisition unit. The AI edge computing unit analyzes interactive behavior using skeletal topological geometric features, constructs interactive binary pairs, and calculates active invasion and defensive collapse feature vectors. Combined with a logic gating module, it outputs suspected bullying trigger signals, optimizes computing power allocation, and performs privacy desensitization processing.
It enables precise deconstruction of multi-person interactive spatial relationships on edge computing devices, eliminates semantic ambiguity between playfulness and bullying, improves the accuracy of bullying behavior recognition, reduces false alarm rate, and ensures real-time response capability in complex scenarios.
Smart Images

Figure CN122135516A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a campus anti-bullying system that supports AI audio and video processing via POE or POF transmission, belonging to the field of image recognition technology. Background Technology
[0002] The current campus security system uses video surveillance combined with computer vision algorithms. It uses convolutional neural networks to detect human targets in video streams, extract limb texture features or movement amplitude to identify punching, kicking or pushing action tags, and has the ability to monitor simple scenarios such as falls or strenuous single-person movements based on single-action classification technology.
[0003] In high-frequency interaction scenarios during school breaks, playful roughhousing and physical bullying exhibit high similarity in single-frame image pixel features and limb movement trajectories. Traditional algorithms rely on limb acceleration or contact determination, lacking in-depth analysis of the relative spatial relationship and posture evolution logic between the interacting parties. Introducing multimodal auxiliary means for logical discrimination still has shortcomings. For example, Chinese invention patent CN119229593A discloses a school anti-bullying system based on speech recognition and big data analysis, which covers monitoring blind spots by matching voiceprints and keywords. This type of technology belongs to surface signal feature statistical matching. Playful roughhousing in the noisy environment of school breaks is often accompanied by high-decibel shouting, which can easily lead to false alarms; and audio analysis or shallow visual features cannot be analyzed. The relative spatial relationships and posture evolution during physical contact, especially in cases of silent malicious suppression or bullying that causes the victim's posture to collapse due to small movements, lack the ability to deconstruct the asymmetric geometry of interaction. It is difficult to distinguish between two-way active interaction and one-way suppression and aggression from the semantic level of images. Introducing a deep learning model based on long-term analysis to resolve semantic ambiguity faces hardware resource constraints. Campus security front-end terminals use Ethernet power supply or plastic fiber optic transmission, and the power envelope can only support low-power edge computing chips, which cannot support high-computing-power 3D convolutional neural network inference tasks. Increasing cloud computing power increases construction costs, and the limitation of data backhaul bandwidth leads to delays in police response.
[0004] Therefore, the technical problem to be solved by this invention is to accurately deconstruct the spatial relationships of multi-person interaction in edge computing-constrained environments based on lightweight geometric structural features and eliminate the semantic ambiguity of playfulness and bullying. Summary of the Invention
[0005] To address the problems mentioned in the background art, the technical solution of the present invention is as follows: A campus anti-bullying system that supports AI audio and video processing via POE or POF transmission, comprising:
[0006] The composite transmission link unit uses Power over Ethernet (POE) or Plastic Optical Fiber (POF) physical media to construct a physical transmission channel that carries uncompressed RAW domain video streams and power supply current.
[0007] A multimodal acquisition unit, coupled to a composite transmission link unit, is used to acquire audiovisual data streams within the monitoring area in real time.
[0008] The AI edge computing processing unit integrates a digital signal processing core to perform interactive behavior analysis on audiovisual data streams based on skeletal topological geometric features. Specifically, the AI edge computing processing unit is used to: extract the coordinates of key skeletal points of multiple human targets in the image using a pose estimation algorithm, and construct an interaction tuple based on adjacent coordinate sets at the same timestamp; for the first and second targets in the interaction tuple, calculate the rate of change of Euclidean distance between the centroid of the first target and the centroid of the second target within a preset time sliding window, and only when the rate of change of Euclidean distance is less than zero and the angle between the torso axis and the vertical axis of the first target is at a preset upright angle... Within the specified range, an active invasion feature vector is generated; the minimum circumscribed convex hull area of all skeletal key points of the second target is calculated simultaneously, and a defensive collapse feature vector is generated when the ratio of the convex hull area of the current frame to the convex hull area of the standard standing state decreases by more than a preset deformation threshold and the vertical height of the centroid decreases by more than a preset height threshold; the active invasion feature vector and the defensive collapse feature vector are input into the logic gating module, and when the interaction tuples in the same time sliding window simultaneously satisfy the geometric constraints of character polarity differentiation and interaction duration exceeding a preset time threshold, a trigger signal representing suspected bullying is output.
[0009] Preferably, the AI edge computing processing unit also includes an audio-visual linkage trigger module, which is used to monitor the ambient sound pressure level data acquired by the multimodal acquisition unit in real time; when the ambient sound pressure level data exceeds a preset decibel threshold or matches a preset high-frequency screaming soundprint feature, an interrupt command is generated; in response to the interrupt command, the image sampling frame rate for the human target area contained in the audio-visual data stream is increased from a first frequency to a second frequency greater than the first frequency, and the rendering weight of the background area is reduced, and computing power resources are preferentially allocated for the extraction of skeletal key point coordinates under the limited power envelope.
[0010] Preferably, when generating the active invasion feature vector, the AI edge computing processing unit is specifically used to obtain the centroid coordinates of the first target and the second target in multiple consecutive frames of images within a preset time sliding window; construct a relative displacement vector and differentiate it with respect to time to obtain the relative approximation velocity between the two; and determine that the spatial invasion condition is met and generate the active invasion feature vector only when the relative approximation velocity is negative and its absolute value is greater than a preset velocity threshold, and at the same time the line connecting the shoulder key points of the first target is detected to be parallel to the horizontal axis.
[0011] Preferably, when generating the defensive collapse feature vector, the AI edge computing processing unit adopts the following posture compactness determination logic: based on the set of key points of the second target, including the top of the head, neck, shoulder, elbow, wrist, hip, knee and ankle, the minimum convex polygon covering the set of key points is constructed using the convex hull algorithm; the posture compactness coefficient representing the degree of body curling is calculated according to the following formula. : ,in, The current pixel area of the smallest convex polygon. The area of the reference convex hull for the second target in a standard upright state; in the calculated... When the value is less than the preset collapse threshold, the second target is confirmed to be in a defensive posture.
[0012] Preferably, the logic gating module is also used to perform asymmetric mutual exclusion verification: when it is detected that the first target and the second target in the interaction tuple both generate active invasion feature vectors or defensive collapse feature vectors within the same time sliding window, it is determined that the interaction behavior has symmetry; in response to the symmetry determination, the output of the trigger signal representing the suspected bullying is shielded by the mutual exclusion logic, and the current interaction behavior is marked as bidirectional play or bidirectional static state.
[0013] Preferably, the plastic optical fiber (POF) physical medium in the composite transmission link unit is specifically used to: provide a first wavelength optical signal channel for transmitting uncompressed RAW domain video streams to preserve the skeletal edge features of the original image; provide a second wavelength optical signal channel or a parallel copper cable channel for transmitting control commands and power supply current; and the AI edge computing processing unit directly reads the uncompressed RAW domain video stream for processing, avoiding the impact of motion blur introduced by video compression coding on the accuracy of skeletal key point extraction.
[0014] Preferably, the system also includes a 3D visualization mapping unit for storing building information model data of the monitored area; in response to a trigger signal indicating suspected bullying, extracting the 3D spatial position of the interaction binary in the camera coordinate system; mapping the 3D spatial position to the building information model data, and overlaying a dynamic red warning sign with height information on the 2D planar map of the display terminal.
[0015] Preferably, the system also includes a tiered early warning management unit, which receives a trigger signal indicating suspected bullying and starts a first-level response timer; if no manual confirmation instruction is detected within the preset first-level response time, a notification signal containing a screenshot of the current scene is automatically sent to the preset security personnel's mobile terminal; if no handling feedback signal is detected within the preset second-level response time, the alarm information and real-time video stream are routed to the superior regional management platform.
[0016] Preferably, the AI edge computing processing unit is also used to perform privacy desensitization processing: before outputting the trigger signal representing suspected bullying, facial feature data is retained only for the first target that is determined to generate an active invasion feature vector; Gaussian blur or pixel block occlusion processing is performed on the facial regions of background personnel targets that did not participate in the interaction and the second target that generated a defensive collapse feature vector.
[0017] Preferably, the AI edge computing processing unit performs interactive behavior analysis locally only, and after generating a trigger signal that represents suspected bullying, it uploads relevant video clips and structured feature data to the cloud server through the composite transmission link unit, so as to reduce the network bandwidth occupation of routine monitoring.
[0018] Compared with the prior art, the beneficial effects of the present invention are:
[0019] 1. In AI audio and video processing supported by POE or POF transmission, a skeleton topology interaction asymmetry judgment mechanism is constructed. It does not rely on single action label recognition, but calculates the active party's spatial invasion vector and the passive party's posture collapse index in real time. By quantifying the spatiotemporal coupling relationship between the active party's centroid approach rate and the passive party's limb structure curling degree, the unique unidirectional suppression and structural collapse characteristics of bullying behavior are accurately defined. Using relative motion and posture geometric logic gating judgment, the symmetrical playful behavior with mutual offense and defense is filtered. It does not rely on the computing power of large cloud models, but eliminates semantic ambiguity in complex multi-person interaction scenarios from the underlying geometric logic, and improves the detection accuracy of alarm systems for real bullying events.
[0020] 2. Utilizing audio signal waveform features as a trigger for allocating computing power in the image processing workflow, this addresses the challenge of real-time processing of complex behavior analysis at high frame rates under the limited power envelope of edge computing devices. In standby or low-risk states, low-power polling is maintained. When the audio terminal detects abnormal sound pressure levels or screams in specific frequency bands, the video processor is triggered to enter a high-frequency sampling mode, prioritizing computing power allocation for skeletal keypoint extraction. A cross-modal audio-visual wake-up frequency conversion sampling mechanism ensures the system does not continuously occupy high bandwidth and computing power resources, capturing millisecond-level instantaneous limb mutation features. This ensures real-time capture and analysis of sudden violent behavior using distributed terminals powered by Ethernet or fiber optics.
[0021] 3. Extract and solve the topological relationships of key points of the human skeleton, and transform the behavior recognition basis from RGB texture features that are easily affected by the environment to stable geometric structure features. By analyzing the changes in the angle of human joint connection and the projection distance, the interference of image pixel values in typical complex lighting environments such as backlit corridors, dim stairwells, or dappled shade on campuses is avoided. The path is analyzed based on structural features rather than color features to ensure that the system can stably maintain human posture locking and solving when the lighting conditions fluctuate drastically or the background is cluttered, and to prevent false alarms and missed alarms due to shadow occlusion or insufficient exposure. Attached Figure Description
[0022] Figure 1 This is a system architecture and data processing flowchart under the POE / POF transmission support of the present invention;
[0023] Figure 2 This is a statistical distribution diagram of the posture compactness coefficient under normal interaction and bullying behavior in this invention;
[0024] Figure 3 This is a flowchart illustrating the audio-visual linkage operation logic and hierarchical early warning interaction of the system of the present invention. Detailed Implementation
[0025] The description of this specific embodiment is intended to provide a detailed explanation of the technical solution of the present invention so that those skilled in the art can understand it. However, the following embodiments are only used to explain the present invention and do not constitute a limitation on the scope of protection of the present invention.
[0026] A campus anti-bullying system supporting AI audio and video processing via PoE or PoF transmission is architecturally composed of a composite transmission link unit, a multimodal acquisition unit, an AI edge computing processing unit, and a back-end hierarchical early warning management unit. The system operates an interactive behavior analysis method based on skeletal topological geometric features. It utilizes a physical transmission channel constructed using PoE Ethernet or PoF plastic optical fiber to connect front-end sensors and edge computing nodes. By extracting key points of the human skeleton from the unstructured pixel stream, it calculates feature vectors representing interactive intentions and outputs alarm signals via a logic gating module, achieving real-time identification and early warning of campus bullying behavior. The composite transmission link unit constructs a physical transmission channel carrying uncompressed RAW domain video streams and power supply current. This channel uses a PoE power supply protocol conforming to the IEEE 802.3bt standard or POF plastic optical fiber with a core diameter of 1.0mm as the transmission medium, providing a speed of at least 1Gbps. The dedicated bandwidth transmits raw image data with a resolution of 1920×1080 and a color depth of 12 bits. This link ensures a signal attenuation rate of less than 0.5dB / 100m, enabling the backend AI edge computing processing unit to directly read the uncompressed RAW domain video stream, avoiding the impact of motion blur introduced by video compression encoding on the accuracy of skeletal keypoint extraction. A multimodal acquisition unit is coupled to the composite transmission link unit to acquire audiovisual data streams in real time within the monitoring area. This unit includes a high-sensitivity microphone and a high-frame-rate image sensor. The microphone is used to capture ambient sound signals, and the image sensor is used to acquire the aforementioned uncompressed RAW domain video stream. The AI edge computing processing unit integrates a digital signal processing core to perform interactive behavior analysis based on skeletal topological geometric features on the audiovisual data stream. This processing unit uses a pose estimation algorithm to process the input RAW domain video frames and detect the coordinates of skeletal keypoints of multiple human targets in the image. The coordinate set includes key parts such as the tip of the nose, neck, shoulder, elbow, wrist, hip, knee and ankle. The system constructs an interactive binary based on adjacent coordinate sets under the same timestamp, and calculates the active invasion feature vector and the defensive collapse feature vector for the first target and the second target in the binary respectively.
[0027] For the first objective, the processor has a length of Within a time sliding window, the Euclidean distance change rate as the centroid of the first target approaches the centroid of the second target is calculated. The system obtains the centroid coordinates of the two targets in multiple consecutive frames within this time window, constructs a relative displacement vector, and calculates its derivative with respect to time to obtain the relative approximation velocity. When the relative approximation velocity is negative and its absolute value is greater than a preset velocity threshold, the system calculates the relative approximation velocity. At the same time, the angle between the torso axis of the first target, i.e., the line connecting the key point of the neck to the midpoint of the hip, and the vertical axis. While maintaining a preset upright angle range, the system generates an active invasion feature vector. For the second target, the processor calculates the minimum circumscribed convex hull area of all skeletal keypoints. The system uses a convex hull algorithm to construct the minimum convex polygon covering the set of keypoints of the second target and calculates the current pixel area of this polygon. Simultaneously, the system calls the pre-stored reference convex hull area of the target in its standard standing state. According to the formula Calculate the attitude compactness coefficient When detected Value in time window The inward descent exceeds the preset deformation threshold, and the vertical height of the second target's centroid... Relative to standard standing height The decline When the height exceeds a preset threshold, the system generates a defensive collapse feature vector. The logic gating module takes the aforementioned active invasion feature vector and defensive collapse feature vector as input and performs an asymmetric mutual exclusion check. This module checks whether the interactive binary pairs within the same sliding window simultaneously satisfy the following geometric constraints: First, there is a role polarity differentiation, i.e., one party generates an active invasion feature vector while the other generates a defensive collapse feature vector; second, the duration of this polarity interaction. If the above conditions are met simultaneously, the logic gating module outputs a trigger signal that indicates suspected bullying. If both targets generate active invasion feature vectors or defensive collapse feature vectors, the system uses mutual exclusion logic to shield the trigger signal output.
[0028] The AI edge computing processing unit also includes an audio-visual linkage trigger module, which monitors ambient sound pressure level data in real time. In non-alarm mode, video processing operates in a low sampling rate polling mode. When the audio module detects that the ambient sound pressure level exceeds a preset decibel threshold or matches a preset high-frequency screaming sound signature, the audio-visual linkage trigger module generates an interrupt command. The video processor responds to this command by increasing the image sampling frame rate for images containing human targets from a first frequency to a second frequency greater than the first frequency, and reducing the rendering weight of background areas. Under a limited power envelope, it prioritizes allocating computing resources for extracting skeletal keypoint coordinates. The system also includes a 3D visualization mapping unit that stores building information model data of the monitored area. In response to a suspected bullying trigger signal, this unit extracts the 3D spatial position of the interaction binary in the camera coordinate system, maps this position to the building information model data using a perspective transformation algorithm, and overlays it onto a 2D map on the display terminal, displaying the position with height information. The system displays dynamic red warning signs; the tiered early warning management unit receives trigger signals indicating suspected bullying and starts a first-level response timer. If no manual confirmation instruction is detected within the preset first-level response time, the unit automatically sends a notification signal containing a screenshot of the scene to the preset security personnel's mobile terminal; if no handling feedback signal is detected within the preset second-level response time, the unit routes the alarm information and real-time video stream to the upper-level regional management platform; the AI edge computing processing unit performs privacy desensitization processing. Before outputting the trigger signal, the system only retains the facial feature data of the first target identified as generating an active invasion feature vector; Gaussian blur or pixel block occlusion processing is performed on the facial areas of background personnel targets who did not participate in the interaction and the second target that generated a defensive collapse feature vector. This unit only performs interaction behavior analysis locally, and after generating the trigger signal, it uploads relevant video clips and structured feature data to the cloud server through the composite transmission link unit.
[0029] Example 1: In a crowded corridor during school breaks, multiple groups of students are engaged in frequent physical contact within the monitored area. A typical application-layer challenge in this scenario is that one group involves two students playfully roughhousing, while the other involves one student pushing and forcing another. Both groups exhibit highly similar limb amplitude and contact characteristics in single-frame images, easily leading to false alarms or missed detections based on traditional action recognition algorithms. When the system encounters this situation, the AI edge computing processing unit acquires the uncompressed RAW domain video stream from the front end via a PoE power supply protocol conforming to the IEEE 802.3bt standard or a POF transmission channel made of PMMA plastic fiber with a core diameter of 1.0mm. It then uses a pose estimation algorithm to extract the skeletal keypoint coordinates of all targets in the image. For the two sets of interaction objectives mentioned above, interaction tuples are constructed respectively, and the interaction asymmetry determination logic is executed in parallel. For the playful interaction tuples, the system slides within a set time window. Internal calculations revealed that although the absolute values of the relative approach velocities of both sides were greater than the preset velocity threshold... However, their roles frequently switch within a short period of time; that is, one appears to be approaching in the current frame but moves backward in subsequent frames, and the pose compactness coefficient of both is... All values remained above 0.8, and no sustained defensive collapse characteristics were observed. Based on this, the logic gating module determined that the interaction behavior was symmetrical and used mutual exclusion logic to shield the trigger signal output and filter false alarms.
[0030] For a binary pair of one-way bullying, the system detects the first target within the time window. The inner velocity continuously approaches the second target, and the relative approach velocity is always negative and its absolute value is greater than 1. The angle between the torso axis and the vertical axis While maintaining an upright angle, the system generates an active invasion feature vector and monitors the attitude compactness coefficient of the second target. Within the same time window, it drops sharply to below 0.6, and the vertical height of the centroid... The threshold is lowered, generating a defensive collapse feature vector. The logic gating module detects that the binary pair simultaneously satisfies the conditions of role polarity differentiation and the duration of the polarity state. When the three geometric constraints of exceeding the preset threshold and not having a role reversal occur, a trigger signal representing suspected bullying is output. This trigger signal then activates the audio-visual linkage trigger module and the 3D visualization mapping unit. If a high-frequency scream is present at this time, the audio module detects that the ambient sound pressure level exceeds the preset decibel threshold and sends an interrupt command, forcibly increasing the image sampling frame rate for that area to ensure that more motion details are captured. The 3D visualization mapping unit extracts the 3D spatial position of the interaction binary and maps it to the building information model data. A dynamic red warning sign with height information is overlaid on the 2D map of the security terminal. The hierarchical early warning management unit then initiates the response process and sends a notification signal containing a screenshot of the scene to the security personnel, realizing a closed-loop processing from abnormal behavior identification to accurate early warning. This embodiment verifies that the system solves the industry pain point of semantic ambiguity between playfulness and bullying in complex interaction scenarios through asymmetric analysis based on skeletal topology.
[0031] Example 2: In this example, a verification platform was built in a controlled simulated school building corridor environment to verify the effectiveness, anti-interference capability, and nonlinear response characteristics of the campus anti-bullying system of the present invention in handling complex interactive behaviors. The platform consists of four PoE-powered network cameras conforming to the IEEE 802.3bt standard, a set of high-sensitivity microphone arrays, and an edge computing processing unit equipped with a BM1688AI computing chip. To simulate the data transmission challenges in a real engineering environment, a 100-meter-long PMMA plastic optical fiber was used for the transmission link. Gaussian white noise with a signal-to-noise ratio of 20dB and power frequency electromagnetic interference at a frequency of 50Hz were artificially introduced into the transmission channel to test the stability of the system under non-ideal channel conditions. Three sets of comparative samples were designed to quantitatively evaluate the interactive behavior analysis method based on skeletal topological geometric features of the present invention. Control group A uses a single-action recognition algorithm that only includes acceleration features; control group B, as a partially missing control group, removes the defensive collapse feature vector verification in the logic gating module and relies only on the active invasion feature vector for judgment; the sample group of this invention uses a complete bidirectional feature vector collaborative judgment logic. In addition, the experiment introduces a problem intensity gradient, using professional actors to simulate different intensities of physical conflict and playful fighting scenarios, namely low, medium, and high. Among them, the high-intensity playfulness approximates real bullying behavior in terms of movement amplitude and speed, constituting a highly difficult interference sample; during the experiment, each group of systems synchronously receives the same uncompressed RAW domain audiovisual data stream. For each group of input data, the system extracts skeletal key points. In the sample group of this invention, the processing unit calculates in real time the relative approach velocity of the first target and the angle between the torso axis, as well as the posture compactness coefficient of the second target. Rate of change of vertical height from the center of mass The table below lists the key intermediate feature values of the present invention sample group and the control group B in a typical high-intensity play scenario (scenario number S-05). In this scenario, although the two targets moved violently, there was no one-way suppression. See Table 1.
[0032] Table 1: Comparison of Key Intermediate Features and Judgment Results
[0033]
[0034] Data shows that control group B, which relies solely on active invasion features, made false alarms when faced with high-intensity play, while the sample group of this invention introduced the posture compactness coefficient of a second target. As a mutual exclusion check condition, it was determined that the second target was in a non-defensive posture. =0.85>0.6), thus successfully suppressing false alarms. This confirms the synergistic effect between the two feature vectors of active invasion and defensive collapse: neither feature alone can effectively eliminate semantic ambiguity; only the logical AND operation of the two can achieve accurate judgment. Furthermore, to verify the parameters... The reasonableness of the (time sliding window) value, the setting of an out-of-range control group in the experiment, and the... The test results, set to 0.2 seconds, 0.5 seconds (preferred values of this invention), and 1.5 seconds respectively, show that when... At 0.2 seconds, the system became overly sensitive to instantaneous actions, causing the false alarm rate to surge to 18.5%; when At 1.5 seconds, although the false alarm rate decreased, the detection rate for short-term, sudden bullying behavior decreased by 22.3%, and the alarm delay increased. Only when... Around 0.5 seconds, the system achieves an optimal balance between detection rate (96.2%) and false alarm rate (1.8%), exhibiting significant peak performance characteristics. Tests on noise interference show that although simulated noise superimposed on the original RAW image data causes fluctuations in the confidence level of single-frame skeleton point extraction, the system's use of multi-frame time-series logic gating minimizes the detection error within the time window. The internal noise is smoothly filtered out, and the final output alarm signal still maintains stable triggering characteristics in a strong noise environment, proving the anti-interference ability of this solution in engineering practice. In summary, this embodiment, through rigorous comparative experiments and gradient stress tests, confirms that the present invention effectively solves the confusion between playful and bullying behaviors in complex scenarios by constructing an asymmetric judgment logic of interactive binary groups, and the selection of key parameters has clear engineering basis and performance advantages.
[0035] Example 3: This example combines Figures 1 to 3 This document describes a campus anti-bullying system that supports AI audio and video processing via PoE or PoF transmission. Figure 1As shown, the system architecture mainly consists of a multimodal acquisition unit, a composite transmission link unit, an AI edge computing processing unit, a logic gating module, and a backend response unit. The processing flow begins with the multimodal acquisition unit acquiring audiovisual data streams and ambient sound pressure levels. This data is transmitted via a composite transmission link unit using POE or POF physical media to carry uncompressed RAW domain video streams to the AI edge computing processing unit. The AI edge computing processing unit, on the one hand, uses a pose estimation algorithm to extract skeletal key points and construct interactive binary pairs; on the other hand, it combines the high-frequency screaming soundprints or abnormal sound pressure data monitored by the audio-visual linkage trigger module to respond to interruption triggers. The sampling frequency is increased, and then the active invasion feature vector, including the relative approach velocity and the angle between the trunk axis and the defensive collapse feature vector, including the ratio of the convex hull area and the decrease in the centroid height, is generated in parallel. The logic gating module receives the above feature vectors, performs asymmetric mutual exclusion verification, and outputs a trigger signal representing suspected bullying under the condition that there is polarity differentiation and the time sliding window eliminates the semantic ambiguity of playfulness and bullying. Finally, it drives the hierarchical early warning management unit to perform privacy desensitization processing and route alarm information to the terminal, and drives the three-dimensional visualization mapping unit to map the dynamic red warning sign with height information into the building information model.
[0036] like Figure 2 As shown, the horizontal axis represents the attitude compactness coefficient. The numerical range is shown on the vertical axis, which represents the frequency (%). The graph compares the distribution differences between normal interactive behavior and bullying behavior. The data shows that the victim's posture under bullying behavior is concentrated in... The values are in the lower range of 0.3-0.6, especially peaking in the 0.4-0.5 range, exhibiting characteristics of structural collapse, while normal interactive behavior is mainly distributed in the lower range of 0.3-0.6. The values are in the higher range of 0.6-0.8, and there is a statistically clear distribution boundary between the two; for example... Figure 3 As shown, the process begins with audio-visual linkage frequency conversion sampling and multimodal real-time monitoring. By extracting skeletal topological features for interactive behavior geometric analysis, the data flow enters the asymmetric mutual exclusion verification stage. This stage strictly follows the logical judgment condition of the existence of polarity differentiation and duration > threshold. When the triggering condition is met, the system triggers a graded early warning process. This process performs privacy desensitization processing and 3D visualization mapping in parallel, uploading the data to the cloud server. On the other hand, it sends alarm information to security personnel to wait for receipt or confirmation of the alarm. If there is a timeout and no processing, it is automatically reported to the superior regional management platform.
[0037] Example 4: This example provides an in-depth analysis and parameterized calibration of the generation logic of the defensive collapse feature vector in the campus anti-bullying system. The aim is to eliminate potential technical black boxes regarding posture compactness determination and threshold setting. In actual system operation, the defensive collapse determination for the second target is not based on a single pixel area change, but rather on a composite calculation model that integrates spatial geometric deformation rate and gravity axis displacement rate. For any second target to be detected, the system constructs a dynamically updated posture state space. This space uses the normalized set of skeletal key points in the target's standard standing state as a benchmark, defining a reference state vector. When the second target enters the interactive area, the system collects the current coordinates of the skeletal key points in real time and maps them into an observation state vector. To quantify the degree of attitude deformation, the system introduces an attitude compactness coefficient. The calculation procedure not only calculates the ratio of the circumscribed convex hull area, but also introduces a correction factor for limb curling degree to calculate the minimum convex polygon area covering all key points of the second target. and with the benchmark area A preliminary ratio was obtained by comparison. The system calculates the average Euclidean distance from key points at the ends of the limbs (wrist, ankle) to the center point of the trunk (midpoint of the line connecting the midpoint of the hip and the midpoint of the neck). and the corresponding average distance in the standard standing state. The distance shrinkage ratio was obtained by comparison. The final attitude compactness coefficient Defined as ,in and These are weighting coefficients; under the default configuration, , .
[0038] Regarding the setting of the judgment threshold, this embodiment provides an adaptive threshold calibration method based on statistical analysis. In the initial stage of system deployment, by collecting no less than 1000 sets of normal interactive behaviors, including posture data of walking, standing and talking, and light play, a normal behavior threshold is constructed. The probability distribution model is calculated by examining the lower quartiles of this distribution. The collapse threshold Set as ,in The interquartile range ensures that only abnormal curling behaviors deviating from the normal posture distribution are marked as collapse features, thereby statistically minimizing the false alarm rate. Furthermore, for determining the magnitude of the vertical drop in centroid height, the system employs a step detection algorithm based on time series analysis. The system maintenance length is... For example, a 30-frame centroid height history buffer, for the current frame's centroid height... The system calculates and the buffer before Average frame height The difference ratio ,when If the height exceeds a preset threshold such as 0.35, and this downward trend continues... The system only confirms that the height descent condition is met when the frame remains stable within 10 frames, i.e., there is no rebound. This timing constraint logic effectively filters out interference from instantaneous height changes caused by short-term actions such as jumping, squatting down to tie shoelaces.
[0039] Example 5: To ensure the stability and consistency of the adaptive threshold calibration and audio-visual linkage triggering logic in the campus anti-bullying system under different engineering deployment environments, this example describes a set of pre-deployment calibration procedures, which are executed before the system is officially put into operation. The aim is to eliminate parameter deviations caused by environmental differences. For the setting of the ambient sound pressure level threshold in the audio-visual linkage triggering logic, an acoustic baseline calibration step is performed. Calibration is conducted in a quiet state in the monitored area during non-teaching hours. The system continuously collects ambient background noise data for no less than 30 minutes and calculates the equivalent continuous sound level during this period. With background noise peak The system automatically sets the sound pressure level trigger threshold to 1. ,in A preset safety margin, such as 15dB, is set to prevent false triggering caused by normal background noise. At the same time, based on the screaming voiceprint characteristics of a specific frequency band, the system plays a set of standardized simulated screaming test audio to verify the frequency domain response characteristics of the audio module and ensure that the signal-to-noise ratio in the target frequency range (1000Hz-3000Hz) is not lower than the preset standard.
[0040] Secondly, regarding the attitude compactness coefficient The system determines the threshold and performs visual geometry calibration. Considering that differences in camera installation height and pitch angle affect the geometric proportions of skeletal projections in 2D images, the system needs to perform scene-adaptive learning. In the initial deployment phase, the system guides testers to perform standard standing, walking, and simulated squatting movements within the monitored area. The processor extracts the skeletal key points of the testers in real time, constructs a reference posture state space for that scene, and records the testers' postures in different positions and positions. and Based on the collected sample set, the system recalculates the data using the aforementioned statistical methods. The probability distribution of values is used to update the collapse threshold. This ensures that the decision logic is adapted to the current camera viewpoint and scene geometry.
[0041] Example 6: This example aims to provide an offline calibration and parameter filling procedure for the actual deployment of a campus anti-bullying system, in order to eliminate the geometric mapping black box in the calculation of spatial encroachment vector and attitude collapse index. This procedure is executed during the system installation and debugging phase. Scene-specific projection parameters are obtained through standardized physical calibration experiments to ensure the calculation accuracy of the algorithm model under different camera installation heights and pitch angles. Perspective projection correction calibration is performed by placing standard calibration poles of known height (e.g., 1.70 meters) at the center of the ground and four edge corners of the monitoring area. The system collects the pixel height of the calibration poles in the two-dimensional image. And combined with actual physical height The pixel-to-physical scale mapping factor for this viewpoint is obtained by using a perspective projection model. The mapping factor It is embedded in the non-volatile memory of the edge computing unit as a reference scaling factor for the relative displacement and approximation velocity in the subsequent computation space encroachment vector, thereby converting the image pixel distance into the real physical space distance.
[0042] Secondly, the system executes the posture reference library construction process, guiding testers to perform standard standing, walking, and simulated squatting movements within the monitored area. The processor extracts the testers' skeletal key points in real time, constructing a standard posture reference space for this scenario. Based on this reference space, the system automatically calculates and updates the posture compactness coefficient. The base area relative to reference height This process generates a set of attitude reference parameters adapted to the current viewpoint. This procedure ensures that when facing individuals of different heights and monitoring images from different viewpoints, the system can accurately determine the posture collapse characteristics using normalized geometric standards, eliminating the risk of misjudgment caused by viewpoint distortion.
[0043] Example 7: The AI edge computing processing unit allocates an independent fixed-length circular buffer for each group of tracked interactive pairs in the on-chip memory. It uses a first-in-first-out strategy to store data containing timestamps, skeletal key point coordinates, and local image feature sequences. The processor uses the current frame timestamp as a reference index to backtrack to historical frames that meet the time threshold T. It fills the frame loss gaps caused by transmission jitter through a linear interpolation algorithm and constructs a time-aligned skeletal key point trajectory matrix as input data for calculating the relative approximation velocity and the rate of change of Euclidean distance.
[0044] After the system is powered on, it automatically executes the gravity axis calibration program, reads the three-axis accelerometer in the multimodal acquisition unit to calculate the projection component of the gravity acceleration vector on the image plane, and defines the projection component as the vertical reference axis for judging the tilt angle of the first target's torso axis. If the target is detected to be in a single-person walking state for a duration exceeding the preset stable period, the arithmetic mean of the convex hull area of all frames within the time period is calculated and fixed as the target's standard reference convex hull area Sstd. The parameter persists with the target ID and is released after the target leaves the monitoring area. The logic gating module adopts a majority voting mechanism based on a time-series cumulative sliding window. Within the time sliding window, geometric constraint verification is performed frame by frame on all sampled frames. When the proportion of frames that meet the character polarity differentiation condition in the total number of frames in the window exceeds the preset confidence threshold, such as 0.8, the counter performs an accumulation operation; otherwise, a linear decay operation is performed. When the cumulative value of the counter exceeds the preset alarm trigger level and the audio-visual linkage module does not output a high-frequency scream interruption signal, the alarm flag is set and the current audio-visual data stream segment is locked.
[0045] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0046] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A campus anti-bullying system that supports AI audio and video processing via PoE or PoF transmission, characterized in that, include: The composite transmission link unit uses Power over Ethernet (POE) or Plastic Optical Fiber (POF) physical media to construct a physical transmission channel that carries uncompressed RAW domain video streams and power supply current. A multimodal acquisition unit, coupled to a composite transmission link unit, is used to acquire audiovisual data streams within the monitoring area in real time. The AI edge computing processing unit integrates a digital signal processing core to perform interactive behavior analysis based on skeletal topological geometric features on audiovisual data streams. Specifically, the AI edge computing processing unit is used to: extract the coordinates of key skeletal points of multiple human targets in the image using a pose estimation algorithm, and construct an interaction tuple based on the set of adjacent coordinates under the same timestamp; for the first target and the second target in the interaction tuple, calculate the rate of change of Euclidean distance between the centroid of the first target and the centroid of the second target within a preset time sliding window, and generate an active invasion feature vector only when the rate of change of Euclidean distance is less than zero and the angle between the axis of the first target's torso and the vertical axis is within a preset upright angle range; Simultaneously calculate the minimum circumscribed convex hull area of all skeletal key points of the second target, and generate a defensive collapse feature vector when the ratio of the convex hull area of the current frame to the convex hull area of the standard standing state decreases by more than a preset deformation threshold and the vertical height of the centroid decreases by more than a preset height threshold. Input the active invasion feature vector and the defensive collapse feature vector into the logic gating module, and output a trigger signal representing suspected bullying when the interaction tuples in the same time sliding window simultaneously satisfy the geometric constraints of character polarity differentiation and interaction duration exceeding a preset time threshold.
2. A campus anti-bullying system supporting AI audio and video processing via POE or POF transmission as described in claim 1, characterized in that, The AI edge computing processing unit also includes an audio-visual linkage trigger module, which is used to monitor the ambient sound pressure level data acquired by the multimodal acquisition unit in real time. When the ambient sound pressure level data exceeds a preset decibel threshold or matches a preset high-frequency screaming soundprint feature, an interrupt command is generated. In response to the interrupt command, the image sampling frame rate for the audio-visual data stream containing the human target area is increased from a first frequency to a second frequency greater than the first frequency, and the rendering weight of the background area is reduced. Under the limited power envelope, computing resources are preferentially allocated for the extraction of skeletal key point coordinates.
3. A campus anti-bullying system supporting AI audio and video processing via POE or POF transmission as described in claim 1, characterized in that, When generating the active invasion feature vector, the AI edge computing processing unit specifically acquires the centroid coordinates of the first target and the second target in multiple consecutive frames within a preset time sliding window; constructs a relative displacement vector and differentiates it with respect to time to obtain the relative approximation velocity between the two; and determines that the spatial invasion condition is met and generates the active invasion feature vector only when the relative approximation velocity is negative and its absolute value is greater than a preset velocity threshold, and at the same time, it detects that the line connecting the shoulder key points of the first target is parallel to the horizontal axis.
4. A campus anti-bullying system supporting AI audio and video processing via POE or POF transmission as described in claim 1, characterized in that, When generating defensive collapse feature vectors, the AI edge computing processing unit employs the following posture compactness determination logic: Based on the set of key points of the second target, including the top of the head, neck, shoulders, elbows, wrists, hips, knees, and ankles, a minimum convex polygon covering this set of key points is constructed using the convex hull algorithm; the posture compactness coefficient, representing the degree of body curling, is calculated according to the following formula. : ,in, The current pixel area of the smallest convex polygon. The area of the reference convex hull for the second target in a standard upright state; in the calculated... When the value is less than the preset collapse threshold, the second target is confirmed to be in a defensive posture.
5. A campus anti-bullying system supporting AI audio and video processing via POE or POF transmission as described in claim 1, characterized in that, The logic gating module is also used to perform asymmetric mutual exclusion verification: when it is detected that the first target and the second target in the interaction tuple both generate active invasion feature vectors or defensive collapse feature vectors within the same time sliding window, the interaction behavior is determined to be symmetric; in response to the symmetry determination, the output of the trigger signal representing the suspected bullying is shielded by the mutual exclusion logic, and the current interaction behavior is marked as bidirectional play or bidirectional static state.
6. A campus anti-bullying system supporting AI audio and video processing via POE or POF transmission as described in claim 1, characterized in that, The plastic optical fiber (POF) physical medium in the composite transmission link unit is specifically used to: provide an optical signal channel of the first wavelength for transmitting uncompressed RAW domain video streams to preserve the skeletal edge features of the original image; A second wavelength optical signal channel or a parallel copper cable channel is provided for transmitting control commands and power supply current; the AI edge computing processing unit directly reads the uncompressed RAW domain video stream for processing.
7. A campus anti-bullying system supporting AI audio and video processing via PoE or PoF transmission as described in claim 1, characterized in that, The system also includes a 3D visualization mapping unit for storing building information model data of the monitored area; in response to a trigger signal indicating suspected bullying, it extracts the 3D spatial position of the interaction binary in the camera coordinate system. The three-dimensional spatial location is mapped into the building information model data, and a dynamic red warning sign with height information is overlaid on the two-dimensional planar map of the display terminal.
8. A campus anti-bullying system supporting AI audio and video processing via POE or POF transmission as described in claim 1, characterized in that, The system also includes a tiered early warning management unit, which receives trigger signals indicating suspected bullying and starts a first-level response timer. If no manual confirmation instruction is detected within the preset first-level response time, a notification signal containing a screenshot of the current scene is automatically sent to the preset security personnel's mobile terminal. If no handling feedback signal is detected within the preset second-level response time, the alarm information and real-time video stream are routed to the superior regional management platform.
9. A campus anti-bullying system supporting AI audio and video processing via POE or POF transmission as described in claim 1, characterized in that, The AI edge computing processing unit is also used to perform privacy desensitization processing: before outputting the trigger signal that represents suspected bullying, facial feature data is retained only for the first target that is determined to generate an active invasion feature vector; Gaussian blur or pixel block occlusion processing is performed on the facial regions of background personnel targets who did not participate in the interaction and the second target that generates a defensive collapse feature vector.
10. A campus anti-bullying system supporting AI audio and video processing via POE or POF transmission as described in claim 1, characterized in that, The AI edge computing processing unit performs interactive behavior analysis locally only, and after generating a trigger signal that indicates suspected bullying, it uploads relevant video clips and structured feature data to the cloud server through the composite transmission link unit.
Citation Information
Patent Citations
Campus bullying prevention system based on voice recognition and big data analysis
CN119229593A