Campus spoofing behavior identification method and device and electronic equipment

By segmenting and extracting multidimensional features from campus audio data, a temporal dialogue representation text is generated. Combined with a large language model for deep reasoning, this solves the problems of low accuracy and subjective dependence in the identification of campus bullying behavior in existing technologies, and achieves efficient and interpretable campus safety monitoring and early warning.

CN121708905APending Publication Date: 2026-03-20HANGZHOU DIANZI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610009176.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-06
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

Existing technologies are insufficient to effectively identify school bullying, especially information such as threats, sarcasm, and fear implied in speech rhythm. Furthermore, existing methods are susceptible to subjective judgment, resulting in low recognition accuracy and difficulty in providing operational guidance.

Method used

By acquiring raw audio data from the campus environment, segmenting it into single-sentence audio segments, extracting multi-dimensional features, and integrating them into a temporal dialogue representation text, deep reasoning analysis is performed using a large language model to generate campus bullying behavior identification results and intervention strategies.

Benefits of technology

It improves the accuracy of identifying school bullying behavior, generates highly interpretable analysis reports, provides behavioral nature analysis, psychological state assessment and specific intervention strategies, reduces reliance on human judgment, and improves monitoring and early warning efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121708905A_ABST
    Figure CN121708905A_ABST
Patent Text Reader

Abstract

The invention provides a campus spoofing behavior identification method and apparatus, and an electronic device. The method comprises the following steps: acquiring original audio data of a campus environment; segmenting the original audio data to obtain a plurality of single-sentence audio clips divided according to natural sentences and timestamp information corresponding to the single-sentence audio clips; performing multi-dimensional feature extraction on each single-sentence audio clip to obtain multi-dimensional feature data of each single-sentence audio clip; according to the timestamp information of each single sentence audio clip, integrating the multi-dimensional feature data into a time-sequenced dialogue representation text; and according to the dialogue representation text, performing campus spoofing behavior identification and strategy analysis to generate a campus spoofing behavior identification result and a corresponding behavior intervention strategy. According to the technical scheme, the audio dialogue is converted into the dialogue representation text fusing the emotion and the side language features, high-precision recognition and interpretable analysis of the campus spoofing behavior can be achieved based on deep reasoning of the dialogue representation text, and operation guidance is provided for educators.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of behavior detection technology, specifically to a method, device, and electronic equipment for identifying school bullying behavior. Background Technology

[0002] With increasing societal attention to students' mental health and safety, the use of artificial intelligence to identify and intervene in school bullying has become an important research direction. Traditional school bullying analysis mainly relies on manual questionnaires, teacher observations, or keyword matching based on text communication records. However, these methods are generally inefficient, susceptible to subjective judgment, and struggle to effectively capture the implicit tone, emotions, and complex linguistic interactions within bullying behavior. Summary of the Invention

[0003] In view of this, this application provides a method, device and electronic device for identifying school bullying behavior, so as to achieve accurate identification of school bullying behavior.

[0004] Specifically, this application is implemented through the following technical solution:

[0005] According to a first aspect of the embodiments of this specification, a method for identifying school bullying behavior is provided, comprising: acquiring raw audio data of a school environment; segmenting the raw audio data to obtain multiple single-sentence audio segments divided according to natural language and their corresponding timestamp information; extracting multi-dimensional features from each single-sentence audio segment to obtain multi-dimensional feature data of each single-sentence audio segment; integrating the multi-dimensional feature data into a time-series dialogue representation text based on the timestamp information of each single-sentence audio segment; and performing school bullying behavior identification and strategy analysis based on the dialogue representation text to generate school bullying behavior identification results and corresponding behavioral intervention strategies.

[0006] According to a second aspect of the embodiments of this specification, a school bullying behavior identification device is provided, comprising: a data acquisition unit for acquiring raw audio data of a school environment; an audio segmentation unit for segmenting the raw audio data to obtain multiple single-sentence audio segments divided according to natural language and their corresponding timestamp information; a feature extraction unit for performing multi-dimensional feature extraction on each single-sentence audio segment to obtain multi-dimensional feature data of each single-sentence audio segment; a data integration unit for integrating the multi-dimensional feature data into a temporally sequenced dialogue representation text based on the timestamp information of each single-sentence audio segment; and a reasoning learning unit for performing school bullying behavior identification and strategy analysis based on the dialogue representation text to generate school bullying behavior identification results and corresponding behavioral intervention strategies.

[0007] According to a third aspect of the embodiments of the present specification, an electronic device is provided, comprising a processor; and a computer readable storage medium having stored therein computer program instructions, which when executed by the processor, cause the processor to perform the method of the first aspect.

[0008] According to a fourth aspect of the embodiments of the present specification, a computer readable storage medium having stored thereon a computer program, which when executed by a processor, performs the method of the first aspect.

[0009] In the embodiments of the present application, by segmenting the original audio data into multiple single-sentence audio segments, multi-dimensional features are extracted for each single-sentence audio segment in parallel, and all single-sentence audio segments are integrated into a unified dialogue representation text based on timestamp information. The key bullying information such as threat, mockery, and fear hidden in the prosody of speech can be texted, so as to enhance the perception and recognition ability of implicit bullying behavior. Thus, the misjudgment and omission caused by pure text analysis or coarse-grained emotional tags can be overcome, and the recognition accuracy can be improved. Moreover, the dialogue representation text containing dialogue scene information is learned by reasoning in the embodiments. The reasoning process can simulate an expert-level logical analysis path, so that the recognition result has high interpretability. In addition to the recognition result of non-bullying, the output information can also include behavior property analysis, psychological state evaluation, risk prediction, and specific intervention strategies, etc., which has operational guidance significance. The embodiments solve the problem of strong practicality and limited application scenarios in the prior art. Furthermore, the natural sentence segmentation, multi-dimensional feature extraction, structured integration, and semantic reasoning in the embodiments can be connected in series as an automated process, which can reduce the dependence on manual judgment, thereby improving the efficiency of campus bullying monitoring and early warning, and providing technical support for building an intelligent and proactive campus security protection system. Therefore, the present application provides a campus bullying behavior recognition method, device, and electronic device to realize accurate recognition of campus bullying behavior. BRIEF DESCRIPTION OF DRAWINGS

[0010] In order to more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the drawings needed in the specific embodiments or prior art description will be briefly introduced below. Some specific embodiments of the present application will be described in detail below with reference to the drawings in an exemplary but non-limiting manner. The same reference signs in the drawings denote the same or similar components or parts. Those skilled in the art should understand that these drawings are not necessarily drawn to scale. In the drawings:

[0011] Figure 1 is a schematic diagram of the architecture of a campus bullying behavior recognition system according to an exemplary embodiment of the present application;

[0012] Figure 2is a flowchart of a campus bullying behavior recognition method according to an example embodiment of the present application;

[0013] Figure 3 is a data processing flowchart in a campus bullying behavior recognition process according to an example embodiment of the present application;

[0014] Figure 4 is a block diagram of an electronic device according to an example embodiment of the present application;

[0015] Figure 5 is a block diagram of a campus bullying behavior recognition device according to an example embodiment of the present application. DETAILED DESCRIPTION

[0016] The example embodiments will be described in detail herein with reference to the attached drawings. In the following description, the same numbers are used to designate the same elements, unless otherwise indicated. The embodiments described in the following example embodiments are not representative of all embodiments consistent with the present application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the present application, as detailed in the appended claims.

[0017] The terminology used in the present application is for the purpose of describing particular embodiments only and is not intended to be limiting of the present application. As used in the present application and the appended claims, the singular forms "a," "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0018] Based on the analysis of the recording data in the campus, it is generally divided into two technical directions of automatic speech recognition (ASR) and speech emotion recognition (SER). The speech recognition technology can convert the dialogue content into text, but the pure text analysis cannot preserve the acoustic information in the speaking process. Many verbal bullying behaviors are not fully reflected in the literal content, but are implied in the emotional prosody features such as tone, volume, and speed, for example, the low-pitched tone when threatening or the high-pitched tone when mocking. The emotion recognition system mostly only outputs coarse-grained emotion labels such as "angry" and "sad", which are difficult to convert into fine-grained and qualitative tone descriptions with operational guidance value, resulting in distortion of audio tone annotation and inability to provide sufficient and accurate context support for subsequent semantic analysis. Moreover, the recognition model of the emotion recognition system has weak interpretability, which limits its actual application scenarios.

[0019] In addition, some reasoning models have obvious limitations in data integration and reasoning learning. For example, the sentence-level semantics, speaker roles and fine sentiment / intonation features in multi-turn dialogues cannot be effectively time-aligned and structuredly fused. Therefore, even if these multi-dimensional data can be obtained, there is a lack of an efficient mechanism to convert these discrete multi-dimensional data into situational context inputs directly available for reasoning models.

[0020] In view of the problems of loose coupling between modules, low information fusion efficiency, difficulty in capturing subtle changes in emotions and intonations, and poor explainability of bullying behavior detection in related reasoning models in multi-modal data processing, an embodiment of the present application provides a campus bullying behavior recognition scheme based on deep situational reasoning. As shown in Figure 1 The recognition scheme of the embodiment of the present application decomposes and executes multi-modal processing tasks by sentence-level high-precision labeling and alignment by specialized models, thereby converting the original dialogue audio into structured, emotion and intonation marked script text data. Finally, relying on the deep reasoning capability of the large language model, an analysis report and a specific coping strategy with high explainability are automatically generated, realizing accurate recognition and effective intervention of campus bullying behavior.

[0021] Next, the embodiments of the present specification will be described in detail.

[0022] The embodiment of the present application provides a campus bullying behavior recognition method. Figure 2 As shown in FIG. 2, the campus bullying behavior recognition method 200 includes at least the following steps S210 and S250: Figure 2

[0023] Step S210, obtaining original audio data of a campus environment.

[0024] The campus bullying behavior recognition method 200 provided by the embodiment of the present application can be executed by a hardware or software entity with data processing capability, including but not limited to a server, a processor, a computer device, a chip, a system on chip (SoC), an application specific integrated circuit (ASIC), a programmable logic device (FPGA), a distributed computing node, a cloud computing instance, an edge computing device, or a functional module / system integrated with corresponding processing logic. For the sake of clarity, the following embodiments will be exemplarily described with the server as the execution subject.

[0025] ​In some embodiments, the original audio data is collected by a data collection terminal deployed in a campus environment. The data collection terminal includes but is not limited to a surveillance camera with recording function, a portable recording device, or a dedicated sound pickup device in a classroom, etc. These terminal devices can continuously or triggered by a specified event (such as detecting a certain decibel value) to collect original audio stream or audio file containing the content of the conversation.

[0026] The collected audio data can be transmitted to a server or storage system at the backend through a wired or wireless network. The server or storage system buffers or temporarily stores the received audio data to form an original audio data source that can be called by subsequent processes.

[0027] In some embodiments, the original audio data collected and uploaded by the data collection terminal meets the preset technical specifications to meet the input requirements of the subsequent processing module. Optionally, the file format of the original audio data is WAV, with a sampling rate of 16 kHz, single-channel recording, and the upper limit of the audio duration of a single continuous collection is 20 minutes. In this way, the frequency range of human voice can be completely preserved, which helps to control the data volume, improve the transmission and processing efficiency, and provide a consistent input basis for subsequent data processing.

[0028] Step S220, the original audio data is segmented to obtain a plurality of single sentence audio segments and their corresponding timestamp information according to natural language sentences.

[0029] In some embodiments, after the server obtains the original audio data, it first performs preprocessing to standardize the input, such as noise reduction, uniform sampling rate, etc., to improve the robustness and accuracy of subsequent analysis.

[0030] In some embodiments, to achieve fine-grained multi-modal analysis, the preprocessed audio data is segmented at the natural language sentence level to divide long audio containing multiple rounds of conversation into independent single sentence units with complete semantics according to natural language semantics and pause rules. The single sentence audio segment refers to an audio segment that is continuous in time and contains only one complete utterance spoken by a speaker. The semantics of each single sentence audio segment is relatively independent, for example, with a clear tone end point or a longer silent interval as the boundary. Each single sentence audio segment is unique in speaker identity and utterance content, thereby laying a foundation for subsequent accurate feature extraction for each single sentence, such as identifying the speaker, emotion category, paralanguage features, etc. It also facilitates strict alignment of feature data in each dimension on the time axis.

[0031] Step S230, multi-dimensional feature extraction is performed on each single sentence audio segment to obtain multi-dimensional feature data of each single sentence audio segment.

[0032] To capture the bullying information that may be implied in the dialogue from different dimensions, the embodiment performs parallel multi-dimensional feature extraction on each single-sentence audio segment obtained by segmentation. Verbal bullying in campus bullying behavior is usually covert and complex, and its aggressiveness may not only be reflected in the literal content, but also often be conveyed through the identity relationship of the speaker, the emotional tendency, and the non-verbal tone features (such as intonation, speech rate, and voice quality changes). Single-dimensional information is difficult to fully restore the real situation of the dialogue and the psychological state of the two parties, and is easy to cause missed judgment or misjudgment. Therefore, the embodiment extracts multi-dimensional features including text content, speaker identity, emotional category, and paralinguistic features, so as to construct a more comprehensive and fine-grained dialogue representation text.

[0033] In some embodiments, the multi-dimensional feature data includes text content, speaker identity, emotional category, and paralinguistic features. The emotional category, for example, includes various emotional categories such as anger, fear, happiness, neutrality, sadness, surprise, etc.; the paralinguistic features refer to the prosodic features, voice quality features, and ways of making sounds in the speech signal, which are used to convey the attitude, intention, etc. of the speaker, in addition to the text content, the prosodic features refer to speech rate, volume, tone fluctuation, etc., the voice quality features refer to the texture of voice tremor, hoarseness, sharpness, etc., whether the breath is unstable, whether there is wheezing, etc.

[0034] In step S240, the multi-dimensional feature data is integrated into a dialogue representation text according to the timestamp information of each single-sentence audio segment.

[0035] To construct an input for large language models to perform deep context reasoning, the embodiment aligns and fuses the dispersed multi-dimensional feature data corresponding to the single-sentence audio segments according to their timestamp information, to generate a structured dialogue representation text. Optionally, the dialogue representation text is a script text.

[0036] In some embodiments, the server time-sequentially sorts and aligns all feature data according to the start and end timestamps of each single-sentence audio segment, ensures that the original order and rhythm of the dialogue are preserved, formats the aligned multi-dimensional features according to a preset structured text template, and outputs them, thereby generating the dialogue representation text. The dialogue representation text is strictly arranged in chronological order, and the speaker identity, paralinguistic features, emotional state, and text content of each single-sentence audio segment are organically combined in a sequence. Its format is similar to a drama script, and the tone features and emotional state are embedded in the form of stage directions between the corresponding speaker identity and text content.

[0037] For example, a single line of data in the dialogue representation text can be represented as:

[0038] Speaker_01: (Deep voice, slow speech, angry) Do you think you're right to do this?

[0039] Speaker_02: (Voice trembling, volume fading, fear) I...I just wanted to finish my homework...

[0040] It is evident that the dialogue representation text not only preserves the literal content of the dialogue and the information of the speakers, but also transforms the acoustic emotions and tone information that are originally difficult to convey directly through text into high-context-density natural language description information that can be directly understood and processed by large language models. This provides input for achieving accurate semantic understanding, behavior analysis and strategy generation in subsequent steps.

[0041] Step S250: Based on the dialogue representation text, identify and analyze school bullying behavior to generate school bullying behavior identification results and corresponding behavioral intervention strategies.

[0042] In some embodiments, a large language model, such as the Qwen-Plus model, is used and Chain of Thought (CoT) cue word engineering is employed to guide the Qwen-Plus model to conduct in-depth reasoning analysis on the deep meaning, conflict level, and psychological state of the dialogue text, ultimately generating an analysis report that includes the identification results of school bullying behavior and corresponding behavioral intervention strategies.

[0043] like Figure 2 As shown in the illustrated method for identifying school bullying, this embodiment segments the original audio data into multiple single-sentence audio segments, extracts multi-dimensional features from each segment in parallel, and integrates all segments into a unified dialogue representation text based on timestamp information. This textualizes key bullying information such as threats, sarcasm, and fear, which are difficult to quantify and are implicit in the prosody of speech, thereby enhancing the perception and identification of covert bullying behavior. This overcomes the misjudgments and omissions caused by pure text analysis or coarse-grained sentiment labeling, improving recognition accuracy. Furthermore, this embodiment performs reasoning learning on the dialogue representation text containing dialogue context information. The reasoning process can simulate expert-level logical analysis paths, making the recognition results highly interpretable. The output information not only includes the bullying identification result but also includes behavioral nature analysis, psychological state assessment, risk prediction, and specific intervention strategies, providing operational guidance. This embodiment solves the problem of limited practicality and application scenarios in existing technologies. Furthermore, the natural language segmentation, multi-dimensional feature extraction, structured integration, and semantic reasoning in this embodiment can be linked into an automated process, which can reduce the reliance on manual judgment, thereby improving the efficiency of campus bullying monitoring and early warning, and providing technical support for building an intelligent and proactive campus safety protection system.

[0044] In some embodiments, the step S210 of segmenting the original audio data to obtain multiple single-sentence audio segments and corresponding timestamp information according to natural sentences includes: using a speech recognition and time alignment model to recognize the original audio data to obtain transcription text data; determining a sentence boundary based on a specified punctuation mark or a silence duration between adjacent words in the transcription text data; and cutting out corresponding single-sentence audio segments from the original audio data according to the sentence boundary and recording the start and end timestamps of each single-sentence audio segment.

[0045] The speech recognition and time alignment model of the present embodiment can be a pre-trained large language model or a professional model such as a WhisperX model. A large language model (LLM) is an artificial intelligence model based on deep learning, which can capture complex relationships and patterns in audio data through training on a large amount of audio data, and thus be used for speech recognition and time alignment tasks.

[0046] In an example, the transcription text data includes timestamps of punctuation marks and words, the specified punctuation marks include a period, a question mark, an exclamation mark, etc., and all word-level timestamps in the transcription text data are traversed. If the specified punctuation mark is detected or the silence duration between adjacent words is greater than a preset duration, for example, greater than 500 ms, the position of the specified punctuation mark and the silence duration greater than the preset duration is determined as a sentence boundary.

[0047] In an example, in the case that there is a blank segment in the transcription text data that is longer than a preset duration, the method further includes: marking the blank segment as an abnormal segment, skipping the multi-dimensional feature extraction step for the abnormal segment, and retaining the corresponding time sequence position of the abnormal segment when generating the dialogue representation text.

[0048] In an example, if the speech recognition and time alignment model outputs an empty text, the present round of campus bullying behavior recognition is ended, and other original audio data is reacquired.

[0049] In some embodiments, the multi-dimensional feature data includes text content, speaker identity, emotion category, and paralanguage features.

[0050] In some embodiments, the speaker identity of each single-sentence audio segment is extracted by the following steps: performing voiceprint recognition and speaker clustering on the original audio data to obtain the speaker identity of each speech segment in the original audio data, and matching the corresponding speaker identity of each single-sentence audio segment based on the position of each single-sentence audio segment in the original audio data.

[0051] The voiceprint recognition model is a pre-trained large language model. The original audio data and each single-sentence audio segment are input into the voiceprint recognition model. The voiceprint recognition model performs voiceprint feature recognition on the original audio data and performs clustering analysis on the voiceprint feature recognition result to obtain the speaker identity of each speech segment in the original audio data. The speaker identity corresponding to the single-sentence audio segment is determined based on the speech segment corresponding to the single-sentence audio segment input into the voiceprint recognition model.

[0052] In some embodiments, the text content of each single-sentence audio segment is extracted by the following steps: each single-sentence audio segment is sequentially input into a speech recognition model to determine the text content corresponding to each single-sentence audio segment. The speech recognition model is a pre-trained large language model. Based on the process of text recognition of the single-sentence audio segment by the speech recognition model, those skilled in the art can refer to the related technical solutions, and this embodiment will not be repeated here.

[0053] In some embodiments, the emotion category of each single-sentence audio segment is extracted by the following steps: each single-sentence audio segment is sequentially input into a speech emotion recognition model to determine the emotion category corresponding to each single-sentence audio segment. The speech emotion recognition model is a pre-trained large language model.

[0054] For example, a speech emotion recognition model using Hidden Unit BERT architecture is used. Each single-sentence audio segment is sequentially input into the speech emotion recognition model. The speech emotion recognition model learns self-supervised speech representation from the currently input single-sentence audio segment and classifies the self-supervised speech representation through a classification head to output a 6-dimensional Softmax probability vector. The 6-dimensional Softmax probability vector corresponds to six emotion categories: anger, fear, joy, neutral, sadness, and surprise. For example, when the highest category confidence is not less than a preset value of 0.7, the emotion category label is adopted, otherwise the currently input single-sentence audio segment is marked as "emotion uncertain". The classification head can be a classification head trained on a multi-emotion dataset (e.g., CASIA-Emotional Corpus).

[0055] In some embodiments, the paralanguage feature of each single-sentence audio segment is extracted by the following steps: each single-sentence audio segment is sequentially input into a paralanguage recognition model. The paralanguage recognition model infers the paralanguage feature of the currently input single-sentence audio segment based on a preset prompt word template to determine the paralanguage feature corresponding to the single-sentence audio segment.

[0056] The paralinguistic recognition model is a pre-trained large language model or a specialized model used to describe paralinguistic features of audio. The prompt word template is used to control the paralinguistic recognition model to output a structured description. For example, the prompt word template instructs the model to play a specific specialized role and constrains its output format to a stage prompt in parentheses. The stage prompt describes at least one or more acoustic features among speech rate, volume, pitch variation, and voice texture.

[0057] In one example, the output of the paralinguistic recognition model undergoes format validation. If it does not conform to the format specified by the prompt word template, a re-inference is triggered until paralinguistic features that meet the format requirements are obtained, or until the retry limit is reached and a preset placeholder is used. Format validation can be performed using regular expressions. When the model's output fails format validation, the model re-performs inference on the input. Optionally, the maximum number of retries is 3.

[0058] In some embodiments, step S240 above, which integrates the multi-dimensional feature data into a temporally sequenced dialogue representation text based on the timestamp information of each single-sentence audio segment, includes: aggregating the multi-dimensional feature data corresponding to each single-sentence audio segment into structured dialogue turn data according to its timestamp information; and combining all dialogue turn data into the dialogue representation text according to the timestamp information; wherein each line in the dialogue representation text corresponds to a single-sentence audio segment and integrates the multi-dimensional feature data corresponding to that single-sentence audio segment.

[0059] This embodiment can aggregate the multi-dimensional feature data corresponding to each single-sentence audio segment into a JSON object. Each JSON object corresponds to a single-sentence audio segment, including start timestamp, end timestamp, speaker ID, text content, emotion category, and paralinguistic features.

[0060] For example, the structure of the JSON object is as follows:

[0061] {

[0062] "Sentence_ID": "1",

[0063] "Speaker_ID": "SPEAKER_01",

[0064] "Start_Time": "00:00:01.500",

[0065] "End_Time": "00:00:03.200",

[0066] Context: "Hello."

[0067] "Emotion": "happy"

[0068] "Tone_Description": "A drawn-out tone with an upward inflection."

[0069] }

[0070] After obtaining the JSON objects of each individual audio segment, iterate through all JSON objects and output them in chronological order, with a uniform output format, for example:

[0071] [Speaker ID]: (Tone description, emotion category) Dialogue text content, which can also be represented as "Speaker_ID: Tone_Description, Emotion, Context".

[0072] In some embodiments, step S250 above, which involves identifying and analyzing bullying behavior based on the dialogue representation text to generate bullying behavior identification results and corresponding behavioral intervention strategies, includes: inputting the dialogue representation text into a pre-trained large language model; guiding the large language model to reason about the dialogue representation text based on a preset thought chain prompting engineering technique; and generating an analysis report based on the reasoning results, wherein the analysis includes bullying behavior identification results, psychological state analysis, risk assessment, and behavioral intervention strategies.

[0073] To illustrate the school bullying behavior identification process in this embodiment in detail, the following embodiments will be used as a reference.

[0074] by Figure 1 Taking the school bullying behavior recognition system shown as an example, the system includes a segmentation module, a labeling module, an integration module, and an analysis module. The segmentation module is used to execute step S1, the labeling module is used to execute step S2, the integration module is used to execute step S3, and the analysis module is used to execute step S4. The following section will combine... Figure 1 and Figure 3 Provide a detailed description of each module of this system.

[0075] Step S1: Preprocessing and audio segmentation.

[0076] like Figure 3 As shown, in this embodiment, the system receives raw audio data uploaded by a campus acquisition terminal, such as a classroom camera or a portable recording device. The data format is WAV, the sampling rate is 16kHz, it is mono, and the maximum duration is no more than 20 minutes.

[0077] After receiving the original audio data, the system first performs standardization preprocessing on the original audio data. For example, if the sampling rate is not 16 kHz, resampling is performed through a professional tool library, and the quantization precision is unified to 16 bits to meet the input specification requirements of subsequent speech recognition and time alignment models. If the signal-to-noise ratio of the original audio data estimated by the joint estimation of short-time energy and zero-crossing rate is lower than 10 dB, a noise reduction model is called to suppress the steady-state background noise in the original audio data to improve the accuracy of subsequent speech recognition. The noise reduction model is, for example, a WaveNet noise reduction model fine-tuned based on the Microsoft DNS Challenge v3 dataset. Of course, those skilled in the art can also use other noise reduction models, and the present embodiment does not particularly limit this.

[0078] After the original audio data is standardized and preprocessed, the preprocessed audio data is segmented using a speech recognition and time alignment model. For example, the speech recognition and time alignment model is a WhisperX model. The WhisperX model first performs automatic speech recognition on the preprocessed audio data and outputs punctuated transcription text data. The transcription text data and audio waveform are simultaneously input into a phoneme or word-level time model to calculate the absolute start time stamp Start_Time and end time stamp End_Time of each word in the audio.

[0079] It is worth noting that if the WhisperX model returns an empty text, the original audio data is marked as abnormal audio data, and the campus bullying behavior recognition is ended, and other original audio data is reacquired.

[0080] The system traverses all word-level time stamps for the transcription text data output by the WhisperX model and determines the sentence boundary based on the logic that the detected specified punctuation symbol position and the detected adjacent word position have a silence duration exceeding a preset threshold (e.g., 500 ms). According to the determined sentence-level time stamps T_sentence_start and T_sentence_end, an audio processing library is called to losslessly cut out the corresponding single-sentence audio segment from the original audio data, with a naming format such as sentence_{ID}.wav, and record its sequence ID and time stamp information. Optionally, the duration of the single-sentence audio segment is between 0.5s and 8s.

[0081] It is worth noting that if the silence between certain words lasts more than 5 seconds, the segment is marked as an abnormal segment, and the abnormal type is long-time silence. For this type of abnormal segment, the subsequent multi-dimensional feature extraction is skipped, but the timestamp is still retained in the data integration stage to maintain the dialogue timing integrity.

[0082] Step S2, multi-dimensional feature extraction.

[0083] The audio cut single sentence audio segments are sent into a parallel processing pipeline, and four independent but collaborative analysis tasks are performed simultaneously to extract features of four key dimensions such as speaker identity (Speaker_ID), text content (Content), emotion state (Emotion) and paralinguistic features (Tone_Description).

[0084] For the speaker identity, the system adopts an ECAPA-TDNN voiceprint embedding network model. The original audio data and each single sentence audio segment are input into the ECAPA-TDNN voiceprint embedding network model, which extracts the voiceprint features of the speaker, and an Agglomerative hierarchical clustering algorithm is used to cluster the extracted voiceprint features. For example, the clustering distance threshold is set to 0.85, so that each speech segment in the original audio data is assigned to a different speaker identity ID according to the similarity of the voiceprint features. If the distance between a new speech segment and all existing speakers is greater than 0.85, a new speaker identity ID is assigned to the new speech segment.

[0085] For the text content, the system adopts a Qwen3-asr-flash model, which is an end-to-end speech recognition model specifically designed for Chinese education scenarios. Each single sentence audio segment is sequentially input into the Qwen3-asr-flash model, which converts the current input single sentence audio segment into high-precision text content. The output is UTF-8 encoded punctuated text content, and then the next single sentence audio segment is input into the model until all single sentence audio segments complete text content recognition. The model has been specifically optimized for Chinese dialect dialogue scenarios and "Kyi", "Tse", and "Qie" mood words, and has good performance, with a character error rate of less than 5.2% on the AISHELL-3 test set.

[0086] For the emotion category, the system adopts a speech emotion recognition model with a Hidden Unit BERT architecture. Each single sentence audio segment is sequentially input into the speech emotion recognition model, which learns self-supervised speech representations from the current input single sentence audio segment. The speech emotion recognition model classifies the self-supervised speech representations through a classification head and outputs a 6-dimensional Softmax probability vector. The 6-dimensional Softmax probability vector corresponds to six emotion categories: anger, fear, happiness, neutral, sadness, and surprise. When the highest category confidence is not less than the preset value 0.7, the emotion category label is adopted, otherwise the current input single sentence audio segment is labeled as "emotion uncertain".

[0087] For paralinguistic features, the system uses the Qwen3-omni-30b-a3b model for paralinguistic feature recognition. The system inputs the following prompt word "Please act as a professional stage director, according to your listening to this audio, describe its tone and expression in natural language, and focus on speech rate, volume, tone fluctuation, voice texture, etc. Output only stage prompt, such as voice tremor, speech rate acceleration". The model outputs paralinguistic features such as "dragging long sound, tone rising" and "voice extremely small, tremor" under the guidance of the prompt word, and then checks the output format through regular expression to ensure that it is output in script format.

[0088] To avoid model hallucination interference, if the returned format does not conform to the format specified by the prompt, discard this output and re-prompt the model to recognize the paralinguistic features of the current single sentence audio segment.

[0089] Step S3, data integration.

[0090] After all single sentence audio segments have completed multi-dimensional feature extraction, all feature extraction results of each single sentence audio segment are collected into a structured JSON object, each JSON object containing complete information of the corresponding single sentence audio segment. The JSON object format can refer to the previous description, and this embodiment will not be repeated here.

[0091] After obtaining the JSON objects of all single sentence audio segments, format the output in chronological order, and the output format is uniform, for example, "Speaker_ID: Tone_Description, Emotion, Context". The generated dialogue representation text is as follows:

[0092] Speaker_01: (dragging long sound, rising tone, happy) Oh, isn't this our class' s college tyrant?

[0093] Speaker_02: (extremely small voice, tremor, unstable breath, fear) I... I just want to finish my homework...

[0094] The dialogue representation text with the above format allows the large language model to accurately perceive the tone, emotion and pressure of the dialogue between the two parties without listening to the original audio data, thereby improving the recognition accuracy of implicit bullying or whispering disputes.

[0095] It is worth noting that during the data integration process, if a field is missing, such as the abnormal segment described in the previous section, the abnormal segment is replaced with the corresponding placeholder, for example, (such as long silence), to ensure that the number of lines of the dialogue representation text is consistent with the number of sentences obtained by audio segmentation, facilitating subsequent alignment analysis.

[0096] Step S4, analysis report generation.

[0097] The system inputs the integrated dialogue representation text as high-dimensional context into the large language model Qwen-Plus for deep semantic analysis and education strategy generation.

[0098] To ensure the professionalism of the analysis, the system uses a thought chain prompting engineering technique. The system provides structured system prompt words to the large language model to guide the large language model to analyze according to the expert logic path.

[0099] For example, the system prompt words are as follows:

[0100] "Please analyze the following dialogue:

[0101] <dialogue representation text>

[0102] Based on the above information, please perform the following analysis steps:

[0103] 1) Dialogue nature and severity level judgment: Determine the nature of the dialogue according to the tone and content of the dialogue, such as bullying, quarrel, misunderstanding, etc.; evaluate the severity of the dialogue and explain the reasons.

[0104] 2) Psychological state portrait: Analyze the deep psychological needs of both parties in the dialogue, such as whether one party is seeking recognition or feeling ignored; identify any potential psychological danger signals, such as self-worth devaluation, excessive sensitivity, or other behavior signs that need attention?

[0105] 3) Prediction and risk assessment: Predict how the situation may develop if not properly handled; assess the impact of this interaction on students' long-term mental health, including but not limited to self-esteem, social skills, etc.

[0106] 4) Response strategy generation: Provide a detailed intervention plan for educators, including which party to talk to first, the specific entry point of the conversation, and follow-up education measures recommendations; and provide some practical examples of speech to help adults more effectively guide students to recognize and solve problems between them."

[0107] The large language model outputs a structured analysis report based on the above system prompt words.

[0108] Based on the above embodiments of the present application, the campus bullying behavior recognition method of the present application at least has the following advantages:

[0109] First, the embodiments of the present application can not only accurately transcribe the text content of the dialogue, but also synchronously extract and describe the paralinguistic features in the voice, such as tone change, rhythm abnormality, and voice quality characteristics, through the cooperation of multiple models. The time-sequenced dialogue representation text generated thereby completely retains the context, emotion, and tone information of the dialogue, providing rich and accurate context input for deep semantic reasoning based on large language models, so that educators can accurately grasp the dialogue dynamics through the generated report without relying on on-site audio playback, thereby improving the recognition ability and analysis accuracy of various campus interpersonal conflicts, including implicit bullying.

[0110] Second, the embodiments of the present application effectively improve the overall recognition performance and robustness of the system through the division of labor and cooperation of specialized models. The system divides the complex multi-dimensional feature extraction task into multiple independently optimized sub-tasks through models such as speaker identity recognition, text recognition, emotion category recognition, and paralinguistic feature recognition. Each model completes a specific function based on its targeted training data and algorithm, thereby avoiding the performance compromise and limitations of a single general-purpose model in multi-task processing, and improving the feature recognition accuracy in complex acoustic environments.

[0111] Third, the embodiments of the present application introduce a thinking chain prompt word engineering, which drives the large language model to perform deep context understanding and logical analysis on the structured dialogue representation text, enabling the system to simulate the logical analysis path of an educational psychology expert or a campus conflict mediation expert, gradually deduce and evaluate the psychological state, behavior motivation, and conflict evolution implied in the dialogue, and automatically generate targeted intervention strategies and action suggestions accordingly, thereby realizing the intelligent upgrade from passive monitoring to active decision support, improving the standardization and execution efficiency of campus safety management, and providing educators with professional support with practical operation guidance significance.

[0112] Fourth, the embodiments of the present application organize discrete paralinguistic features, text content, speaker identity, and emotion labels according to the time sequence of dialogue occurrence into structured dialogue representation text with high readability, which not only facilitates efficient access to key semantic information by educational managers, but also supports standardized archiving of dialogue records, thereby providing traceable and analyzable data basis for subsequent long-term student behavior analysis, psychological health assessment, and campus safety management decision-making, thereby improving the transparency of the behavior recognition process and the explainability of the results.

[0113] Figure 4 FIG. 1 is a schematic diagram of an electronic device according to an example embodiment. Please refer to Figure 4At the hardware level, the device includes a processor 402, an internal bus 404, a network interface 406, a memory 408, a hardware acceleration device 410, and a non-volatile memory 412, and of course, other hardware required by functions. One or more embodiments of the present application can be implemented in a software manner, such as reading a corresponding computer program from the non-volatile memory 412 into the memory 408 by the processor 402 and then running. Of course, in addition to the software implementation, one or more embodiments of the present application do not exclude other implementation manners, such as logic devices or a combination of software and hardware, etc., that is, the execution subject of the above processing flow is not limited to each logic unit, but can also be hardware or logic devices.

[0114] Figure 5 is a block diagram of a campus bullying behavior recognition device shown by an exemplary embodiment of the present application, which can be applied to an electronic device as shown in Figure 5 to implement the technical solutions of the present application. Among them, the campus bullying behavior recognition device includes a data acquisition unit 510, an audio segmentation unit 520, a feature extraction unit 530, a data integration unit 540, and an inference learning unit 550, wherein:

[0115] The data acquisition unit 510 is configured to acquire raw audio data of a campus environment.

[0116] The audio segmentation unit 520 is configured to perform segmentation processing on the raw audio data to obtain a plurality of single-sentence audio segments and corresponding timestamp information according to natural sentence division.

[0117] The feature extraction unit 530 is configured to perform multi-dimensional feature extraction on each single-sentence audio segment to obtain multi-dimensional feature data of each single-sentence audio segment.

[0118] The data integration unit 540 is configured to integrate the multi-dimensional feature data into a time-sequenced dialogue representation text according to the timestamp information of each single-sentence audio segment.

[0119] The inference learning unit 550 is configured to perform campus bullying behavior recognition and strategy analysis according to the dialogue representation text to generate a campus bullying behavior recognition result and a corresponding behavior intervention strategy.

[0120] In some embodiments, the audio segmentation unit 520 is configured to identify the raw audio data using a speech recognition and time alignment model to obtain transcription text data; determine a sentence boundary based on a specified punctuation mark or a silence duration between adjacent words in the transcription text data; cut out corresponding single-sentence audio segments from the raw audio data according to the sentence boundary, and record the start and end timestamps of each single-sentence audio segment.

[0121] In some embodiments, the audio segmentation unit 520 is further configured to, in a case where there is a blank segment in the transcribed text data that exceeds a preset time length, mark the blank segment as an abnormal segment, skip the multi-dimensional feature extraction step for the abnormal segment, and retain the corresponding time sequence position of the abnormal segment when generating the dialogue representation text.

[0122] In some embodiments, the feature extraction unit 530 is configured to perform voiceprint recognition and speaker clustering on the original audio data to obtain the speaker identity of each speech segment in the original audio data, match the speaker identity corresponding to each single-sentence audio segment based on the position of each single-sentence audio segment in the original audio data; and / or input each single-sentence audio segment into a speech recognition model in sequence to determine the text content corresponding to each single-sentence audio segment; and / or input each single-sentence audio segment into a speech emotion recognition model in sequence to determine the emotion category corresponding to each single-sentence audio segment; and / or input each single-sentence audio segment into a paralanguage recognition model in sequence, wherein the paralanguage recognition model performs paralanguage feature reasoning on the currently input single-sentence audio segment based on a preset cue word template to determine the paralanguage feature corresponding to the single-sentence audio segment.

[0123] In some embodiments, the data integration unit 540 is configured to aggregate the multi-dimensional feature data corresponding to each single-sentence audio segment into structured dialogue turn data according to the timestamp information thereof; and combine all dialogue turn data into the dialogue representation text according to the timestamp information; wherein each row in the dialogue representation text corresponds to a single-sentence audio segment and integrates the multi-dimensional feature data corresponding to the single-sentence audio segment.

[0124] In some embodiments, the inference learning unit 550 is configured to input the dialogue representation text into a pre-trained large language model, guide the large language model to perform inference on the dialogue representation text based on a preset thinking chain prompting engineering technology; and generate an analysis report according to the inference result, wherein the analysis includes campus bullying behavior recognition result, psychological state analysis, risk assessment, and behavior intervention strategy.

[0125] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts are described in the method embodiment. The device embodiments described above are only illustrative, wherein the units described as separate components can or can not be physically separated, and the components displayed as units can or can not be physical units, i.e., they can be located in one place or distributed on multiple network units. Some or all of the modules can be selected to achieve the purposes of the present application according to actual needs. Those skilled in the art can understand and implement without creative labor.

[0126] Correspondingly, the present application also provides a computer readable storage medium, having stored thereon a computer program, which, when executed by a processor, implements the method of any one of the above embodiments.

[0127] Correspondingly, the present application also provides a computer program product configured to perform the method of any one of the above embodiments.

[0128] The system, apparatus, module or unit illustrated in the above embodiments can be specifically implemented by a computer chip or entity, or by a product with certain functions. A typical implementation device is a computer, and the specific form of the computer can be a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an e-mail device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0129] In a typical configuration, a computer includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0130] The memory can include non-persistent memory in computer readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or Flash memory. The memory is an example of computer readable media.

[0131] Computer readable media includes permanent and non-permanent, removable and non-removable media implemented by any method or technology for storing information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, disk storage, quantum memory, graphene-based storage medium or other magnetic storage device, or any other non-transmission medium that can be used to store information accessible by a computing device. According to the definition herein, computer readable media does not include transitory computer readable media, such as modulated data signals and carriers.

[0132] Although this description contains many specifics, these should not be construed as limiting the scope of any invention or application, but merely as describing features that can be incorporated into a specific embodiment of the application. The description herein of certain features can also describe features that can be combined with features of other embodiments. Similarly, the description herein of certain embodiments can also describe features that can be combined with features of other embodiments. Furthermore, while features can be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination and the claimed combination can be directed to a subcombination or variation of a subcombination.

[0133] Similarly, while operations are depicted in the drawings in a particular order, this should not be understood as requiring such order, nor that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing can be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0134] Accordingly, particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. In some cases, actions recited in the claims can be performed in a different order and still achieve desirable results. In addition, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve the desired results. In some implementations, multitasking and parallel processing can be advantageous.

[0135] It is also noted that the terms "comprising", "including", or any other similar phrase means "including but not limited to", and the singular forms "a", "an", and "the" include plural references unless the context clearly dictates otherwise. It is further noted that the claims can be drafted to exclude any optional element or step. As such, any element or step recited in the claims that does not specifically exclude any other element or step can be assumed to be optional unless otherwise specified.

[0136] The above description is merely illustrative of the application and not restrictive. Since certain changes can be made in the application without departing from the spirit and scope of the application, it is intended that all of the subject matter of the above description be covered by the following claims.

Claims

1. A method for identifying school bullying behavior, characterized in that, The method includes: Obtain raw audio data of the campus environment; The original audio data is segmented to obtain multiple single-sentence audio segments divided according to natural sentences and their corresponding timestamp information; Multidimensional feature extraction is performed on each individual audio segment to obtain multidimensional feature data of each individual audio segment. Based on the timestamp information of each individual audio segment, the multi-dimensional feature data is integrated into a temporal dialogue representation text; Based on the dialogue text, school bullying behavior identification and strategy analysis are performed to generate school bullying behavior identification results and corresponding behavioral intervention strategies.

2. The method according to claim 1, characterized in that, The multi-dimensional feature data includes text content, the identity of the speaker, the emotion category, and paralinguistic features.

3. The method according to claim 1, characterized in that, The process of segmenting the original audio data to obtain multiple single-sentence audio segments divided according to natural language and their corresponding timestamp information includes: The original audio data is identified using a speech recognition and time alignment model to obtain transcribed text data; Sentence boundaries are determined based on specified punctuation marks or the duration of silence between adjacent words in the transcribed text data. Based on the sentence boundaries, the corresponding single-sentence audio segments are cut out from the original audio data, and the start and end timestamps of each single-sentence audio segment are recorded.

4. The method according to claim 3, characterized in that, In the event that blank segments exceeding a preset duration exist in the transcribed text data, the method further includes: The blank fragments are marked as anomalous fragments, the multidimensional feature extraction step is skipped for the anomalous fragments, and their corresponding temporal positions are preserved when generating the dialogue representation text.

5. The method according to claim 1, characterized in that, The process of extracting multi-dimensional features from each individual audio segment to obtain multi-dimensional feature data for each individual audio segment includes: Voiceprint recognition and speech object clustering are performed on the original audio data to obtain the speech object identity of each speech segment in the original audio data. Based on the position of each single-sentence audio segment in the original audio data, the speech object identity corresponding to each single-sentence audio segment is matched. And / or, input each single-sentence audio segment sequentially into the speech recognition model to determine the text content corresponding to each single-sentence audio segment; And / or, input each single audio segment sequentially into the speech emotion recognition model to determine the emotion category corresponding to each single audio segment; And / or, each single-sentence audio segment is sequentially input into the paralinguistic recognition model, which performs paralinguistic feature inference on the currently input single-sentence audio segment based on a preset prompt word template, and determines the paralinguistic features corresponding to the single-sentence audio segment.

6. The method according to claim 1, characterized in that, The step of integrating the multi-dimensional feature data into a temporal dialogue representation text based on the timestamp information of each individual audio segment includes: The multi-dimensional feature data corresponding to each single audio segment are aggregated into structured dialogue turn data according to their timestamp information; All dialogue rounds data are combined into the dialogue representation text according to the timestamp information; In this context, each line in the dialogue representation text corresponds to a single-sentence audio segment, and integrates the multi-dimensional feature data corresponding to that single-sentence audio segment.

7. The method according to claim 1, characterized in that, The step of identifying and analyzing bullying behavior and strategies based on the dialogue representation text to generate bullying behavior identification results and corresponding behavioral intervention strategies includes: The dialogue representation text is input into a pre-trained large language model, and the large language model is guided to reason about the dialogue representation text based on a preset thought chain prompting engineering technique. An analysis report is generated based on the reasoning results. The analysis includes the identification results of school bullying behavior, psychological state analysis, risk assessment, and behavioral intervention strategies.

8. A device for identifying school bullying behavior, characterized in that, The device includes: The data acquisition unit is used to acquire raw audio data from the campus environment. The audio segmentation unit is used to segment the original audio data to obtain multiple single-sentence audio segments divided according to natural sentences and their corresponding timestamp information; The feature extraction unit is used to extract multi-dimensional features from each single-sentence audio segment to obtain multi-dimensional feature data of each single-sentence audio segment. The data integration unit is used to integrate the multi-dimensional feature data into a time-series dialogue representation text based on the timestamp information of each single audio segment. The reasoning learning unit is used to identify and analyze school bullying behavior based on the dialogue representation text, so as to generate school bullying behavior identification results and corresponding behavioral intervention strategies.

9. An electronic device, characterized in that, include: processor; as well as A computer-readable storage medium storing computer program instructions that, when executed by the processor, cause the processor to perform the method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which is executed by a processor according to any one of claims 1 to 7.