Non-suicide self-injury risk early warning method and system for teenagers

By collecting voice information from adolescents using portable sensors, performing multi-dimensional feature extraction and neural network analysis of self-attention mechanisms, the problem of concealment of non-suicidal self-harm behaviors among adolescents has been solved, enabling efficient risk warning and early intervention.

CN121260184APending Publication Date: 2026-01-02淮安市第三人民医院
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511824553.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-05
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Non-suicidal self-harm behaviors among adolescents are often concealed, making early identification and intervention difficult, and current technologies lack effective early warning methods.

Method used

By collecting speech information from adolescents using portable miniature sensors, performing multi-dimensional feature extraction and self-attention mechanism neural network model analysis, the risk of non-suicidal self-harm can be predicted, and corresponding intervention programs can be provided.

Benefits of technology

It significantly improves the accuracy and timeliness of risk warnings for non-suicidal self-harm behaviors, provides a reliable basis for early intervention, and enhances the accuracy and comprehensiveness of prediction results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121260184A_ABST
    Figure CN121260184A_ABST
Patent Text Reader

Abstract

The invention discloses a teenager non-suicide self-injury risk early warning method and system, and the method comprises the steps: firstly collecting teenager voice information, carrying out the information preprocessing, carrying out the feature extraction of a preprocessed voice signal through three independent modules, extracting an emotional tendency feature through a first module, extracting an intonation feature through a second module, and carrying out the early warning of the teenager non-suicide self-injury risk. The third module extracts speech speed features. And based on the information extracted by the three modules, predicting the non-suicide self-injury behavior probability of the teenagers by using an artificial intelligence algorithm. And the non-suicide self-injury probability is graded, and related intervention measures are given. According to the method and the system, through sequential voice information acquisition, multi-dimensional feature hierarchical extraction and self-attention mechanism driven feature association modeling, and in combination with dialogue object classification analysis, potential signals of non-suicide self-injury behaviors of teenagers can be effectively captured, the accuracy and timeliness of risk early warning are remarkably improved, and the risk early warning efficiency is improved. And a reliable basis is provided for early intervention.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of mental health monitoring, and particularly relates to a youth nonsuicidal self-injury risk early warning method and system. BACKGROUND

[0002] Nonsuicidal self-injury (NSSI) is one of the most common mental health and problem behaviors in the youth group. As a non-adaptive behavior of individuals (such as anxiety, depression, interpersonal communication and academic difficulties), NSSI is not recognized by society, which not only damages the physical and mental health of the youth, but also increases the risk of suicide. The universality and harmfulness of NSSI make it a more and more urgent public health problem and attract more and more research attention.

[0003] In view of the strong concealment of NSSI, it is difficult to discover many NSSI occurring in the youth group, so early identification and intervention cannot be well achieved. In view of the obvious contrast between the core performance and daily state of NSSI in language: angry / frustrated emotional language signals, depressed / depressed emotional language signals, anxious / tense emotional language signals, etc., the present application monitors the bad mood of the youth nonsuicidal self-injury through the language characteristics of the youth nonsuicidal self-injury, including emotional tendency, tone, and speed, etc. Multimodal information, classifies and warns the youth nonsuicidal self-injury behavior and intervenes. SUMMARY

[0004] The purpose of the present application is to provide a youth nonsuicidal self-injury risk early warning method and system, which effectively predicts the youth nonsuicidal self-injury behavior, so as to intervene in advance to reduce the adverse consequences caused thereby.

[0005] The technical solution of the present application is: A youth nonsuicidal self-injury risk early warning method, comprising the following steps: S1, collecting voice information of the youth in a set period, and converting the voice information into text, and then classifying and counting the dialogue objects, the dialogue objects including six categories: self-talk, dialogue with parents, dialogue with teachers, dialogue with friends, dialogue with relatives, and dialogue with strangers; S2, preprocessing the collected youth voice information, and finally cutting into short frames of fixed length for subsequent feature extraction; S3, multi-dimensional feature extraction is performed on the collected youth voice information, three modules are adopted, the first module extracts emotional features in the voice through an emotional tendency extraction module, the emotional features including positive, negative, and neutral, the second module extracts the fundamental frequency and its dynamic change features through a tone extraction module, and the third module calculates the number of phonemes or the proportion of effective voice per unit time through a speech rate extraction module, forming a multi-dimensional voice feature set; S4, input the emotional features, tone, and speech rate multi-dimensional voice features extracted in step S3 into a self-attention mechanism neural network model, focus on the correlation between the features through the self-attention mechanism, perform self-adaptive weight distribution and deep feature fusion on the input features, and then output the risk probability of the existence of non-suicidal self-injury behavior of the teenager through reasoning calculation, and finally realize the accurate early warning of the non-suicidal self-injury risk of the teenager; S5, give a corresponding intervention scheme according to the early warning risk level.

[0006] Preferably, in step S1, a portable micro sensor is used to collect voice information. The portable micro sensor is adapted to the daily wearing scene of teenagers, so as to realize continuous and stable collection of voice information.

[0007] Preferably, in step S2, when the extracted voice signal is preprocessed, first, the blank segments in the signal are removed through endpoint detection, then environmental noise interference is suppressed through a noise reduction algorithm, and finally the voice components of other people are separated through voice separation technology. Finally, the processed pure voice continuous signal is cut into short frames of fixed length, which is used for subsequent feature extraction.

[0008] Preferably, in step S3, the emotion tendency extraction module of the first module uses a convolutional neural network (CNN) or a long short-term memory network (LSTM) as the core algorithm architecture to perform feature learning and classification on the preprocessed voice data, and finally outputs the emotional tendency result in the voice of the teenager, providing the core input features in the emotional dimension for the subsequent risk early warning model.

[0009] Preferably, in step S3, the tone extraction module of the second module uses extraction parameters adapted to the fundamental frequency range of voice signals of different genders through a preset gender classification logic for targeted processing to improve the accuracy and stability of the tone feature extraction, and provides more reliable tone dimension features for the subsequent risk early warning model.

[0010] Preferably, in step S3, the speech rate extraction module of the third module distinguishes male and female voices, and uses the ratio of the number of pronunciation units to the effective time length in the effective voice segment as the core logic of speech rate quantization to output the numerical value reflecting the speech rate features of the voice of the teenager, and provides input data in the speech rate dimension for the subsequent risk early warning model.

[0011] Preferably, in step S4, the self-attention mechanism neural network model automatically assigns differentiated weights to the three-dimensional features by learning the correlation between the features, and after deep feature fusion and reasoning calculation, outputs the risk probability of the existence of non-suicidal self-injury behavior of the teenager and outputs the source of negative emotions, realizing early warning driven by multi-dimensional features.

[0012] Preferably, in step S5, the probability of non-suicidal self-injury of adolescents is divided into 4 levels, respectively: Probability below 40% indicates almost no risk, no self-injury behavior; Probability of 40-60% indicates risk, possible self-injury behavior; Probability of 60-80% indicates more serious, self-injury behavior; Probability of 80-100% indicates very serious, self-injury behavior is serious, and immediate medical treatment is needed.

[0013] Preferably, in step S5, according to the risk level and the analysis result of the source of bad mood, online intervention or offline intervention is given, wherein the online intervention sets relevant psychological intervention countermeasures in the system.

[0014] A risk warning system for non-suicidal self-injury of adolescents, comprising: An information collection module collects voice information of adolescents within a set period through an acoustic-electric conversion sensor, and converts the voice information into text; A dialogue object classification module classifies the collected voice information according to dialogue objects, which include six categories: self-talk, dialogue with parents, dialogue with teachers, dialogue with friends, dialogue with relatives, and dialogue with strangers; A signal preprocessing module removes blank segments in the signal, reduces signal noise, separates other people's voices, and finally cuts the processed pure voice continuous signal into short frames of fixed length; A multi-modal voice signal analysis module connected to the signal preprocessing module is configured to extract multi-dimensional features from the collected voice information of adolescents, specifically including: An emotional tendency extraction submodule for extracting emotional features from voice information, including positive tendency, negative tendency, and neutral tendency; A tone extraction submodule for extracting the fundamental frequency and its dynamic change characteristics of the voice; A speech rate extraction submodule for calculating the number of phonemes or the proportion of effective speech in a unit of time to represent the speech rate feature; The multi-modal voice signal analysis module integrates the above-mentioned emotional tendency, tone, and speech rate features to form a multi-dimensional voice feature set; A risk prediction module connected to the multi-modal voice signal analysis module is configured to: Receive the multi-dimensional voice features as input; adopt a self-attention mechanism neural network model to process the input features, focus on the correlation between features through the self-attention mechanism, realize adaptive weight distribution and deep feature fusion of the input features; after model inference calculation, output the risk probability of non-suicidal self-injury behavior of adolescents, and complete accurate risk warning.

[0015] An intervention scheme generation module connected to the risk prediction module and configured to automatically match and generate a corresponding intervention scheme according to the early warning risk level output by the risk prediction module.

[0016] A computer readable storage medium having stored thereon a computer program, wherein the program, when executed by a processor, implements the early warning method for non-suicidal self-injury risk of adolescents. Advantages

[0017] The method and system can effectively capture the potential signals of non-suicidal self-injury behavior of adolescents, significantly improve the accuracy and timeliness of risk early warning, and provide a reliable basis for early intervention, by means of time-series voice information collection, multi-dimensional feature hierarchical extraction, feature correlation modeling driven by self-attention mechanism, and combined with dialogue object classification analysis. The specific advantages include the following: 1. The present application uses voice information in a set period as input data to predict the risk of non-suicidal self-injury behavior of adolescents. Considering the limitations of voice information collected in a short period (difficult to fully reflect the true state), which is not representative, the present application collects voice data in multiple periods (such as one day, one week, one month or one year), avoiding the limitations of single-period voice, thereby improving the accuracy of the prediction result.

[0018] 2. The present application fuses 3-dimensional features in the voice information of adolescents, including emotional features, tone features, and speech rate features. By increasing the information dimension, the prediction model is closer to the practical experience of professionals or professional institutions in evaluating non-suicidal self-injury behavior.

[0019] 3. The self-attention mechanism used in the prediction of non-suicidal self-injury behavior of adolescents in the present application greatly improves the accuracy of the model in evaluating non-suicidal self-injury behavior. BRIEF DESCRIPTION OF DRAWINGS

[0020] The present application will be further described below in conjunction with the accompanying drawings and embodiments: Figure 1 The flowchart of the early warning method for non-suicidal self-injury risk of adolescents of the present application; Figure 2 The flowchart of multi-modal voice signal analysis; Figure 3 The structure diagram of the non-suicidal self-injury behavior risk early warning system of adolescents.

[0021] Figure 3 In the figure, 1, information collection module; 2, dialogue object classification module; 3, signal preprocessing module; 4, multi-modal voice signal analysis module; 4a, emotional tendency extraction submodule; 4b, tone extraction submodule; 4c, speech rate extraction submodule; 5, risk prediction module; 6, intervention scheme generation module. Detailed Implementation Example

[0022] like Figure 1 As shown, this embodiment provides a method for early warning of non-suicidal self-harm behavior risks in adolescents, including the following steps: S1. Collect the voice information of teenagers within a set period, convert the voice information into text, and then classify and statistically analyze it according to the dialogue object. The dialogue object includes six categories: talking to oneself, talking to parents, talking to teachers, talking to friends, talking to relatives, and talking to strangers. It uses a portable miniature sensor to collect voice information, which is suitable for teenagers to wear in daily life, and can achieve continuous and stable collection of voice information.

[0023] Set a period of one day, one week, one month, or one year. By comparing the changes every day, week, month, and year, the prediction results become more reliable, and the changes in emotions over time are captured. Pack a bag each day, record a set of evaluation results, form a time series chart, and finally summarize the data over a period of time for a comprehensive evaluation.

[0024] S2. The collected adolescent speech information is preprocessed and finally segmented into short frames of fixed length for subsequent feature extraction. When preprocessing the extracted speech signal, firstly, blank segments in the signal are removed by endpoint detection, then noise reduction algorithm is used to suppress environmental noise interference, and finally, speech separation technology is used to separate the speech components of other people. Finally, the processed clean speech continuous signal is divided into short frames of fixed length for subsequent feature extraction.

[0025] S3. Multi-dimensional feature extraction is performed on the collected adolescent speech information. Three modules are used: the first module extracts the emotional features in the speech through the emotion tendency extraction module, and the emotional features include positive, negative and neutral. The second module extracts the fundamental frequency and its dynamic change features through the intonation extraction module. The third module calculates the number of syllables or the effective speech ratio per unit time through the speech rate extraction module to form a multi-dimensional speech feature set. The sentiment extraction module of the first module uses a convolutional neural network (CNN) or a long short-term memory network (LSTM) as the core algorithm architecture to perform feature learning and classification on the preprocessed speech data, and finally output the sentiment results of adolescent speech, providing core input features of the sentiment dimension for the subsequent risk warning model.

[0026] The tone extraction module of the second module is responsible for tone extraction. The core of extracting the tone height is to obtain the fundamental frequency of the voice signal. In view of the significant difference in the fundamental frequency distribution range of males and females in the youth group (males are usually 80-250 Hz, and females are 160-500 Hz), the module uses the extraction parameters suitable for the fundamental frequency range of voice signals of different genders through preset gender classification logic for targeted processing to improve the accuracy and stability of the tone height feature extraction, and provides more reliable tone dimension features for the subsequent risk warning model.

[0027] The speech rate extraction module of the third module is responsible for speech rate extraction. In view of the different speech rates of men and women, the male and female voices are distinguished, and the ratio of the number of pronunciation units to the effective time length in the effective speech segment is used as the core logic of speech rate quantization. The value reflecting the speech rate characteristics of the youth voice is output, and the input data of the speech rate dimension for the subsequent risk warning model is provided.

[0028] S4, as Figure 2 shown, the emotional features, tone, and speech rate multi-dimensional voice features extracted in step S3 are input into a self-attention mechanism neural network model. The self-attention mechanism focuses on the correlation between the features, such as the coupling of negative emotions and slowed speech rate. The input features are adaptively weighted and deep feature fusion, and then the risk probability of the youth having non-suicidal self-injury behavior is output through reasoning calculation. Finally, the precise warning of the risk of non-suicidal self-injury of the youth is realized; The self-attention mechanism neural network model automatically assigns different weights to the three dimensional features by learning the correlation between the features. After deep feature fusion and reasoning calculation, the risk probability of the youth having non-suicidal self-injury behavior is output, and the source of negative emotions is also output, such as from self-talk or from communication with others, realizing the warning driven by multi-dimensional features.

[0029] S5, for the warning risk level, the corresponding intervention scheme is given. The probability of non-suicidal self-injury of the youth is divided into four levels, which are: Probability below 40%, indicating almost no risk, no self-injury behavior; Probability of 40-60%, indicating risk, possible self-injury behavior; Probability of 60-80%, indicating more serious, self-injury behavior; Probability of 80-100%, indicating very serious, self-injury behavior is serious, and immediate medical treatment is needed.

[0030] According to the risk level and the analysis result of the source of negative emotions, online intervention or offline intervention is given, and the online intervention sets relevant psychological intervention countermeasures in the system. Embodiment

[0031] As Figure 3As shown, the embodiment provides a risk early warning system for adolescent non-suicidal self-injury behavior, comprising the following modules.

[0032] The information collection module 1, the dialogue object classification module 2, the signal preprocessing module 3, the multi-modal speech signal analysis module 4, the risk prediction module 5 and the intervention scheme generation module 6 realize the risk early warning and intelligent generation of intervention scheme target through cooperative division and efficient linkage, and the specific process is as follows: The sound-electricity conversion sensor used by the information collection module 1 needs to meet the design requirements of being small and light, so as to adapt to the shape limitation of wearable devices such as bracelets, and at the same time have the function characteristics of realizing continuous collection of speech information in the device. The selection requirements are as follows: volume: diameter ≤6mm, thickness ≤3mm (adapt to the narrow space of bracelet), power consumption: working current ≤1mA (avoid frequent charging and support continuous collection), sensitivity: -38dBV~-42dBV (ensure clear pickup of speech signal at close distance (30-50cm)), such as micro-electromechanical system microphone.

[0033] When classifying dialogue objects, speech information is collected at a sampling rate of 10kHz, and four sampling periods are set: one day, one week, one month and one year. The daily sampled speech information is packaged into an audio file, the file is named "adolescent name-X year X month X day", and is stored in the bracelet.

[0034] The dialogue object classification module 2 performs speech-to-text processing on the daily audio file: uses the Whisper-large-v2 model based on the Transformer encoder-decoder architecture to convert speech into text sequences, and at the same time identifies the dialogue object through voiceprint recognition technology (such as CNN-based voiceprint classification model) -pre-recorded voiceprint feature library of the adolescent himself, parents, teachers, friends, relatives and strangers, through calculating the cosine similarity (threshold set to 0.85) between real-time voiceprint and library features, classifying the speech into "soliloquy" (no other voiceprint matching), "dialogue with parents", "dialogue with teachers", "dialogue with friends", "dialogue with relatives" and "dialogue with strangers" a total of 6 categories, and each category separately counts the basic data such as daily speech duration and dialogue frequency, and stores them in a relational database (such as MySQL).

[0035] The signal preprocessing module 3 carries out fine preprocessing on the collected original speech signal according to the progressive process of "noise suppression-interference separation-framing processing", and provides high-quality data support for subsequent feature extraction. (1) Using endpoint detection and blank segment rejection: using a double-threshold method (based on speech signal energy and zero-crossing rate features) for endpoint detection, identifying and rejecting the silent blank segments (including non-speech background blank, device idle noise segment) in the speech signal, retaining the speech interval containing effective semantics, and reducing the interference of invalid data on subsequent processing; (2) Adaptive noise suppression: using the spectral subtraction method combined with the wavelet threshold denoising algorithm, first extracting pure noise samples (environmental noise in the blank segment) through the voice activity detection (VAD) module, constructing a noise spectrum model, and then performing spectral subtraction operation on the effective speech interval to suppress the steady-state noise (such as device running sound, background human voice) and non-steady-state noise (such as sudden interference sound) in the environment, while retaining the detail features (such as tone changes) of the speech signal through wavelet threshold processing, ensuring the integrity of the speech information after noise reduction; (3) Multi-speaker speech separation: introducing the Conv-TasNet speech separation model, extracting the time-frequency features of the mixed speech through the encoder, using the attention mechanism to distinguish the speech feature differences (such as voiceprint, tone, speech rate) of different speakers, generating exclusive speech masks, separating the speech components of the target youth and other dialogues, and realizing accurate extraction of single-target speech; (4) Speech framing processing: using the Hanning window function to frame the separated pure speech continuous signal, cutting the continuous speech into a fixed-length short frame sequence, which not only ensures the stationarity of the speech signal in a single frame, but also retains the time sequence correlation features between frames through frame overlap, laying a foundation for subsequent multi-dimensional feature extraction of emotional tendency, tone, speech rate, etc.

[0036] The multi-modal speech signal analysis module 4 receives the audio files and classification results output by the information collection and classification module, and carries out multi-dimensional feature extraction: Emotional tendency extraction sub-module 4a: using a pre-trained speech emotion recognition model (such as an LSTM-based SER model), inputting the pre-processed speech information, outputting the emotional tendency probability distribution (positive, negative, neutral, etc.), taking the class with the highest probability as the emotional label of the speech segment, and calculating the proportion of negative emotions in each classification scenario (such as conversation with parents) per day; Tone extraction sub-module 4b: extracting the fundamental frequency (F0) sequence of the audio through the Praat speech analysis tool, calculating the mean, standard deviation, maximum, minimum, and dynamic change rate (such as the F0 difference within 500ms) of the fundamental frequency, representing the stability and fluctuation characteristics of the tone; Speech rate extraction sub-module 4c: combined with the voice information, the number of effective syllables (after removing pauses and emotional words) in a unit of time (1 minute) is counted, and the proportion of effective speech duration in the total collection duration is calculated (effective speech refers to the speech segment containing semantics, which is identified by the voice activity detection VAD model), which is used as a speech rate feature; The above features are integrated into a multi-dimensional feature vector (such as [negative sentiment proportion, fundamental frequency standard deviation, number of effective syllables per minute, dialogue object category code]), which is used for subsequent model input.

[0037] Risk prediction module 5, using a neural network model based on self-attention mechanism (such as Transformer encoder), the specific structure includes input layer, embedding layer, 4 layers of self-attention coding layer and output layer: The input layer receives the multi-dimensional feature vector; the embedding layer converts the dialogue object classification features (discrete values) into a 256-dimensional embedding vector, which is concatenated with the continuous features (standardized) such as emotion, tone, and speech rate to form a 512-dimensional feature matrix; The self-attention coding layer calculates the correlation weight between features (such as the coupling coefficient of "negative emotion in conversation with strangers" and "slowed speech rate") through the multi-head attention mechanism (8 heads), adaptively strengthens the influence of key features (such as the weight of the negative emotion + low speech rate combination is increased by 20%), and stabilizes the training through residual connection and Layer Normalization; The output layer uses a Sigmoid activation function to output a risk probability value between 0 and 1 (0 indicates no risk, and 1 indicates extremely high risk), and is divided into 4 risk levels according to the probability value (below 40% is almost risk-free, 40-60% is risky, 60-80% is more serious, and 80-100% is very serious). Model training phase: a voice dataset containing 5000 teenagers (12-18 years old) (including non-suicidal self-injury behavior labels) is used, the training set and the test set are divided according to 7:3, the model parameters are optimized with cross-entropy loss function, and the AUC value of the final test set is ≥0.92, ensuring the prediction accuracy.

[0038] Intervention scheme generation module 6, which has a risk level-intervention scheme mapping library built in, outputs corresponding schemes for different risk levels: Almost no risk (below 40%): generate a "daily attention scheme", suggesting that parents have 1 in-depth communication with teenagers every week, and record emotional changes; At risk (40%-60%): generate a "family guidance scheme" containing 3 sets of parent-child interaction activities (such as joint exercise, interest cultivation) and emotional management tips; Severe (60%-80%): Develop a "school-family collaborative program", inform the class teacher to pay attention to the performance of adolescents in school, and jointly carry out counseling with the psychological teacher or the mental health promotion center once every 1-2 weeks, and dynamically evaluate the psychosomatic condition.

[0039] Very severe (80%-100%): Start "emergency intervention process", inform parents, schools and mental health promotion centers at the same time, coordinate professional assessment and intervention within 24 hours, and go to the hospital in time.

[0040] The embodiment realizes the full-process automation from voice information collection, multi-dimensional feature extraction, intelligent risk prediction to precise intervention scheme generation through the cooperative work of the above-mentioned modules, can effectively capture the potential signals of non-suicidal self-injury behavior of adolescents, and provides a scientific basis for early intervention.

[0041] The embodiment of the application also provides a terminal, which comprises one or more processors and a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors realize the non-suicidal self-injury risk early warning method of adolescents as described in any one of the above embodiments.

[0042] The embodiment of the application also provides a computer readable storage medium, which stores a computer program, and the program is characterized in that when the program is executed by a processor, the non-suicidal self-injury risk early warning method of adolescents as described in any one of the above embodiments is realized.

[0043] The above embodiments are only for illustrating the technical concept and characteristics of the application, and the purpose is to enable those skilled in the art to understand the content of the application and implement it, and cannot limit the protection scope of the application. Any modification made according to the spirit and essence of the main technical solution of the application should be covered within the protection scope of the application.

Claims

1. A method for early warning of non-suicidal self-harm risks in adolescents, characterized in that, Includes the following steps: S1. Collect the voice information of teenagers within a set period, convert the voice information into text, and then classify and statistically analyze it according to the dialogue object. The dialogue object includes six categories: talking to oneself, talking to parents, talking to teachers, talking to friends, talking to relatives, and talking to strangers. S2. The collected adolescent speech information is preprocessed and finally segmented into short frames of fixed length for subsequent feature extraction. S3. Multi-dimensional feature extraction is performed on the collected adolescent speech information. Three modules are used: the first module extracts the emotional features in the speech through the emotion tendency extraction module, and the emotional features include positive, negative and neutral. The second module extracts the fundamental frequency and its dynamic change features through the intonation extraction module. The third module calculates the number of syllables or the effective speech ratio per unit time through the speech rate extraction module to form a multi-dimensional speech feature set. S4. Input the multi-dimensional speech features of emotion, tone, and speech rate extracted in step S3 into the self-attention mechanism neural network model. By focusing on the correlation between features through the self-attention mechanism, adaptive weight allocation and deep feature fusion are performed on the input features. Then, through inference calculation, the probability of non-suicidal self-harm behavior in adolescents is output, and finally, accurate early warning of non-suicidal self-harm risk in adolescents is achieved. S5. Provide corresponding intervention plans based on the warning risk level.

2. The method for early warning of non-suicidal self-harm risks in adolescents according to claim 1, characterized in that, In step S1, a portable miniature sensor is used to collect voice information. The portable miniature sensor is suitable for daily wear scenarios of teenagers to achieve continuous and stable collection of voice information.

3. The method for early warning of non-suicidal self-harm risks in adolescents according to claim 1, characterized in that, In step S2, when preprocessing the extracted speech signal, firstly, blank segments in the signal are removed by endpoint detection, then a noise reduction algorithm is used to suppress environmental noise interference, and finally, speech separation technology is used to separate the speech components of other people. Finally, the processed clean speech continuous signal is divided into short frames of fixed length for subsequent feature extraction.

4. The method for early warning of non-suicidal self-harm risks in adolescents according to claim 1, characterized in that, In step S3, the sentiment tendency extraction module of the first module uses a convolutional neural network (CNN) or a long short-term memory network (LSTM) as the core algorithm architecture to perform feature learning and classification on the preprocessed speech data, and finally outputs the sentiment tendency results in the speech of teenagers, providing core input features of the sentiment dimension for the subsequent risk warning model.

5. The method for early warning of non-suicidal self-harm risks in adolescents according to claim 1, characterized in that, In step S3, the intonation extraction module of the second module uses preset gender classification logic to perform targeted processing on the speech signals of different genders by adopting extraction parameters adapted to their fundamental frequency range, so as to improve the accuracy and stability of intonation feature extraction and provide more reliable intonation dimension features for subsequent risk warning models.

6. The method for early warning of non-suicidal self-harm risks in adolescents according to claim 1, characterized in that, In step S3, the speech rate extraction module of the third module distinguishes between male and female speech, uses the ratio of the number of pronunciation units in the effective speech segment to the effective duration as the core logic for speech rate quantification, and outputs a value that reflects the speech rate characteristics of teenagers, providing input data for the speech rate dimension of the subsequent risk warning model.

7. The method for early warning of non-suicidal self-harm risks in adolescents according to claim 1, characterized in that, In step S4, the self-attention mechanism neural network model learns the correlation between features, automatically assigns differentiated weights to the three-dimensional features, and outputs the risk probability of non-suicidal self-harm behavior in adolescents after deep feature fusion and inference calculation, and outputs the source of negative emotions, thereby realizing early warning driven by multi-dimensional features.

8. The method for early warning of non-suicidal self-harm risks in adolescents according to claim 1, characterized in that, In step S5, the probability of non-suicidal self-harm among adolescents is divided into four levels: A probability of less than 40% indicates almost no risk and no risk of self-harm. A probability of 40-60% indicates a risk, with the possibility of self-harm. A probability of 60-80% indicates a relatively serious condition, with self-harming behavior. A probability of 80-100% indicates a very serious condition; the self-harm behavior is severe and requires immediate medical attention.

9. The method for early warning of non-suicidal self-harm risks in adolescents according to claim 1, characterized in that, In step S5, based on the analysis results of risk level and sources of negative emotions, online or offline interventions are provided, with online interventions involving setting up relevant psychological intervention strategies within the system.

10. A risk warning system for non-suicidal self-harm among adolescents, characterized in that, include: The information acquisition module (1) acquires the speech information of teenagers within a set period through the sound-to-electric conversion sensor and converts the speech information into text; The dialogue object classification module (2) classifies the collected voice information according to the dialogue object. The dialogue objects include six categories: talking to oneself, talking to parents, talking to teachers, talking to friends, talking to relatives, and talking to strangers. The signal preprocessing module (3) removes blank segments from the signal, reduces signal noise, separates other people's speech, and finally divides the processed pure speech continuous signal into short frames of fixed length. The multimodal speech signal analysis module (4), connected to the signal preprocessing module (3), is configured to perform multi-dimensional feature extraction on the collected adolescent speech information, specifically including: The sentiment extraction submodule (4a) is used to extract sentiment features from speech information, including positive sentiment, negative sentiment, neutral sentiment, etc. The intonation extraction submodule (4b) is used to extract the fundamental frequency of speech and its dynamic variation features; The speech rate extraction submodule (4c) is used to calculate the number of syllables or the percentage of effective speech per unit time to characterize speech rate features; The multimodal speech signal analysis module (4) integrates the above-mentioned emotional tendency, intonation, and speech rate features to form a multi-dimensional speech feature set; The risk prediction module (5), connected to the multimodal speech signal analysis module (4), is configured as follows: The system receives the multi-dimensional voice features as input; it processes the input features using a self-attention mechanism neural network model, focusing on the correlation between features through the self-attention mechanism to achieve adaptive weight allocation and deep feature fusion of the input features; and through model inference calculation, it outputs the risk probability of adolescents exhibiting non-suicidal self-harm behavior, thus completing accurate risk warning. The intervention plan generation module (6) is connected to the risk prediction module (5) and is configured to automatically match and generate the corresponding intervention plan based on the early warning risk level output by the risk prediction module (5).

11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by the processor, the program implements the method for early warning of non-suicidal self-harm risks in adolescents as described in any one of claims 1-9.

Citation Information

Patent Citations

  • Self-injury suicide risk assessment method and device, electronic equipment and storage medium

    CN115631772A

  • Bad emotion reminding system

    CN119323978A

  • Negative emotion characterization analysis system for social contact in colleges and universities

    CN119476311A

  • Intelligent psychological intervention system based on multi-modal fusion

    CN120337162A

  • Cognitive disease intelligent risk management method based on multi-mode voiceprint data analysis

    CN120452481A