Multi-mode suicide risk detection method, electronic equipment, storage medium and program product
By extracting audio and text features from user voice data and combining speech, emotion, language and clinically related features for suicide risk detection, the problem of failure to fully utilize multimodal information in the prior art is solved, and higher detection accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510445266.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-08-01
AI Technical Summary
The prior art uses only voice or text for suicide risk assessment, and fails to fully combine information from both, resulting in limited accuracy.
By extracting audio features and text features from user voice data, suicide risk detection is performed using preset classifiers, and multimodal detection is performed in combination with voice features, emotional features, language features and clinically related features.
It improves the accuracy and robustness of suicide risk detection, allows a more comprehensive understanding of the potential factors leading to suicide thoughts, and improves the accuracy of detection.
Smart Images

Figure CN120408264A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of suicide risk detection, and particularly to a multi-modal suicide risk detection method, an electronic device, a storage medium, and a program product. Background Art
[0002] Currently, there are multiple studies and products related to mental health monitoring based on speech and text, especially for depression, anxiety, and suicide risk assessment. However, the related technologies only use speech or text and fail to fully combine the information of both, resulting in limited accuracy. Summary of the Invention
[0003] Embodiments of the present application provide a multi-modal suicide risk detection method, an electronic device, a storage medium, and a program product, which are used to solve at least one of the above technical problems.
[0004] In a first aspect, an embodiment of the present application provides a multi-modal suicide risk detection method, including: Extracting audio features from user speech data; Converting the user speech data into corresponding user text data, and extracting text features from the user text data; Performing suicide risk detection using a preset classifier based on at least the audio features and the text features. Embodiments of the present application simultaneously consider two modalities of data, namely audio features and text features related to user speech data, to perform suicide risk detection on users, achieving multi-modal detection, thereby improving the accuracy and robustness of suicide risk detection.
[0005] In some embodiments, the audio features include speech features and emotional features; The performing suicide risk detection using a preset classifier based on at least the audio features and the text features includes: Performing suicide risk detection using a preset classifier based on at least the speech features, the emotional features, and the text features.
[0006] In some embodiments, the text features include: language features and clinically relevant features; The performing suicide risk detection using a preset classifier based on at least the speech features, the emotional features, and the text features includes: Performing suicide risk detection using a preset classifier based on at least the speech features, the emotional features, the language features, and the clinically relevant features.
[0007] In some embodiments, the preset classifier includes: a preset audio classifier and a preset text classifier; The suicide risk detection using the preset classifier at least based on the audio feature and the text feature includes: Using the preset audio classifier to detect the audio feature to obtain an audio detection result; Using the preset text classifier to detect the text feature to obtain a text detection result; Determining a suicide risk detection result at least based on the audio detection result and the text detection result.
[0008] In some embodiments, the preset audio classifier includes at least one of a first audio classifier and a second audio classifier; The first audio classifier is used to detect the speech feature to obtain a first audio detection result; The second audio classifier is used to detect the emotion feature to obtain a second audio detection result.
[0009] In some embodiments, the clinically relevant features include life event features and symptom features; the preset text classifier includes at least one of a first text classifier, a second text classifier, and a third text classifier; The first text classifier is used to detect the language feature to obtain a first text detection result; The second text classifier is used to detect the life event feature to obtain a second text detection result The third text classifier is used to detect the symptom feature to obtain a third text detection result.
[0010] In some embodiments, determining a suicide risk detection result at least based on the audio detection result and the text detection result includes: Determining a suicide risk detection result according to the first audio detection result, the second audio detection result, the first text detection result, the second text detection result, and the third text detection result.
[0011] In a second aspect, an embodiment of the present application provides a storage medium, in which one or more programs including execution instructions are stored, and the execution instructions can be read and executed by an electronic device (including but not limited to a computer, a server, or a network device, etc.) to execute any one of the above multi-modal suicide risk detection methods of the present application.
[0012] In a third aspect, an electronic device is provided, which includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute any one of the above multi-modal suicide risk detection methods of the present application.
[0013] In a fourth aspect, an embodiment of the present application further provides a computer program product, which includes a computer program stored on a storage medium. The computer program includes program instructions that, when executed by a computer, cause the computer to execute any one of the above-mentioned multi-modal suicide risk detection methods.
[0014] The embodiment of the present application simultaneously considers two modalities of data, namely audio features and text features, related to user voice data to detect the suicide risk of the user, achieving multi-modal detection, thereby improving the accuracy and robustness of suicide risk detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0016] Figure 1 It is a flowchart of an embodiment of the multi-modal suicide risk detection method of the present application; Figure 2 It is a flowchart of another embodiment of the multi-modal suicide risk detection method of the present application; Figure 3 It is a schematic structural diagram of an embodiment of the electronic device of the present application; Figure 4 It is a working flowchart of an embodiment of the expert consultation strategy of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present application. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.
[0018] It should also be noted that in this text, the terms "comprising" and "including" not only include those elements, but also other elements not explicitly listed, or elements inherent in such a process, method, article or device. Without further limitation, an element defined by the statement "comprising..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the said element.
[0019] As Figure 1 shown, an embodiment of the present application provides a multi-modal suicide risk detection method, including: S10. Extract audio features from user voice data.
[0020] Exemplarily, the user voice can include spontaneous voice or reading voice. For example, the spontaneous voice is the voice of the user speaking in daily life recorded, and the reading voice is the voice of the user reading a preset text recorded. In the present application, a pre-trained audio feature extraction model can be used to extract audio features from user voice data.
[0021] S20. Convert the user voice data into corresponding user text data, and extract text features from the user text data.
[0022] Exemplarily, the automatic speech recognition technology is used to automatically recognize and convert the user voice data into corresponding user text data. Then, a pre-trained text feature extraction model is used to extract text features from the user text data.
[0023] S30. Use a preset classifier to perform suicide risk detection at least based on the audio features and the text features. Exemplarily, a pre-trained preset classifier is used to implement suicide risk detection of the user according to data of two modalities, namely audio features and text features.
[0024] The embodiment of the present application simultaneously considers data of two modalities, namely audio features and text features, related to user voice data to perform suicide risk detection on the user, realizing multi-modal detection, thereby improving the accuracy and robustness of suicide risk detection.
[0025] In some embodiments, the audio features include voice features and emotion features. The corresponding pre-trained audio feature extraction models include the HuBERT model and Emotion2vec. Among them, the HuBERT model is used to extract voice features from user voice data, and Emotion2vec is used to extract emotion features from user voice data.
[0026] The suicide risk detection using a preset classifier is at least based on the audio features and the text features, including: using a preset classifier to perform suicide risk detection at least based on the speech features, the emotional features, and the text features.
[0027] Considering that the suicide risk of the subjects when answering questions may be related to their emotional state, in this embodiment, the speech features and the emotional features are further obtained from the user audio data through the HuBERT model and Emotion2vec respectively, and combined with the text features to jointly perform suicide risk detection, thereby improving the accuracy of suicide risk detection.
[0028] In some embodiments, the text features include: linguistic feature B and clinically relevant features (L and S); the corresponding pre-trained text feature extraction models include mental-BERT and RoBERTa-large, which are used to extract linguistic features and clinically relevant features from the user text data.
[0029] The suicide risk detection using a preset classifier is at least based on the speech features, the emotional features, and the text features, including: using a preset classifier to perform suicide risk detection at least based on the speech features, the emotional features, the linguistic features, and the clinically relevant features.
[0030] This application incorporates clinically relevant features and combines linguistic features for suicide risk detection, aiming to enable the model to more comprehensively understand the potential factors leading to suicidal thoughts. This approach allows us to capture the linguistic and clinical nuances that are crucial for accurate detection.
[0031] In some embodiments, the preset classifier includes: a preset audio classifier and a preset text classifier; wherein, the preset audio classifier is used to classify the audio features, and the preset text classifier is used to classify the text features.
[0032] As Figure 2 shown, it is a schematic flowchart of another embodiment of the multi-modal suicide risk detection method of this application. In this embodiment, the suicide risk detection using a preset classifier is at least based on the audio features and the text features, including: S31. Use the preset audio classifier to detect the audio features to obtain an audio detection result.
[0033] Exemplarily, the audio features are processed by a preset audio classifier to obtain an audio detection result. Exemplarily, the audio detection result includes the probability that the corresponding user has a suicide risk or whether there is a suicide risk.
[0034] S32. Use the preset text classifier to detect the text features to obtain a text detection result.
[0035] Exemplarily, a preset text classifier is used to process text features to obtain a text detection result. Exemplarily, the text detection includes the probability that a corresponding user has a suicide risk or whether there is a suicide risk.
[0036] S33. Determine a suicide risk detection result based on at least the audio detection result and the text detection result. Exemplarily, the audio detection result and the text detection result are comprehensively considered to determine the suicide risk detection result, so as to obtain whether the user has a suicide risk. For example, a soft voting and / or hard voting method is used to determine the suicide risk detection result according to the audio detection result and the text detection result.
[0037] In some embodiments, the preset audio classifier includes at least one of a first audio classifier A1 and a second audio classifier A2; the first audio classifier is used to detect the speech features to obtain a first audio detection result; the second audio classifier is used to detect the emotion features to obtain a second audio detection result. Exemplarily, the first audio classifier and the second audio classifier adopt a multi-layer perceptron (MLP) and / or a transformer.
[0038] In some embodiments, the clinically relevant features include life event features L and symptom features S; the preset text classifier includes at least one of a first text classifier T1, a second text classifier T2, and a third text classifier T3; the first text classifier is used to detect the language features to obtain a first text detection result; the second text classifier is used to detect the life event features to obtain a second text detection result; the third text classifier is used to detect the symptom features to obtain a third text detection result. Exemplarily, the first text classifier, the second text classifier, and the third text classifier adopt a multi-layer perceptron (MLP) and / or a transformer. The second text classifier can be implemented as a life event model, and the third text classifier can be implemented as a symptom classifier.
[0039] In the embodiments of the present application, we recognize that mental symptoms and life stress events are crucial for detecting suicidal intent. These features are clinically considered to be strong indicators of suicide risk. By incorporating these clinically relevant features, we aim to enable our model to more comprehensively understand the potential factors leading to suicidal thoughts. This approach allows us to capture the linguistic and clinical nuances that are crucial for accurate detection.
[0040] In some embodiments, the suicide risk detection result is determined at least based on the audio detection result and the text detection result, including: determining the suicide risk detection result based on the first audio detection result, the second audio detection result, the first text detection result, the second text detection result, and the third text detection result.
[0041] Exemplarily, the above first audio detection result r1, second audio detection result r2, first text detection result r3, second text detection result r4, and third text detection result r5 are comprehensively considered in a voting manner to determine the suicide risk detection result.
[0042] Suppose the detection results corresponding to five classifiers (R1, R2, A1, A2, A3) are as shown in Table 1 below:
[0043] Hard Voting Step description: Label conversion: Round the prediction probability of each model to a class label (0 or 1).
[0044] Classifier R1, R2, A2 → Label 1 Classifier A1, A3 → Label 0 Count the votes: Votes for Label 1: 3 votes (classifiers R1, R2, A2) Votes for Label 0: 2 votes (classifiers A1, A3) Majority vote decision: Select the label with the most votes as the final result.
[0045] Final result: Hard voting result = 1 (at risk) The above hard voting result can be used as the suicide risk detection result, which means that the current user is at risk of suicide.
[0046] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a combination of a series of actions. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application. In the above embodiments, the descriptions of each embodiment have their own focuses. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0047] In some embodiments, an embodiment of the present application provides a non - volatile computer - readable storage medium. One or more programs including execution instructions are stored in the storage medium, and the execution instructions can be read and executed by an electronic device (including but not limited to a computer, a server, or a network device, etc.) to execute any one of the above - mentioned multi - modal suicide risk detection methods of the present application.
[0048] In some embodiments, an embodiment of the present application further provides a computer program product. The computer program product includes a computer program stored on a non - volatile computer - readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer is enabled to execute any one of the above - mentioned multi - modal suicide risk detection methods.
[0049] In some embodiments, an embodiment of the present application further provides an electronic device, which includes: at least one processor, and a memory communicatively connected to the at least one processor. Wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the multi - modal suicide risk detection method.
[0050] Figure 3 FIG. is a schematic hardware structure diagram of an electronic device for executing the multi - modal suicide risk detection method provided by another embodiment of the present application. As Figure 3 shown, the device includes: One or more processors 310 and a memory 320. Figure 3 Here, one processor 310 is taken as an example.
[0051] The device for executing the multi - modal suicide risk detection method may further include: an input device 330 and an output device 340.
[0052] The processor 310, the memory 320, the input device 330, and the output device 340 may be connected through a bus or other means. Figure 3 Here, connection through a bus is taken as an example.
[0053] The memory 320, as a non - volatile computer - readable storage medium, can be used to store non - volatile software programs, non - volatile computer - executable programs, and modules, such as the program instructions / modules corresponding to the multi - modal suicide risk detection method in the embodiments of the present application. The processor 310 executes various functional applications and data processing of the server by running the non - volatile software programs, instructions, and modules stored in the memory 320, that is, implements the multi - modal suicide risk detection method in the above - mentioned method embodiments.
[0054] The memory 320 may include a program storage area and a data storage area. The program storage area may store an operating system and application programs required for at least one function. The data storage area may store data created according to the use of the multi-modal suicide risk detection device, etc. In addition, the memory 320 may include a high-speed random access memory and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 320 may optionally include a memory remotely provided with respect to the processor 310, and these remote memories may be connected to the multi-modal suicide risk detection device through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0055] The input device 330 may receive input digital or character information and generate signals related to user settings and function control of the multi-modal suicide risk detection device. The output device 340 may include a display device such as a display screen.
[0056] The one or more modules are stored in the memory 320 and, when executed by the one or more processors 310, perform the multi-modal suicide risk detection method in any of the above method embodiments.
[0057] The above product may execute the method provided in the embodiments of the present application and has function modules and beneficial effects corresponding to the execution of the method. For technical details not described in detail in this embodiment, reference may be made to the method provided in the embodiments of the present application.
[0058] The electronic device in the embodiments of the present application exists in various forms, including but not limited to: (1) Mobile communication devices: These devices are characterized by having mobile communication functions and mainly aim to provide voice and data communication. Such terminals include: smart phones (such as iPhone), multimedia phones, functional phones, and low-end phones, etc.
[0059] (2) Ultra-mobile personal computer devices: These devices belong to the category of personal computers, have computing and processing functions, and generally also have the characteristic of mobile Internet access. Such terminals include: PDAs, MIDs, and UMPC devices, etc., such as iPad.
[0060] (3) Portable entertainment devices: These devices can display and play multimedia content. Such devices include: audio and video players (such as iPod), handheld game consoles, e-books, and smart toys and portable vehicle navigation devices.
[0061] (4) Server: A device that provides computing services. The server consists of a processor, hard disk, memory, system bus, etc. The server is similar to a general computer architecture, but due to the need to provide highly reliable services, it has higher requirements in terms of processing power, stability, reliability, security, scalability, manageability, etc.
[0062] (5) Other electronic devices with data interaction functions.
[0063] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0064] To make the technical solutions and effects of this application clearer and more understandable, the research, exploration, implementation, and experimental demonstration processes of this invention are presented as follows: This application relates to a multi-modal suicide risk detection method and system that utilizes audio and text features from anonymous voice recordings. Given the data scarcity, this application explores various pre-trained and fine-tuned audio and text features, and combines eRisk data as well as clinically relevant features such as stressful life events and psychiatric symptoms. This application proposes an expert consultation process that includes three different systems: an election system, a panel discussion system, and a plenary discussion system. Each system adopts different classification and integration strategies. Finally, the classification results of all three designed systems are significantly higher than the baseline accuracy. By using the hard voting method on the development set to combine the classification results of all features, the best performance is achieved, with an accuracy of 0.79. However, the distribution difference between the test set and the development set poses a challenge to achieving consistent performance on both sets.
[0065] 1. 1. Introduction Preventing adolescent suicide is a crucial global mental health issue, which highlights the need for innovative strategies to identify high-risk populations early. Although traditional methods such as self-report questionnaires and clinical interviews are valuable, they are often limited by issues such as accessibility, stigma, and delayed intervention. Recent advances in computational speech analysis hold great promise for overcoming these limitations, as they can uncover subtle voice biomarkers related to psychological distress. This application aims to utilize a unique anonymous voice dataset from 600 Chinese adolescents (aged 10 - 18), half of whom are clinically determined to be at risk of suicide. By developing machine learning models to detect risk signals in spontaneous and read-aloud speech, we strive to provide a more effective early detection method for suicide prevention.
[0066] Since speech analysis can provide non-invasive, low-cost, and remote assessment, it has become a promising tool for detecting psychological distress and suicidal ideation. Several studies have shown that audio features such as prosody, voice source, formants, and spectrum are significantly correlated with suicidal ideation. In addition, the speech of those at risk of suicide tends to exhibit more disfluencies, including more frequent hesitations and speech errors, which further differentiates it from non-suicidal speech. From a semantic perspective, suicidal speech tends to involve a different set of primary words compared to non-suicidal speech.
[0067] The dataset of the first Speech Health Challenge consists of speech recordings of 600 adolescents aged 10 to 18 years old, with each participant providing three audio segments. The aim is to predict suicidal intent based on these recordings. The dataset is divided into a training set, a test set, and a development set, ensuring consistent gender ratios and age distributions across all subsets, while maintaining a balanced ratio of label 0 and label 1, i.e., 1:1. The training set includes 400 subjects, and the development set and test set each include 100 subjects. Each subject provides three audio segments corresponding to three different tasks: Task 0 - Emotional Regulation (ER), Task 1 - Passage Reading (PR), and Task 2 - Expression Description (ED). This structure naturally provides two modes of data: acoustic features and text descriptions. To make full use of this bimodal information, the challenge was initially divided into two parallel tracks: the audio track and the text track. In the ER and ED tasks, participants are required to answer specific questions. In the PR task, they are required to read a standardized poem, so the text descriptions are uniform for all participants in this task. Therefore, in the subsequent chapters discussing the models and results of the text mode, only the ER and ED tasks were considered. Against this background, our research questions are as follows: 1. Given the scarcity of data, which acoustic and text features are most effective for each individual task? Are there significant differences in the feature performance across different tasks? 2. Faced with three tasks and two modes, how can we effectively combine these models to create a robust and high-performance system that can integrate the most relevant information? Our contribution lies in addressing these research questions through a comprehensive approach. We explored various pre-trained and fine-tuned audio and text features and found that different tasks benefit from specific features. Notably, we incorporated semantic features extracted from task-related data as well as clinically significant features such as life stress events and psychiatric symptoms. Additionally, we proposed an expert consultation process for integrating different features and tasks. This process includes three different classification and integration methods, each designed to leverage the strengths of individual models and features. By effectively combining the most effective information from multiple sources, our method aims to improve the overall performance and robustness of the system. The accuracy of our system on the development set was 0.79, 0.76, and 0.74 respectively, while the accuracy of the baseline systems was 0.53 and 0.56 respectively.
[0068] 2. Methodology To address the current challenges, we explored integrating two modalities: audio and text, in three tasks (Task 0 - Emotion Regulation, Task 1 - Paragraph Reading, Task 2 - Expressive Description, corresponding to three audio segments). It is crucial to effectively utilize and combine these modalities and tasks. We proposed three different systems, each similar to an expert consultation process, as Figure 4 shown in the flowchart of the expert consultation strategy of this application: 1. Election System (System 1). We trained each modality (audio and text) independently for each task. This approach ensures that each modality is optimized for its specific task, enabling us to fully utilize the unique advantages of both modalities.
[0069] 2. Panel Discussion System (System 2). We trained confusion matrices for text and audio separately. Then these confusion matrices were combined through a voting mechanism. By this method, we can systematically integrate the results of both modalities, making full use of the unique insights provided by each modality.
[0070] 3. Plenary Discussion System (System 3). We integrated the results of both text and audio modalities into a unified confusion matrix. This comprehensive approach allows for a thorough evaluation, leading to a final classification result. This method aims to maximize the synergy between the two modalities.
[0071] 2.1. Using Data for Pre-training and Fine-tuning Features Given the limited available data for this challenge, we propose using pre-trained data and model fine-tuning.
[0072] 2.1.1. Audio Features In recent years, many models for speech feature extraction and representation have emerged. These models are all trained on large-scale datasets. Given that our dataset is relatively small, it is more reasonable to use a mature pre-trained model for feature extraction than to train a feature extraction model from scratch using the training set. Compared with wav2vec2 used in the reward baseline, we selected HuBERT and emotion2vec to achieve a generalized understanding and extraction of audio features.
[0073] HuBERT: For the HuBERT model, considering that the dataset for this challenge consists entirely of Chinese audio data, we adopted the HuBert base model for Chinese speech (chinese-hubert-base model). This speech processing model is based on the HuBERT model architecture and was pre-trained on the WenetSpeech L subset, which contains 10,000 hours of Chinese speech data. In addition, we also fine-tuned the HuBERT model using the dataset for this challenge.
[0074] Emotion2vec: Models trained on the IEMOCAP dataset can extract emotional features from speech data. Considering that the suicide risk of subjects when answering questions may be related to their emotional state, we also adopted the Emotion2vec base model (emotion2vec-base model) to extract features for training.
[0075] 2.1.2. Text Features Similarly, for the text pattern, more relevant information must be sought from other fields to improve the performance and robustness of our model. We first identified the eRisk 2021 Task 2: Self-harm dataset as a valuable resource. This dataset has previously been used for suicide detection on social media and is highly relevant to our task. By leveraging the eRisk dataset, we can supplement our limited dataset with other examples with similar features and challenges. This not only increases the amount of training data but also introduces a wider range of language patterns (referring to some regularities, patterns, or structures in language, including grammar, vocabulary, pronunciation, etc.) and contexts related to suicidal ideation, thus improving the generality of our model.
[0076] Secondly, we recognize that psychiatric symptoms and life stress events are crucial for detecting suicidal intent. These features are clinically considered strong indicators of suicide risk. By incorporating these clinically relevant features, we aim to enable our model to have a more comprehensive understanding of the potential factors leading to suicidal thoughts. This approach allows us to capture the linguistic and clinical nuances that are essential for accurate detection.
[0077] Psychiatric symptom model: The DSM-5 is a symptom-based diagnostic criterion commonly used in clinical practice. In related technologies, based on the DSM-5, 38 common symptoms of mental disorders were defined, and a high-quality symptom dataset was labeled. Using this dataset, they trained a symptom classifier capable of extracting symptom features from text. We used their dataset to train the symptom classifier in the same way, extracting symptom features from the subjects in Task 0 and Task 2 for classification. Similar to the life event model, we used this probability as a feature in subsequent classification.
[0078] Life event model: Importantly, life events are also important risk factors for suicidal ideation and behavior. Related technologies show that stressful life events increase the suicide risk by 37% - 45%, with men and young people being more affected. Our definition of life events is based on previous research, which proposed 121 life events in 12 categories. We modified its classification by adding new events and reorganizing overlapping content, finally obtaining 12 categories containing 127 life events. We recruited annotators and adopted strict training methods and a series of quality control protocols to finally obtain a high-quality life event dataset. We used this dataset to train the life event classification model. After processing using this classification model, we were able to obtain the probability of life events occurring in the text of each subject. We used this probability as the life event probability feature for each subject.
[0079] 2.2. Classification and integration strategies To classify suicide risk, we explored the following strategies: - Independent training and voting: Each mode (text and audio) is independently trained for each task. The results are combined by voting to determine the final classification.
[0080] - Independent confusion matrices and voting: Confusion matrices are trained separately for text and audio results. These matrices are combined by voting to obtain the final classification result.
[0081] - Unified confusion matrix and voting: Text and audio results are combined into a unified confusion matrix. Then this confusion matrix is used to train the final classification model, which is determined by voting.
[0082] 3. Experiments 3.1. Data preprocessing In the ER and ED tasks, the subjects have to answer specific questions. However, in the PR task, the subjects have to read a standard poem, so the text descriptions of different subjects are unified. Therefore, we mainly conduct text analysis for the ER and ED tasks. We use the large V3 Whisper model for automatic speech recognition (ASR) to convert the WAV files into text. Translating the text into English can improve the content, expand the selection range of models for text processing and feature extraction, and utilize a larger dataset for fine-tuning, which is crucial for mental health risk assessment.
[0083] 3.2. Audio Feature Extraction Similar to the baseline method, we adopt the sliding window method with a window length of 10 seconds and a step size of 6 seconds to divide the entire audio file into multiple segments. Features are extracted from each segment for training. Finally, the predicted label of the audio is determined by performing hard voting and soft voting on the results of multiple segments.
[0084] Data Augmentation: We use two augmentation methods in audio feature extraction: Mixup creates new data by randomly mixing the features and labels of two samples, which can enhance the generality of the model. Speed perturbation is a time warping technique that creates new data by changing the speed of the original audio. This method has been verified in detecting speech-related suicide risks.
[0085] HuBERT Model Fine-tuning: According to the structure of the baseline model, we use HuBERT and Emotion2vec instead of wav2vec2 as the encoder, and use two fully connected layers as the classifier to extract audio features. During fine-tuning, the training is combined. The learning rates for tasks 0 and 2 are 2e-5 and 5e-5 respectively. In task 2, a sliding window algorithm with a window length of 30 seconds and a step size of 18 seconds is used to clip the audio, and the batch size is changed to 4.
[0086] 3.3. Text Feature Extraction Data Screening: In addition to the dataset provided by the challenge, we also used the data in eRisk2021 Task 2: Self-Harm to fine-tune the model. The eRisk2021 t2 dataset contains 1,421 users, among which 152 are positive samples and 1,269 are negative samples. Considering the unbalanced ratio of positive and negative samples, we designed a balanced sampler to ensure that the number of positive and negative samples added to the model for fine-tuning is always the same. In addition, to improve the quality of the data used for fine-tuning, we filtered the eRisk data using the data provided by the challenge. To filter the eRisk data, we first extracted embeddings from the eRisk data using sentence-BERT, and then filtered the data according to different cosine similarity thresholds (exemplarily, calculating the cosine similarity between the training set data and the fine-tuning data and comparing it with the cosine similarity threshold to filter the data). In the experiment, we selected three thresholds of 0.5, 0.6, and 0.7 and tried different proportions of eRisk data. Too low a threshold may result in the selection of poor-quality data, which is not very helpful for fine-tuning. On the contrary, too high a threshold may lead to insufficient data being selected.
[0087] BERT Model Fine-Tuning: Similar to the baseline, we adopted a fine-tuning model for text features. We fine-tuned the model using the original data (i.e., the dataset of the First Voice Health Challenge) and the filtered eRisk data. We tried several BERT-based fine-tuning models, including BERT (base and large), mental-BERT, and RoBERTa-large. Finally, we selected the two best-performing models: mental-BERT and RoBERTa-large. The learning rates used for Task 0 and Task 2 are 2e-5 and 5e-5 respectively.
[0088] 3.4. Classification After feature extraction, classification is required. Due to different feature dimensions, we adopted two classifiers: multi-layer perceptron (MLP) and transformer. The former is suitable for fine-tuning models and classifying low-dimensional features, while the latter is more suitable for continuous data such as text and audio features. Exemplarily, in all systems, the separately trained audio model (H E) uses MLP, and the model trained by combining text models and text audio features uses transformer.
[0089] MLP Classifier: We used an MLP consisting of two fully connected layers, with the learning rate adjusted using CosineAnnealingLR, starting from an initial value of 5e-5. The input features were fed into a hidden layer with a dimension of 128, and a dropout rate of 0.1 was applied. During training, we set the batch size to 8 and trained for at least 200 epochs, then selected the model with the highest accuracy on the validation set for the final prediction.
[0090] Transformer Classifier: We adopted a transformer-based classifier with a four-layer network structure and eight attention heads. The hidden dimension of this model was set to 128, and to prevent overfitting, the dropout rate was 0.25. To balance the limited amount of training data and the high-dimensional features of each sample, a linear layer for dimensionality reduction was added to the model (exemplarily, this linear layer was added before the transformer classifier for dimensionality reduction).
[0091] Considering multi-modal data and multiple tasks, it is crucial to effectively integrate the results of different modalities. We explored two methods: The first method is to train the model according to the characteristics of a single task and modality, and then aggregate the results using hard voting or soft voting. Hard voting selects the majority prediction, while soft voting averages the probabilities. The probabilities that classifiers of Model 1 to Model n label a sample as Label 1 are respectively defined as p 1,... p n .
[0092] Soft Voting:
[0093] Hard Voting:
[0094] Among them, for the formula of soft voting, the explanation is as follows: Input: The predicted probabilities that each model (from Model 1 to Model n) predicts a certain sample belongs to Label 1 p 1, p 2, …, p n .
[0095] Calculation: For p 1 to p n Take the arithmetic mean to obtain the average probability p s .
[0096] Through round( p s) rounds the average probability and converts it to the final class label (0 or 1). For example, if p s ≥0.5, then round( p s ) = 1 (classified as label 1); otherwise, 0.
[0097] Essence: Soft voting reflects the "collective confidence" of the model through the average value of the probability, and then determines the category based on the threshold (0.5).
[0098] Example: The predicted probabilities for the three models are p 1=0.7, p 2=0.6, p 3=0.4, then:
[0099] The final classification is label 1.
[0100] The formula for hard voting is as follows: Input: The predicted probability of each model for a sample p 1, p 2,…, p n .
[0101] calculate: First, the probability of each model p i Round to the class label (0 or 1), i.e. round( p i ).
[0102] Count the total number of times all models voted 1 and calculate their proportion p h .
[0103] The last pair p h Round off again to determine the final category. For example, if more than half of the models vote for 1 (i.e. p h ≥0.5), it is classified as label 1.
[0104] Essence: Hard voting follows the principle of "the minority obeys the majority", ignores the specific value of the probability, and only focuses on the statistical results of the category label.
[0105] Example: The predicted probabilities for the three models are p 1=0.7, p 2=0.6, p 3=0.4, then:
[0106] Finally classified as label 1.
[0107] This method can improve the overall performance, but the model selected based on the development set accuracy may perform poorly on the test set, thus reducing the reliability. The second method is to concatenate the features extracted by the fine-tuned model and then input them into the classifier for the final prediction.
[0108] 3.5. Evaluation Metrics Consistent with the baseline, we use "Accuracy" to compare the models and evaluation results. In addition, for the voting results obtained from the three systems, we incorporated the F1 score, precision, and recall for further comparison.
[0109] 4. Results To systematically evaluate the effectiveness of various acoustic and text features, we conducted a comprehensive comparison of different feature sets and classification strategies.
[0110] 4.1. Feature Comparison Table 2 lists the results of all models. We found that the prediction performance of the audio and text modes is comparable in single-task scenarios. In addition, applying data augmentation techniques in HuBERT can improve the performance of Task 1 and Task 2. Similarly, pre-training the BERT model with eRisk data also improves the performance of the text mode.
[0111] Table 2: Model accuracies of System 1 on Tasks 0, 1, and 2 in the development set
[0112] 4.2. Classification and Integration Paradigms For System 1, we selected the best combination of models in Table 2 (i.e., Hubert - enhanced version, Emotion2Vec classifier, BERT classifier using eRisk, life event classifier, and symptom classifier). On the development set, through hard voting, we obtained a model with an accuracy of 0.79, which is significantly better than any single model and the baseline model. Further, the above best combination of models was evaluated on the test set, and an accuracy of 0.56 was obtained. This indicates that the results may be overfitted to the development set and cannot well generalize the data distribution of the test set. In System 2, the results of the audio group and the text group are shown in the first two rows of Table 3. The results of System 3 (i.e., all combinations) are listed in the third row of Table 3. These results are the best accuracies obtained by the MLP classifier on the development set.
[0113] Table 3: Model precisions on the development set when System 2 and 3 adopt different concatenation methods
[0114] 4.3. Different Voting Strategies We tried two methods, soft voting and hard voting, and found that the effect of hard voting is significantly better than that of soft voting. This indicates that hard voting shows stronger robustness, and also indicates that different models often produce low-quality predictions for certain samples, especially high-confidence incorrect predictions. The results are shown in Table 4, and all the results obtained by our system are better than the two baseline results.
[0115] Table 4: Performance Comparison of Models on the Development Set (HV = Hard Voting; SV = Soft Voting)
[0116] Exemplarily, the models adopted in System 2 and System 3 in Table 4 above are the best combinations of the models in Table 2, and more than one model is used for each modality in System 1. Table 4 is a vote for three tasks because each subject has speech containing 3 tasks, and the ultimate goal is to correctly predict the label of the subject.
[0117] The results of this application are based on the scoring framework of the MINI-KID scale, which assesses the current suicide risk as either at risk or not at risk.
[0118] 5. Conclusion This application uses the dataset of The 1st SpeechWellness Challenge to perform the suicide risk prediction task. We explored the audio and text modalities, adopted techniques such as model fine-tuning and combining eRisk data. We used multiple models for feature extraction and designed three different methods to combine model features and voting strategies. Finally, the accuracy of each system on the development set exceeded that of the two baseline systems.
[0119] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solutions in essence or the part that contributes to the related technologies can be embodied in the form of a software product, and this computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0120] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A multimodal suicide risk detection method, characterized in that, Comprising: Extracting audio features from user voice data; Converting the user voice data into corresponding user text data, and extracting text features from the user text data; Performing suicide risk detection using a preset classifier based at least on the audio features and the text features.
2. The method according to claim 1, wherein The audio features include voice features and emotional features; The performing suicide risk detection using a preset classifier based at least on the audio features and the text features includes: Performing suicide risk detection using a preset classifier based at least on the voice features, the emotional features, and the text features.
3. The method according to claim 2, wherein The text features include: language features and clinically relevant features; The performing suicide risk detection using a preset classifier based at least on the voice features, the emotional features, and the text features includes: Performing suicide risk detection using a preset classifier based at least on the voice features, the emotional features, the language features, and the clinically relevant features.
4. The method according to claim 3, characterized in that, The preset classifier includes: a preset audio classifier and a preset text classifier; The performing suicide risk detection using a preset classifier based at least on the audio features and the text features includes: Detecting the audio features using the preset audio classifier to obtain an audio detection result; Detecting the text features using the preset text classifier to obtain a text detection result; Determining a suicide risk detection result based at least on the audio detection result and the text detection result.
5. The method according to claim 4, characterized in that The preset audio classifier includes at least one of a first audio classifier and a second audio classifier; The first audio classifier is used to detect the voice features to obtain a first audio detection result; The second audio classifier is used to detect the emotional features to obtain a second audio detection result.
6. The method according to claim 5, wherein The clinically relevant features include life event features and symptom features; The preset text classifier includes at least one of a first text classifier, a second text classifier, and a third text classifier; The first text classifier is used to detect the language features to obtain a first text detection result; The second text classifier is used to detect the life event features to obtain a second text detection result The third text classifier is used to detect the symptom features to obtain a third text detection result.
7. The method according to claim 6, characterized in that The determining a suicide risk detection result based at least on the audio detection result and the text detection result includes: Determining a suicide risk detection result based on the first audio detection result, the second audio detection result, the first text detection result, the second text detection result, and the third text detection result.
8. An electronic device, comprising: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the steps of the method according to any one of claims 1-7.
9. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1-7 are implemented.
10. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1-7 are implemented.