Interview voice mental health detection early warning method and device based on artificial intelligence

By extracting speech paralinguistic features from interview audio and utilizing a neural network model, this method solves the problem of automatically extracting mental health features from natural dialogue in existing technologies, achieving efficient mental health detection and early warning, and is applicable to various resource-limited or large-scale screening scenarios.

CN121884869APending Publication Date: 2026-04-17北京市中医药研究所 +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
北京市中医药研究所
Filing Date
2026-01-28
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing technologies struggle to reliably extract relevant features related to mental health status from natural dialogue in real-world scenarios, resulting in low efficiency of mental health testing, reliance on professionals, and difficulty in large-scale deployment.

Method used

By acquiring audio from two-person interviews, extracting statistical measures of speech paralinguistic features, constructing a feature set, and using a neural network model for mental health detection, automated early warning of psychological risks can be achieved.

Benefits of technology

It enables rapid and effective mental health testing in real-world scenarios, reduces reliance on professional interviewers, is suitable for resource-limited or large-scale screening scenarios, and improves the objectivity and efficiency of the testing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121884869A_ABST
    Figure CN121884869A_ABST
Patent Text Reader

Abstract

The invention discloses an interview voice mental health detection early warning method and device based on artificial intelligence, and relates to the technical field of voice signal processing. The method comprises the following steps: acquiring interview audio data between a subject and an interviewer; preprocessing the interview audio data to separate out continuous speech segments uttered by the subject; and extracting statistics of the voice side language features, inputting feature vectors constructed based on the voice feature set into a pre-trained psychological health detection model, and obtaining a probability value that the subject is in a psychological risk state. The technical problem that in the prior art, voice input is usually required to be an isolated statement with clear content and clear pronunciation, and the complexity of natural dialogues in a real scene is not fully considered is solved, psychological health detection and early warning based on the natural dialogues in the real scene are achieved, dependence on professional interviews is reduced, and the psychological health early warning effect is improved. The method is suitable for resource-limited or large-scale screening scenes, and mental health detection can be quickly and effectively carried out.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech signal processing technology, and in particular to an artificial intelligence-based method and device for detecting and warning of mental health through interview speech. Background Technology

[0002] Traditional mental health assessments primarily rely on scale screening and clinical interviews. Scale screening is easily influenced by the subject's subjective will and social desirability, resulting in unstable reliability and validity; while clinical interviews heavily depend on professionals, are costly and inefficient, and are difficult to implement on a large scale.

[0003] With the development of speech analysis technology, methods for emotion recognition and psychological state assessment based on speech features have emerged. However, existing technologies mostly focus on emotion recognition for read-aloud text or short voice commands, such as extracting acoustic features like fundamental frequency, energy, and Mel-frequency cepstral coefficients, and then using machine learning models for classification. These methods typically assume that the speech input is a single person with clear content and articulated pronunciation, and their technical framework does not fully consider the complexity of natural dialogue in real-world scenarios.

[0004] Specifically, existing technologies struggle to reliably extract effective features strongly correlated with mental health from real interview audio containing multi-person dialogues, overlapping speech, background noise, non-standard pronunciation, and long periods of continuous speech. Therefore, current technologies lack methods for mental health detection and early warning based on natural dialogue in real-world interview scenarios. Summary of the Invention

[0005] The purpose of this application is to provide an AI-based method and device for detecting and warning about mental health through interview audio. By acquiring audio from a two-person interview, statistical measures of paralinguistic features are extracted to construct a feature set, and then a model outputs the probability of psychological risk. This solves the technical problem of existing technologies that rely on standard spoken language and fail to consider the complexity of natural dialogue in real-world scenarios. This method can achieve mental health detection and warning based on natural dialogue in real interview scenarios, reducing reliance on professional interviewers. It is suitable for resource-limited or large-scale screening scenarios and can quickly and effectively detect mental health issues.

[0006] To achieve the above objectives, this application provides the following solution: In a first aspect, this application provides an artificial intelligence-based method for detecting and warning of mental health through interview audio, comprising: acquiring interview audio data between a subject and an interviewer; preprocessing the interview audio data to separate continuous speech segments uttered by the subject; extracting statistical measures of paralinguistic features from the continuous speech segments uttered by the subject to construct a speech feature set for mental health assessment; inputting a feature vector constructed based on the speech feature set into a pre-trained mental health detection model, wherein the mental health detection model is a neural network model; and obtaining a probability value of the subject being in a state of psychological risk based on the output of the mental health detection model.

[0007] Optionally, the preprocessing of the interview audio data includes: performing speech activity detection on the interview audio data to identify mixed human voice segments; and performing speaker differentiation on the mixed human voice segments to filter out the interviewer's speech and obtain continuous speech segments spoken by the subject.

[0008] Optionally, the step of extracting the statistical measures of speech sub-language features includes: performing frame segmentation and windowing processing on the continuous speech segments spoken by the subject to obtain multiple short-time speech frames; calculating the frame-level feature value of at least one type of speech sub-language feature in each short-time speech frame to obtain the frame-level feature sequence of each type of speech sub-language feature; calculating the descriptive statistics of the frame-level feature sequence and using it as the statistical measures of the speech sub-language features.

[0009] Optionally, the categories of the speech sub-language features include at least one of the following: power spectral density, zero-crossing rate, zero-crossing rate based on power spectral density Z-score, Mel frequency cepstral coefficient, fundamental frequency, formant frequency, spectral centroid, spectral flatness, and energy percentage of a specific frequency band.

[0010] Optionally, the statistics include mean, standard deviation, median, maximum value, and minimum value.

[0011] Optionally, the feature vector constructed based on the speech feature set includes: concatenating the statistics corresponding to all categories of speech sub-language features to form the feature vector.

[0012] Optionally, the mental health detection model is trained based on multiple interview audio data that have been labeled with mental health status tags; wherein each data sample used to train the mental health detection model is a feature vector extracted from the interview audio data labeled with mental health status tags and constructed from statistics of speech paralinguistic features.

[0013] Secondly, this application provides an artificial intelligence-based interview voice psychological health detection and early warning device, comprising: an acquisition module configured to acquire interview audio data between a subject and an interviewer; a preprocessing module configured to preprocess the interview audio data to separate continuous speech segments uttered by the subject; a statistics extraction module configured to extract statistics of speech paralinguistic features from the continuous speech segments uttered by the subject to construct a speech feature set for psychological health assessment; a model prediction module configured to input a feature vector constructed based on the speech feature set into a pre-trained psychological health detection model, wherein the psychological health detection model is a neural network model; and a probability value output module configured to obtain a probability value of the subject being in a psychological risk state based on the output of the psychological health detection model.

[0014] Thirdly, this application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the artificial intelligence-based interview voice psychological health detection and early warning method described in any one of the above.

[0015] Fourthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the artificial intelligence-based interview voice psychological health detection and early warning method described above.

[0016] According to the specific embodiments provided in this application, the following technical effects are disclosed: This application provides an AI-based method and device for detecting and warning about mental health through interview audio data. By preprocessing the interview audio data, continuous speech segments spoken by the subject are separated, accurately extracting clean continuous speech segments from the original recording. This provides a reliable data foundation for subsequent analysis and solves the problem of feature contamination caused by mixed speech. By extracting statistical quantities of paralinguistic features to construct a speech feature set, a quantitative representation of the subject's overall speech pattern is achieved. This constructs a feature set that is independent of semantic content and can stably reflect the psychological state, solving the problem of extracting effective mental health-related features from free interview audio. By inputting the constructed feature vector into a pre-trained mental health detection model, the probability value of the psychological risk state is output, achieving automated probability prediction of mental health risk and improving the objectivity and efficiency of detection. Therefore, the AI-based interview audio mental health detection and warning method in this embodiment can complete mental health detection and warning based on interview audio data, and is suitable for resource-limited or large-scale screening scenarios, achieving rapid and effective mental health detection. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is an application environment diagram of an artificial intelligence-based interview voice psychological health detection and early warning method according to an embodiment of this application; Figure 2 A flowchart illustrating an artificial intelligence-based interview voice-based mental health detection and early warning method provided in an embodiment of this application; Figure 3 for Figure 2 A detailed flowchart illustrating the preprocessing steps for interview audio data. Figure 4 for Figure 2 A detailed flowchart illustrating the steps for extracting statistical measures of speech paralinguistic features. Figure 5 A schematic diagram of the functional modules of an artificial intelligence-based interview voice psychological health detection and early warning device provided in an embodiment of this application; Figure 6 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0019] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0020] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0021] The AI-based interview voice psychological health detection and early warning method provided in this application can be applied to, for example... Figure 1In the application environment shown, terminal 101 communicates with server 102 via a network. A data storage system can store the data that server 102 needs to process. The data storage system can be set up independently, integrated into server 102, or placed in the cloud or on another server. Terminal 101 can send the interview audio data between the subject and the interviewer to server 102. Server 102 receives the interview audio data, preprocesses it to separate continuous speech segments spoken by the subject, extracts statistical measures of paralinguistic features from these segments to form a speech feature set for mental health assessment, and inputs the feature vector constructed based on the speech feature set into a pre-trained mental health detection model. Based on the output of the mental health detection model, the probability value of the subject being in a psychological risk state is obtained. Server 102 can then feed back the obtained probability value of the subject being in a psychological risk state to terminal 101. In addition, in some embodiments, the probability value of the subject being in a state of psychological risk can also be implemented by the server 102 or the terminal 101 alone. For example, the terminal 101 can directly process the interview audio data between the subject and the interviewer, or the server 102 can obtain the interview audio data between the subject and the interviewer from the data storage system.

[0022] The terminal 101 can be, but is not limited to, various desktop computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. The server 102 can be a standalone server, a server cluster, or a cloud server.

[0023] In one exemplary embodiment, such as Figure 2 As shown, an artificial intelligence-based method for detecting and warning about mental health through interview voice recordings is provided. This method is executed by a computer device, specifically a terminal or server, or both. In this embodiment, the method is applied to... Figure 1 We will use server 102 as an example to illustrate this.

[0024] like Figure 2 As shown, the AI-based interview voice-based mental health detection and early warning method includes the following steps: S11, Obtain audio data of the interview between the subject and the interviewer; S12, preprocess the interview audio data to separate continuous speech segments spoken by the subject; S13, extract statistics of speech paralinguistic features from continuous speech segments spoken by the subject to construct a speech feature set for mental health assessment; S14, input the feature vector constructed based on the speech feature set into the pre-trained mental health detection model; S15. Based on the output of the mental health testing model, obtain the probability value of the subject being in a state of psychological risk.

[0025] Among them, the mental health detection model is a neural network model, specifically including a one-dimensional convolutional layer and a fully connected layer.

[0026] Through the above steps, this method can accurately separate the subject's speech from the original interview recording, extract paralinguistic features related to psychological state, and realize probabilistic prediction of psychological risk through a pre-trained neural network model. This enables efficient and objective psychological health detection in real dialogue scenarios, effectively solving the technical problem of automatically extracting effective psychological state features from natural dialogue and conducting assessments.

[0027] In practice, step 11 involves acquiring audio data from the interview between the subject and the interviewer. Specifically, the audio data is collected from structured or semi-structured interviews conducted with the subject for mental health assessment. The interviews are performed by trained personnel or interactive systems that follow pre-defined logic, guiding the subject to deliver a continuous and natural narrative. This interview method can effectively stimulate the subject to generate speech data within a controlled framework. The speech data is of sufficient duration and rich in paralinguistic features, such as intonation, rhythm, and pauses.

[0028] Interview questions are usually pre-designed to elicit specific ranges of thought and emotional responses from the participants, guiding them to engage in a continuous, natural verbal expression rather than a simple "yes / no" answer. Examples include: emotional state: "Please describe how you've felt in the past week"; stressful events: "Talk about the most stressful thing that happened to you recently"; free narration: "Please tell a story about an experience that has had a significant impact on your life."

[0029] The interview can be conducted in a relatively quiet and controlled environment, such as a psychological counseling room, a soundproof laboratory, or a room with very low background noise, in order to minimize the contamination of speech features by environmental noise.

[0030] The interview equipment can use high-quality audio recording devices, such as professional microphones, voice recorders, or high-quality microphone arrays integrated into mobile devices or computers. The audio recording equipment should ensure that the frequency response range covers the main frequency bands of human voice, typically 80Hz-8000Hz; the sampling rate can be set to 16kHz, which is sufficient to clearly capture the main features of human voice with less computation; the bit depth can be set to 16bit to provide sufficient dynamic range.

[0031] Furthermore, after acquiring the interview audio data in step 11 and before preprocessing the interview audio data in step 12, the interview audio data can be de-identified. All data processing in this method complies with the "Personal Information Protection Law of the People's Republic of China" and related ethical review requirements. Specifically, all personally identifiable information, such as name, ID number, and phone number, is removed from the audio file name, and only a randomly generated unique ID is used to associate it with the subject's other data, such as psychological assessment scale results. The anonymized interview audio data is stored in encrypted form on a secure server or designated storage system with access control. Data access is limited to authorized research or operational personnel, and access logs are recorded to prevent data leakage and unauthorized use.

[0032] In specific implementation, such as Figure 3 As shown, the specific steps for preprocessing the interview audio data in step 12 above include: S121, Perform speech activity detection on the interview audio data and identify mixed human voice segments.

[0033] Specifically, a speech activity detection algorithm based on energy and spectral features can be employed. This algorithm analyzes the audio signal in short time frames, calculating the signal energy and zero-crossing rate of each frame and comparing them with an adaptive threshold. This automatically identifies and locates all "speech segments" containing human speech in the recording, distinguishing them from "non-speech segments" such as silent segments, background noise, and excessively short pauses. Once identified, the "non-speech segments" are removed.

[0034] S122, Speaker differentiation is performed on the mixed human voice segments to filter out the interviewer's voice and obtain the continuous speech segments spoken by the subject.

[0035] Specifically, voiceprint recognition, channel separation, or a combination of both can be used to further distinguish speakers after identifying "speech segments" in order to separate one or more continuous speech segments that contain only the subject's voice.

[0036] By performing steps 121 to 122 above, one or more continuous speech segments containing only the subject's voice can be continuously output from the original interview audio data. These segments constitute the direct data source for subsequent speech feature analysis.

[0037] In specific implementation, after separating the continuous speech segments of the subject in step 12 and before feature extraction in step 13, an optional signal enhancement step may be included: pre-emphasizing the continuous speech segments produced by the subject. Specifically, the pre-emphasization process involves filtering the speech signal using a first-order finite impulse response high-pass filter. The transfer function of this filter is typically: ,in, This is the pre-emphasis coefficient, typically ranging from 0.9 to 0.97. Pre-emphasis processing can compensate for the attenuation of high-frequency components in the speech signal during propagation, enhance the prominence of high-frequency features (such as formants), and make the signal spectrum flatter, thus facilitating subsequent spectrum analysis.

[0038] In specific implementation, such as Figure 4 As shown, step 13 above, the step of extracting statistical measures of speech paralinguistic features, includes: S131, the continuous speech segments spoken by the subject are framed and windowed to obtain multiple short speech frames.

[0039] Specifically, because speech segments exhibit short-term stationarity, their characteristics typically remain largely unchanged within a short timeframe of 10-30 milliseconds. Therefore, segmenting the continuous speech segments produced by the subject into short-time speech frames is a standard practice, for example, with a frame length of 20-30 milliseconds and a frame shift of 10 milliseconds. Subsequently, each frame is windowed, i.e., multiplied by a window function, such as a Hamming window, to reduce spectral leakage caused by signal truncation and to ensure smooth transitions at frame edges.

[0040] Furthermore, the amplitude of all frames within each speech segment can be Z-score normalized by subtracting the signal mean of the speech segment and then dividing by the standard deviation, as shown in the formula: , The signal is a segment of speech. The signal mean of the speech segment. The standard deviation is denoted by Z-score. By standardizing the amplitude using Z-score, the overall volume difference caused by factors such as recording equipment gain and the speaker's distance from the microphone can be eliminated, ensuring that the signal energy between different samples is within a comparable range.

[0041] S132, calculate the frame-level feature value of at least one type of speech sub-language feature in each short-time speech frame to obtain the frame-level feature sequence of each type of speech sub-language feature.

[0042] Specifically, paralinguistic features refer to acoustic features independent of semantic content, aiming to quantify the prosody, timbre, and spectral characteristics of speech. For each type of speech paralinguistic feature, calculations are performed by traversing all short-time speech frames, resulting in a frame-level feature sequence arranged in chronological order. The categories of speech paralinguistic features include: power spectral density, zero-crossing rate, Mel-frequency cepstral coefficients, fundamental frequency, formant frequencies, spectral centroid, spectral flatness, and energy percentage of a specific frequency band.

[0043] Power spectral density (PSD) characterizes the distribution of signal power in the frequency domain. The characteristic calculation method for PSD is as follows: The periodogram method is used, which involves performing a Fast Fourier Transform (FFT) on each frame of the signal and then taking the square of its magnitude to obtain the PSD estimate vector. For a smoother estimate, Welch's method can also be used to calculate the PSD estimate vector by performing periodogram calculations on overlapping segments and then averaging the results. After obtaining the PSD estimate vector for each frame of the signal, higher-order statistics of the PSD estimate vector, such as skewness and kurtosis, are further calculated.

[0044] Zero-crossing rate (ZCR) characterizes the number of times an audio signal crosses zero points per frame and is highly sensitive to background noise. The formula for calculating the zero-crossing rate is: In the formula, This is an indicator function that takes the value 1 when the condition inside the parentheses is true, and 0 otherwise; This indicates the total number of sampling points in the signal frame; Indicates the first The amplitude of each sampling point.

[0045] The zero-crossing rate based on the power spectral density Z-fraction (PSD-ZCR) is a composite feature used to reflect the volatility or instability of the spectrum. Its calculation involves two consecutive steps: calculating... The vector is then standardized using the Z-score formula. In the formula, Represents the original power spectral density vector , express The arithmetic mean of vectors; express The standard deviation of a vector; This represents the new vector obtained after standardization.

[0046] Mel-frequency cepstral coefficients (MFCCs) are features based on the human auditory mechanism and are one of the most crucial features in speech recognition and analysis, effectively representing the spectral envelope of speech. The extraction process for Mel-frequency cepstral coefficients includes: performing a Fast Fourier Transform (FFT) on each frame of signal to obtain the spectrum; squaring the spectral amplitude to obtain the power spectrum; passing the power spectrum through a set of Mel-scale triangular filters; taking the logarithm of the energy output from each filter; performing a Discrete Cosine Transform (DCT) on the resulting logarithmic energy sequence; and retaining the first 12-13 coefficients as MFCC features.

[0047] The fundamental frequency, also known as the dominant frequency, is the basic frequency of vocal cord vibration and is directly related to the pitch of speech. It can be estimated using autocorrelation methods or peak detection algorithms based on computational auditory scenarios. These algorithms estimate the fundamental frequency by searching for periodic peaks in the time domain (autocorrelation) or the frequency domain (harmonic structure).

[0048] The formant frequencies were estimated using the Linear Predictive Coding (LPC) analysis method. This involved performing LPC analysis on the speech frame to obtain a full-pole model representing the vocal tract, and then solving for the roots of this LPC polynomial to estimate the frequencies of the first, second, and third formants. The formants represent the energy concentration bands in the speech spectrum, and the frequencies of the first, second, and third formants are related to the shape of the vocal tract and reflect the vowel timbre and articulation.

[0049] The centroid of the spectrum, as the average frequency, characterizes the "center of gravity" of the spectral energy and is an indicator of the brightness of a sound. Its calculation formula is: In the formula, Indicates frequency index; This indicates that after the Fast Fourier Transform, at the frequency index... The corresponding power spectrum value; molecule The denominator represents the weighted sum of power with frequency index as the weight; This indicates the total power of the signal.

[0050] Spectral flatness characterizes whether a spectrum resembles noise (flat) or contains significant harmonic or tonal components. For a given frame of signal, spectral flatness is obtained by calculating the ratio of its geometric mean to its arithmetic mean, using the following formula: In the formula, This is expressed as spectral flatness; Indicates the first Power spectral values ​​at each frequency index; The numerator represents the total number of frequency points involved in the summation; Represented as the geometric mean of the power spectral values; denominator It is expressed as the arithmetic mean of the power spectrum values.

[0051] The energy percentage of a specific frequency band is characterized as the energy distribution of the spectrum. It is calculated as follows: calculate the total energy of the signal in the frame; find the index position of the corresponding frequency according to the frequency resolution of FFT, such as 200Hz, 500Hz, 700Hz, 1000Hz, 2000Hz; calculate the energy from the index position to the highest frequency; finally, calculate the energy percentage, percentage = (high frequency band energy) / (total energy).

[0052] S133, calculate the descriptive statistics of the frame-level feature sequence and use them as statistics of speech sub-language features.

[0053] Specifically, after obtaining the frame-level feature sequences of various speech sub-language features, these sequences need to be statistically aggregated to obtain statistics that can characterize the entire speech segment in a fixed dimension. These statistics include the mean, standard deviation, median, maximum, and minimum. Furthermore, for the zero-crossing rate feature based on the power spectral density Z-fraction itself, its mean, standard deviation, median, maximum, and minimum values ​​are calculated again from the numerical sequences obtained across all frames, denoted as ZCR zPSDsw.

[0054] After obtaining the statistics for all categories of speech sub-language features, they need to be organized into a standard input format for machine learning models, namely a fixed-dimensional feature vector. Specifically, this can be achieved by concatenating and combining all descriptive statistics calculated for each category of speech sub-language features in a predetermined and uniform order. Through this concatenation operation, all statistics are integrated into a one-dimensional and structured numerical vector, which constitutes the final feature vector used as input for subsequent mental health testing models.

[0055] Implementing steps 131 to 133 above solves the problem that existing technologies struggle to extract effective mental health-related features from long-duration free-flowing interviews. By extracting a set of paralinguistic features unrelated to semantic content, such as zero-crossing rate, MFCC, fundamental frequency, formants, spectral centroid and flatness, PSD, and their derived features, a multidimensional feature sequence capable of quantifying speech prosody, tone quality, and spectral complexity is constructed, providing a data foundation for mental state assessment. Furthermore, by aggregating the sequence into a fixed-dimensional high-order statistical feature vector, the data dimensionality is reduced and the influence of varying time lengths is eliminated, generating the standardized input format required by the model. This enables subsequent mental health detection models to perform stable and efficient pattern recognition and risk status assessment based on this format.

[0056] In practice, the mental health detection model is a neural network model consisting of one-dimensional convolutional layers and fully connected layers. Specifically, two one-dimensional convolutional layers are used, each with 256 filters and a kernel size of 3. A one-dimensional max-pooling layer with a pooling size of 3 is placed between the two convolutional layers to help eliminate time-frequency variability caused by speech variability in each recording. The fully connected layers consist of four fully connected hidden layers with progressively decreasing dimensions, followed by an output layer. Furthermore, the ReLU activation function is applied to the outputs of the two convolutional layers and each fully connected hidden layer, and a dropout rate of 0.5 is set after the first two fully connected layers to prevent overfitting.

[0057] The mental health detection model was trained on multiple interview audio datasets labeled with mental health status, with four samples per batch for a total of 400 training iterations. Each data sample used to train the mental health detection model is a feature vector extracted from a single set of interview audio datasets labeled with mental health status, constructed from statistical measures of speech paralinguistic features.

[0058] After the interview audio data undergoes the preprocessing described above and statistical extraction of speech paralinguistic features, a final feature vector representing the interview is obtained. This feature vector has a fixed dimension and is used as input to the trained mental health detection model. The process is as follows: First, the input feature vectors are prepared by organizing them into a four-dimensional tensor to meet the model's input requirements. Since the feature vectors are static vectors constructed based on statistics from the entire interview speech, the dimension representing the sequence length in their input shape is 1. Specifically, the shape of the input tensor is (B, 1, N), where B is the batch size, set according to training, but may be 1 or padded to 4 when predicting a single sample; N is the total dimension of the feature vectors, for example, 1250 dimensions. This input tensor will be fed into the neural network model for forward propagation computation.

[0059] The input tensor is then fed into the first one-dimensional convolutional layer, which uses 256 filters of size 3 to convolve along the feature dimension to capture local combination patterns between features. Subsequently, a ReLU activation function is applied to introduce non-linearity. The output tensor of this layer has a shape of (B, 1, 256).

[0060] Then, the first one-dimensional maximization layer is entered, using a max pooling operation with a pooling size of 3 to reduce the dimensionality of the features, thereby enhancing the model's robustness to changes in feature location and suppressing minor fluctuations. The output tensor shape of this layer is (B, 1, 85).

[0061] Then, the second one-dimensional convolutional layer is used for processing. This layer also uses 256 filters of size 3 and the ReLU activation function to perform a higher-level abstract combination of the pooled features. The output tensor of this layer has a shape of (B, 1, 256). Then, the flattening operation is performed, which flattens the multidimensional tensor output by the second one-dimensional convolutional layer into a one-dimensional vector with a shape of (B, 256) so that it can be input into the fully connected layer.

[0062] Then, the data is passed through multiple fully connected hidden layers with progressively decreasing dimensions. Each layer is followed by a ReLU activation function. The fully connected layers perform a global nonlinear transformation and mapping on the flattened features.

[0063] Finally, the output of the last fully connected layer is passed to the output layer. For binary classification problems, the output layer typically has one neuron; for multi-class classification, the number of neurons equals the number of classes. The output layer uses either the sigmoid activation function (for binary classification) or the softmax function (for multi-class classification). These functions map the raw output value of the neuron to a probability value between 0 and 1, representing the probability that a subject is in a state of psychological risk. For example, in binary classification, an output of 0.85 indicates an 85% confidence level that the subject is in a state of psychological risk.

[0064] The AI-based interview voice psychological health detection and early warning method provided in this application has a wide range of applications, including but not limited to the following fields: School Student Psychological Screening: Applicable to large-scale student mental health surveys conducted regularly. It can guide students to complete a standardized short voice interview in a counseling room or through a secure student terminal. The system automatically analyzes and generates a risk assessment report, helping schools quickly identify students who need special attention, achieve early intervention, and establish dynamic psychological profiles.

[0065] Employee Assistance Program (EAP): Integrated into the company's EAP service system, employees can complete confidential interviews through the company's internal system. The system provides instant feedback or suggestions to help company managers understand the trends in the group's psychological state.

[0066] Initial screening in telemedicine and internet hospitals: Before formal treatment, patients can complete a brief voice interview through a mobile application. The system automatically conducts a preliminary psychological assessment and provides the results as background information to the attending physician, improving the efficiency of remote consultations and helping doctors quickly grasp the core issues.

[0067] In all the aforementioned scenarios, this invention effectively avoids assessment biases caused by language organization skills, cultural background, or subjective concealment by analyzing content-irrelevant paralinguistic features. It provides various institutions with an efficient initial mental health screening tool, significantly reducing reliance on professional human resources and screening costs. Furthermore, this method is also applicable to various occasions requiring preliminary mental health assessments, such as community health surveys, forensic psychological evaluations, and monitoring of special populations.

[0068] In one exemplary embodiment, such as Figure 5 As shown, an artificial intelligence-based interview voice psychological health detection and early warning device is provided, including: an acquisition module 201, a preprocessing module 202, a statistics extraction module 203, a model prediction module 204, and a probability value output module 205.

[0069] The acquisition module 201 is configured to acquire audio data of the interview between the subject and the interviewer.

[0070] The preprocessing module 202 is configured to perform preprocessing on the interview audio data to separate continuous speech segments spoken by the subject.

[0071] The statistics extraction module 203 is configured to extract statistics of speech paralinguistic features from continuous speech segments spoken by the subject to form a speech feature set for mental health assessment.

[0072] The model prediction module 204 is configured to input a feature vector constructed based on a set of speech features into a pre-trained mental health detection model, wherein the mental health detection model is a neural network model.

[0073] The probability value output module 205 is configured to execute the output based on the mental health testing model to obtain the probability value of the subject being in a state of psychological risk.

[0074] In one exemplary embodiment, a computer device is provided, which may be a server or a terminal, and its internal structure diagram may be as follows. Figure 6 As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operation of the operating system and computer programs in the non-volatile storage media. The database stores audio data from interviews between the subject and the interviewer. The I / O interfaces are used for information exchange between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements an artificial intelligence-based interview voice-based mental health detection and early warning method.

[0075] Those skilled in the art will understand that Figure 6 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0076] In one exemplary embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.

[0077] In one exemplary embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0078] In one exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0079] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0080] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).

[0081] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0082] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0083] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for detecting and warning of mental health through interview voice recordings based on artificial intelligence, characterized in that, include: Acquire audio data of interviews between subjects and interviewers; The interview audio data is preprocessed to separate continuous speech segments spoken by the subject; Statistical measures of speech paralinguistic features are extracted from continuous speech segments spoken by the subjects to form a speech feature set for mental health assessment. The feature vector constructed based on the speech feature set is input into a pre-trained mental health detection model, wherein the mental health detection model is a neural network model; Based on the output of the mental health testing model, the probability value of the subject being in a state of psychological risk is obtained.

2. The AI-based interview voice psychological health detection and early warning method according to claim 1, characterized in that, The preprocessing of the interview audio data includes: Speech activity detection was performed on the interview audio data to identify mixed human voice segments; The mixed human voice segments are distinguished by speaker to filter out the interviewer's voice, resulting in continuous speech segments produced by the subject.

3. The AI-based interview voice psychological health detection and early warning method according to claim 1, characterized in that, The statistical measures for extracting speech paralinguistic features include: The continuous speech segments spoken by the subject are segmented and windowed to obtain multiple short speech frames; Calculate the frame-level feature value of at least one type of speech sub-language feature in each of the short-time speech frames to obtain the frame-level feature sequence of each type of speech sub-language feature; Calculate the descriptive statistics of the frame-level feature sequence and use them as the statistics of the speech sub-language features.

4. The AI-based interview voice psychological health detection and early warning method according to claim 3, characterized in that, The categories of the speech sub-language features include at least one of the following: power spectral density, zero-crossing rate, zero-crossing rate based on power spectral density Z-fraction, Mel frequency cepstral coefficient, fundamental frequency, formant frequency, spectral centroid, spectral flatness, and energy percentage of a specific frequency band.

5. The AI-based interview voice psychological health detection and early warning method according to claim 3, characterized in that, The statistics include mean, standard deviation, median, maximum, and minimum.

6. The artificial intelligence-based interview voice psychological health detection and early warning method according to any one of claims 3 to 5, characterized in that, The feature vector constructed based on the speech feature set includes: The statistics corresponding to all categories of speech sub-language features are concatenated to form the feature vector.

7. The artificial intelligence-based interview voice psychological health detection and early warning method according to claim 1, characterized in that, The mental health detection model is trained based on multiple interview audio data that have been labeled with mental health status tags; wherein, each data sample used to train the mental health detection model is a feature vector extracted from the interview audio data labeled with mental health status tags and constructed from statistical measures of speech paralinguistic features.

8. An artificial intelligence-based interview voice psychological health detection and early warning device, characterized in that, include: The acquisition module is configured to acquire audio data from interviews between the subject and the interviewer. The preprocessing module is configured to perform preprocessing on the interview audio data to separate continuous speech segments spoken by the subject. The statistics extraction module is configured to extract statistics of speech paralinguistic features from continuous speech segments spoken by the subject to form a speech feature set for mental health assessment. The model prediction module is configured to input a feature vector constructed based on the speech feature set into a pre-trained mental health detection model, wherein the mental health detection model is a neural network model; The probability value output module is configured to execute the output of the mental health testing model to obtain the probability value of the subject being in a state of psychological risk.

9. A computer device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the artificial intelligence-based interview voice psychological health detection and early warning method according to any one of claims 1-7.

10. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the AI-based interview voice psychological health detection and early warning method as described in any one of claims 1-7.