Graph intuitionistic fuzzy-based depression identification method, apparatus and device, and readable storage medium

By adopting a depth model based on graph intuitive fuzzy in depression recognition, the problem of low accuracy and efficiency of depression recognition in the prior art is solved, and more efficient recognition effects and processing power for noise and outliers are achieved.

CN120048290APending Publication Date: 2025-05-27THE CHINESE UNIV OF HONG KONG (SHENZHEN)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510118249.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The prior art has problems of accuracy and inefficiency in the identification of depression, especially when processing physiological signal data of noise or outliers.

Method used

The depth model based on graph intuitive fuzziness is used to build the model by increasing the incremental method of node-by-node-by-hidden layer by node, and trained based on the intuitive fuzzy values ​​of each training sample dataset to improve the accuracy and efficiency of depression recognition.

Benefits of technology

Improves the accuracy and efficiency of depression identification, can handle noise and outliers more effectively, and reduces the risk of overfitting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048290A_ABST
    Figure CN120048290A_ABST
Patent Text Reader

Abstract

The invention provides a depression recognition method, device and equipment based on graph intuitionistic fuzziness and a readable storage medium. The depression recognition method comprises the steps of obtaining voice data of a to-be-processed object; extracting and determining an audio feature sample set corresponding to the voice data; constructing a depth model based on an increment mode of increasing nodes one by one and hidden layers one by one, and training the constructed depth model based on the intuitionistic fuzzy numerical value corresponding to each training sample in the unbalanced training sample data set to obtain a trained depression recognition model; the depth model comprises a plurality of hidden layers, and each hidden layer comprises a plurality of nodes; the unbalanced training sample comprises a plurality of positive class samples and a plurality of negative class samples; and inputting the audio feature sample set into a trained depression identification model to carry out identification and prediction processing so as to obtain an identification and prediction result of the to-be-processed object. According to the scheme, the accuracy and efficiency of depression identification can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer information processing, and particularly to a method, device, equipment and readable storage medium for depression recognition based on graph intuitionistic fuzzy. Background Art

[0002] Depression is the most common affective disorder, manifested as persistent low mood. It not only causes problems to the physical and mental health of patients, but also increases the risk of suicide and self-harm. In recent years, the application of artificial intelligence technology in the medical field has become more and more extensive. Using the voice signals of patients to diagnose depression is of great significance, which can not only help doctors diagnose depression more accurately, but also improve the efficiency of early detection and prevention.

[0003] In recent years, neural networks based on random weights have been widely applied compared with neural networks based on gradient descent, with fewer parameters and faster training efficiency. Different from traditional neural networks based on random weights, Stochastic Configuration Networks (SCNs) and Deep Stochastic Configuration Networks (DSCNs) use data-dependent inequality constraints to allocate neuron parameters and have good generalization performance. However, in practical applications, in addition to sample imbalance, disease data based on physiological signals may contain noise or outliers, which easily affect the performance of classifiers and reduce the accuracy of depression recognition.

[0004] Generally, the cost-sensitive least squares loss function can improve the robustness of data analysis problems containing outliers. Among them, Intuitionistic Fuzzy Sets (IFS) was proposed in 1983. By combining the membership function and non-membership function, it extends the traditional fuzzy set, enabling IFS to more comprehensively describe fuzzy concepts in practical applications, providing more information and more effectively capturing uncertainty. Subsequently, the IFS theory has been effectively applied to mitigate the adverse effects of noise and outliers on the performance of machine learning models. At present, the cost-sensitive l2-norm loss function based on IFS has been used in stochastic configuration networks. Intuitionistic Fuzzy Stochastic Configuration Networks (IFSCNs) use a cost-sensitive learning framework based on intuitionistic fuzzy to determine network parameters, ensuring the universal approximation performance of IFSCNs.

[0005] However, in practical applications, the calculation methods of membership functions or non-membership functions are crucial for the results of fuzzy classifiers. Most intuitionistic fuzzy-based models mainly use kernel functions to calculate membership functions and non-membership functions, ignoring the analysis of the relative neighborhood density of samples and the influence of hesitancy on the results. In addition, although IFSCNs can combine multi-layer feature mapping or multi-layer structures to improve the accuracy performance of single-layer models, as the number of nodes or hidden layers increases, the model is prone to overfitting problems to varying degrees. Summary of the Invention

[0006] The technical problem to be solved by the present invention is to provide a method, device, equipment and readable storage medium for depression recognition based on graph intuitionistic fuzzy, so as to improve the accuracy and efficiency of depression recognition.

[0007] To solve the above technical problem, an embodiment of the present invention provides a method for depression recognition based on graph intuitionistic fuzzy, including:

[0008] Obtain the speech data of the object to be processed;

[0009] Extract and determine the audio feature sample set corresponding to the speech data;

[0010] Construct a deep model based on an incremental method of adding one node and one hidden layer at a time, and train the constructed deep model based on the intuitionistic fuzzy numerical pairs corresponding to each training sample in the imbalanced training sample dataset to obtain a trained depression recognition model; the deep model includes multiple hidden layers, and each hidden layer includes multiple nodes; the imbalanced training samples include multiple positive class samples and multiple negative class samples;

[0011] Input the audio feature sample set into the trained depression recognition model for recognition and prediction processing to obtain the recognition and prediction result of the object to be processed.

[0012] In one embodiment, extracting and determining the audio feature sample set corresponding to the speech data includes:

[0013] Decompose the speech data to obtain multiple speech frames corresponding to the speech data;

[0014] Extract features from the multiple speech frames according to a preset audio feature extraction algorithm to obtain the audio feature sample set corresponding to the speech data.

[0015] In one embodiment, training the deep model based on the intuitionistic fuzzy numerical pairs corresponding to each training sample in the imbalanced training sample dataset to obtain a depression recognition model includes:

[0016] Determine the membership degree and non - membership degree corresponding to each training sample in the unbalanced training sample dataset;

[0017] Take the sample value, label value, the membership degree and the non - membership degree corresponding to each training sample as the training array of the current training sample and input it into the constructed deep model, and perform training processing on each hidden layer in turn to obtain the depression recognition model.

[0018] In one embodiment, determining the membership degree and non - membership degree corresponding to each sample in the unbalanced training sample dataset includes:

[0019] According to the distance between the current training sample in the unbalanced training sample dataset and its k - nearest neighbor similar training samples, determine the similarity between the current training sample and its k - nearest neighbor similar training samples;

[0020] According to the similarity between the current training sample and its k - nearest neighbor similar training samples, determine the membership degree and non - membership degree corresponding to each training sample in the unbalanced training sample dataset.

[0021] In one embodiment, taking the sample value, label value, the membership degree and the non - membership degree corresponding to each training sample as the training array of the current training sample and inputting it into the constructed deep model, and performing training processing on each hidden layer in turn to obtain the depression recognition model includes:

[0022] According to the membership degree and non - membership degree corresponding to each training sample, determine the scoring value corresponding to each training sample;

[0023] Input the sample value, label value and scoring value corresponding to each training sample into multiple hidden layers of the deep model for weighted processing, and obtain the training residual corresponding to each hidden layer and the output matrix of each node in each hidden layer; wherein, the input weight of each input node in the first hidden layer of the deep model is determined according to the scoring value corresponding to each training sample;

[0024] According to the output matrix and residual of each node in each hidden layer, obtain the training prediction result and the trained depression recognition model.

[0025] In one embodiment, obtaining the trained depression recognition model according to the output matrix and residual of each node in each hidden layer includes:

[0026] Based on the objective loss function of graph intuitionistic fuzzy, update the model parameters of the deep model to obtain the updated model parameters, and the model parameters include the input weight of the hidden layer and the output weight of the output layer corresponding to the hidden layer;

[0027] Determine the trained depression recognition model according to the updated model parameters.

[0028] In one embodiment, based on the objective loss function of graph intuitionistic fuzzy, update the model parameters of the deep model to obtain the updated model parameters, including:

[0029] Determine the input weights of the next hidden layer based on the objective loss function of graph intuitionistic fuzzy, the output matrix and the residual of each node in the current hidden layer;

[0030] Determine the output weights of each hidden layer based on the objective loss function of graph intuitionistic fuzzy, the output matrix of each node in each hidden layer, and the scoring function corresponding to the training samples.

[0031] An embodiment of the present invention further provides a depression recognition device based on graph intuitionistic fuzzy, including:

[0032] A data acquisition module, configured to acquire the speech data of the object to be processed;

[0033] A feature extraction module, configured to extract and determine the audio feature sample set corresponding to the speech data;

[0034] A model construction module, configured to construct a deep model based on an incremental manner of adding one node by one node and one hidden layer by one hidden layer, and train the constructed deep model based on the intuitionistic fuzzy numerical pairs corresponding to each training sample in the unbalanced training sample dataset to obtain a trained depression recognition model; the deep model includes multiple hidden layers, and each hidden layer includes multiple nodes; the unbalanced training samples include multiple positive class samples and multiple negative class samples;

[0035] An identification and prediction module, configured to input the audio feature sample set into the trained depression recognition model for identification and prediction processing to obtain the identification and prediction result of the object to be processed.

[0036] An embodiment of the present invention further provides a computing device, including:

[0037] A memory, configured to store one or more programs;

[0038] One or more processors, configured to execute the one or more programs to implement the method described in any one of the above.

[0039] An embodiment of the present invention further provides a computer-readable storage medium, characterized in that it stores instructions, and when the instructions run on a computer, the computer is made to execute the method described in any one of the above.

[0040] The above solution of the present invention has at least the following beneficial effects:

[0041] The method, device, equipment and readable storage medium for depression recognition based on graph intuitionistic fuzzy provided by the above solution of the present invention include: obtaining the speech data of an object to be processed; extracting and determining the audio feature sample set corresponding to the speech data; constructing a deep model based on an incremental manner of adding one node by one node and one hidden layer by one hidden layer, and training the constructed deep model based on the intuitionistic fuzzy numerical pairs corresponding to each training sample in the unbalanced training sample dataset to obtain a trained depression recognition model; the deep model includes multiple hidden layers, and each hidden layer includes multiple nodes; the unbalanced training sample includes multiple positive class samples and multiple negative class samples; inputting the audio feature sample set into the trained depression recognition model for recognition and prediction processing to obtain the recognition and prediction result of the object to be processed, thereby improving the accuracy and efficiency of depression recognition. Description of the Drawings

[0042] Figure 1 is a flowchart of the method for depression recognition based on graph intuitionistic fuzzy provided by an embodiment of the present invention;

[0043] Figure 2 is an architecture diagram of the depression recognition model provided by an optional embodiment of the present invention;

[0044] Figure 3 is a schematic block diagram of the modules of the device for depression recognition based on graph intuitionistic fuzzy provided by an embodiment of the present invention;

[0045] Figure 4 is a schematic block diagram of an electronic device provided by an embodiment of the present invention; and

[0046] Figure 5 is a schematic block diagram of a computing device provided by an embodiment of the present invention. Detailed Embodiments

[0047] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be completely conveyed to those skilled in the art.

[0048] In the following description, for the purpose of explaining various disclosed embodiments, certain specific details are set forth to provide a thorough understanding of the various disclosed embodiments. However, those skilled in the relevant art will recognize that the embodiments can be practiced without one or more of these specific details. In other instances, well-known devices, structures, and techniques associated with the present application may not be shown or described in detail so as not to unnecessarily obscure the description of the embodiments.

[0049] References to "one embodiment" or "an embodiment" throughout the specification mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. Thus, the appearances of "in one embodiment" or "in an embodiment" in various places throughout the specification are not necessarily all referring to the same embodiment. Additionally, the particular features, structures, or characteristics may be combined in any manner in one or more embodiments.

[0050] In the following description, for the purpose of clearly showing the structure and working mode of the present invention, many directional terms will be used for description. However, terms such as "front", "rear", "left", "right", "outer", "inner", "outward", "inward", "up", "down", etc. should be understood as convenient terms and should not be understood as limiting terms.

[0051] As Figure 1 shown, an embodiment of the present disclosure provides a method 10 for identifying depression based on graph intuitionistic fuzzy, including:

[0052] Step 11: Obtain the speech data of the object to be processed;

[0053] Step 12: Extract and determine the audio feature sample set corresponding to the speech data;

[0054] Step 13: Construct a deep model in an incremental manner by adding one node and one hidden layer at a time, and train the constructed deep model based on the intuitionistic fuzzy numerical pairs corresponding to each training sample in the unbalanced training sample dataset to obtain a trained depression recognition model; the deep model includes multiple hidden layers, and each hidden layer includes multiple nodes; the unbalanced training sample includes multiple positive class samples and multiple negative class samples;

[0055] Step 14: Input the audio feature sample set into the trained depression recognition model for recognition and prediction processing to obtain the recognition and prediction result of the object to be processed.

[0056] In this embodiment, the speech data corresponding to the object to be processed is collected and data perception is performed to construct an audio feature sample set corresponding to the speech cognitive ability of the object to be processed, preparing for subsequent recognition and prediction;

[0057] The depth model is constructed in a node increment manner. The depth model consists of multiple base learners. Each base learner contains a hidden layer, and each hidden layer includes multiple nodes. At the same time, each hidden layer corresponds to an output layer. After the audio feature sample set is input into the trained depression recognition model, each audio feature sample will first be weighted processed on each node of the hidden layer of the first base learner. The corresponding output layer outputs a recognition prediction result according to the weighted processing result. At the same time, the weighted processing results on each node of the hidden layer of the first base learner are input into each node of the hidden layer of the second base learner for weighted processing. The output layer of the second base learner outputs a recognition prediction result according to the weighted processing result of the corresponding hidden layer. By analogy, assuming there are M base learners, corresponding to M hidden layers, the first layer calculates an output result, the first and second layers calculate an output result, and from the first layer to the Mth hidden layer, there will be M recognition prediction results (the input data of the next hidden layer is the weighted processing result of the previous hidden layer; the input data of the first hidden layer is the audio feature sample corresponding to the voice data of the object to be recognized). Further, the M recognition prediction results are used to determine the final recognition prediction result according to the principle of the minority obeying the majority, so as to ensure the accuracy of the recognition prediction result.

[0058] Here, the positive class samples in the imbalanced training sample dataset represent the training feature samples corresponding to healthy objects, and the negative class samples represent the training feature samples corresponding to diseased (depressed) objects. During the model training process, the constructed depth model is trained based on the imbalanced training sample dataset, and the intuitionistic fuzzy values determined based on the relative neighborhood density between the training sample data are considered during the training process. At the same time, the objective loss function based on graph intuitionistic fuzzy and the weighted supervision mechanism are used to determine the model parameters of the depth model, so as to solve the problems of sample imbalance and noise, and further ensure the accuracy of the recognition prediction result of the trained depression recognition model.

[0059] In an optional embodiment of the present disclosure, step 12 above may include:

[0060] Step 121, decompose the voice data to obtain multiple voice frames corresponding to the voice data;

[0061] Step 122, extract features from the multiple voice frames according to a preset audio feature extraction algorithm to obtain an audio feature sample set corresponding to the voice data.

[0062] In this embodiment, due to clinical manifestations such as lack of energy, reduced movement, and slowed thinking, patients with depression usually exhibit a slower speech rate, monotonous pitch changes, and weaker volume compared to healthy individuals. In this embodiment, the syllable speech in the speech frame can be processed through a preset audio feature extraction algorithm. Preferably, the preset audio feature extraction algorithm can be the Mel Frequency Cepstral Coefficients (MFCC). By using MFCC, the subtle differences in the syllable speech can be captured, which helps to more accurately identify and analyze the speech. Further, calculate the mean and standard deviation of the audio features captured by MFCC, and use the mean and standard deviation of the audio features to describe the overall spectral characteristics and variability of the speech data. And use the mean and standard deviation of the audio features as the feature data of different dimensions for each audio feature sample for subsequent identification and prediction processing.

[0063] In an alternative embodiment of the present disclosure, training a deep model based on the intuitionistic fuzzy numerical pairs corresponding to each training sample in the imbalanced training sample dataset to obtain a depression recognition model may include:

[0064] Step 21, determine the membership degree and non-membership degree corresponding to each training sample in the imbalanced training sample dataset.

[0065] Here, an intuitionistic fuzzy numerical value can be assigned to each sample based on a direct fuzzy classifier. The intuitionistic fuzzy numerical value includes a membership degree and a non-membership degree. Each positive class sample is assigned a membership degree and a non-membership degree, and each negative class sample is assigned a membership degree and a non-membership degree.

[0066] The membership degree represents the ratio of the number of positive class samples within the neighborhood of the current training sample to the total number of training samples, and the non-membership degree represents the ratio of the number of negative class samples within the neighborhood of the current training sample to the total number of training samples. To improve the robustness of intuitionistic fuzzy in solving outlier and class imbalance learning problems, here, the intuitionistic fuzzy numerical value of each training sample can be determined according to the relative neighborhood density of the imbalanced training samples.

[0067] Specifically, in an alternative embodiment of the present disclosure, the above step 21 may include:

[0068] Step 211, determine the similarity between the current training sample in the imbalanced training sample dataset and its k nearest neighbor's same-class training samples according to the distance between them.

[0069] Step 212, determine the membership degree and non-membership degree corresponding to each training sample in the imbalanced training sample dataset according to the similarity between the current training sample and its k nearest neighbor's same-class training samples.

[0070] In this embodiment, the similarity between the current training sample and its k nearest neighbor training samples of the same class is the relative neighborhood density. Let x i represent the current training sample, and let x j represent the k nearest neighbor training samples that are of the same class as x i (both are positive class samples or both are negative class samples). The distance between the current training sample and its k nearest neighbor training samples of the same class can be represented by the Euclidean distance between the two, and the specific calculation process is shown in formulas (1) and (2):

[0071] P(x i , x j ) = ||x i - x j || 2 , subject: y i = +1, y j = +1; (1)

[0072] N(x i , x j ) = ||x i - x j || 2 , subject: y i = -1, y j = -1; (2)

[0073] Among them, y i = +1 indicates that x i is a positive class sample, y i = -1 indicates that x i is a negative class sample, y j = +1 indicates that x j is a positive class sample, y j = -1 indicates that x j is a negative class sample, P(x i , x j ) represents the distance between the two when x i and x j are both positive class samples; N(x i , x j ) represents the distance between the two when x i and x j are both negative class samples;

[0074] For unbalanced positive and negative samples, when and are k nearest neighbors in the same class, the preset kernel function K(·) can be used to represent the similarities G i and G j of x ij in positive class samples and negative class samples respectively. The specific calculation process is shown in formula (3):

[0075]

[0076] Furthermore, based on the similarity obtained from the above calculations, the membership function of the current sample x can be calculated through formulas (4) and (5). i of the membership function

[0077]

[0078] wherein,

[0079] wherein, k1 and k2 respectively represent the k-nearest neighbors of the positive-class samples and negative-class samples; G i represents the sum of the similarities between the current training sample x i and other training samples of the same class in the neighborhood; Based on the above membership function, the membership degree of each training sample can be calculated and obtained.

[0080] Furthermore, the non-membership function of the current training sample x i of the non-membership function can be expressed as formula (6):

[0081]

[0082] wherein,

[0083] Here, A represents the coefficient of non-membership. The numerator in formula (7) represents the number of training samples of different classes whose kernel space distances are less than η, and the denominator in formula (7) represents the number of all training samples whose kernel space distances are less than η; Based on the above non-membership function, the non-membership degree of each training sample can be calculated and obtained.

[0084] Furthermore, after determining the membership function and non-membership function of each training sample, the training sample set can be defined as respectively represent the membership degree and non-membership degree of the sample x i of the membership degree and non-membership degree.

[0085] In order to obtain richer discriminant information, a scoring function Θ combining membership degree, non-membership degree and hesitation degree is defined for each training sample i , specifically as shown in formula (8):

[0086]

[0087] wherein, θ and υ are respectively preset weight parameters, and P represents a parameter for controlling the shape of the scoring function. It represents the hesitation degree; the scoring value of each training sample can be calculated according to the corresponding scoring function of each training sample, so as to prepare for subsequent model training.

[0088] In an optional embodiment of the present disclosure, training a deep model based on the intuitionistic fuzzy numerical pairs corresponding to each training sample in the unbalanced training sample dataset to obtain a depression recognition model may include:

[0089] Step 22: Input the sample value, label value, membership degree, and non-membership degree corresponding to each training sample as the training array of the current training sample into the constructed deep model, and perform training processing on each hidden layer in turn to obtain a depression recognition model.

[0090] In this embodiment, according to the sample value and label value corresponding to each training sample, the model training objective is defined and respectively represent the sample value feature and label value of the training sample, d and m respectively represent the corresponding dimensions, and N represents the number of training samples.

[0091] In an optional embodiment of the present disclosure, the above step 22 may include:

[0092] Step 221: Determine the scoring value corresponding to each training sample according to the membership degree and non-membership degree corresponding to each training sample;

[0093] Step 222: Input the sample value, label value, and scoring value corresponding to each training sample into multiple hidden layers of the deep model for weighted processing, and obtain the training residual corresponding to each hidden layer and the output matrix of each node in each hidden layer; among them, the input weight of each input node in the first hidden layer of the deep model is determined according to the scoring value corresponding to each training sample;

[0094] Step 223: Obtain the training prediction result and the trained depression recognition model according to the output matrix and residual of each node in each hidden layer.

[0095] In this embodiment, different from the deep random configuration network, as Figure 2 shown, the deep model may have M hidden layers. On each node in each hidden layer, the input data is weighted processed in turn, and the corresponding weighted processing result (output matrix) is output. Further, the weighted processing result of the current hidden layer is input into the corresponding output layer for processing to obtain a training output result; in turn, the weighted processing result of the previous hidden layer is used as the input data of each node in the next hidden layer, and the above training steps are repeated until the set M hidden layers are trained and the target loss function has converged, and the training ends to obtain the trained depression recognition model;

[0096] Preferably, here, the target loss function can be a cost-sensitive target loss function based on graph intuitionistic fuzzy to solve the problem of outliers or noise values occurring during the training process:

[0097] The cost-sensitive target loss function based on graph intuitionistic fuzzy is specifically expressed as shown in formula (9):

[0098]

[0099] where, Θ = diag(Θ 1 , Θ 2 ,..., Θ N ) represents the graph intuitionistic fuzzy diagonal matrix, which is calculated from the scoring function Θ i corresponding to the initially input training samples; represents the activation function of the hidden layer n, where n = 1, 2,..., M; and respectively represent the input weight and bias of node j in the hidden layer n; represents the output weight of node j in the hidden layer n, where j = 1, 2,..., L n ; represents the regularization term of the l2 norm to avoid overfitting; C represents the regularization parameter.

[0100] In an alternative embodiment of the present disclosure, in step 223 above, obtaining the trained depression recognition model according to the output matrix and residual of each node in each hidden layer may include:

[0101] Step 2231, based on the target loss function based on graph intuitionistic fuzzy, update the model parameters of the deep model to obtain the updated model parameters, where the model parameters include the input weights of the hidden layer and the output weights of the corresponding output layer of the hidden layer;

[0102] Step 2232, determine the trained depression recognition model according to the updated model parameters.

[0103] In this embodiment, after the deep model is constructed in an incremental manner by adding one node by one node and one hidden layer by one hidden layer, the model starts to be trained; and during the model training, the initial parameters of the constructed deep model are updated; here, the scoring values of the training samples in the initially input deep model can be used as the input weights of the nodes in the first hidden layer. After the weighted processing is completed on each node in the first hidden layer, the output result of the first hidden layer is used to determine the input weights of the upper nodes in the second hidden layer and the output weights of this hidden layer, and so on. The input and output weights of other hidden layers except the first hidden layer are updated in turn;

[0104] Here, a hidden layer and its corresponding output layer can be regarded as a basic model. According to the training output result of the first basic model (including the first hidden layer and its corresponding output layer) (the result obtained by the output layer performing weighted processing on the output matrix of the nodes in the corresponding hidden layer according to the output weights), the corresponding input and output weights of the subsequent hidden layer and output layer are determined to improve the generalization performance of the model; subsequently, an ensemble model is constructed based on the output results of all hidden layers and their corresponding output layers, that is, the finally trained depression recognition model is obtained.

[0105] In an optional embodiment of the present disclosure, the above step 2231 may include:

[0106] Step 22311: Determine the input weights of the next hidden layer based on the objective loss function of graph intuitionistic fuzzy, the output matrix and the residual of each node in the current hidden layer;

[0107] Step 22312: Determine the output weights of each hidden layer based on the objective loss function of graph intuitionistic fuzzy, the output matrix of each node in each hidden layer, and the scoring function corresponding to the training samples.

[0108] Here, the residual is the residual corresponding to the end of the training of the current hidden layer; it is assumed that there are M hidden layers (corresponding to M basic models) and L M -1 nodes. To effectively allocate the input weights of the nodes in the next hidden layer, a graph intuitionistic fuzzy weighting strategy is adopted to weight the residual of the current training and L M the output matrix of the nodes to obtain the weighted result:

[0109] Furthermore, an intuitionistic fuzzy-based constraint inequality is defined, as shown in formula (10). Among the candidate node parameters that satisfy the maximum value will be used as the input weight of the node L M in the new hidden layer;

[0110] Here,

[0111] where q = 1, 2,..., m represents the calculation dimension of the weighted residual, r represents the shrinkage factor, and its value range is between 0 and 1, which is used to avoid overfitting of the model;

[0112] Furthermore, according to the output matrix of each node in each hidden layer and the scoring function corresponding to the training samples, the output weights of each hidden layer can be determined, which can be calculated by formula (11):

[0113]

[0114] Among them, represents the weighted output matrix of all hidden layers of the n-th layer; represents the weighted output matrix of all L n nodes; C represents the regularization parameter; I represents the identity matrix;

[0115] Since the deep model adopts an incremental depth architecture, here, an ensemble model can be constructed by combining the output results of M basic models, and then the finally trained depression recognition model can be obtained; specifically, it can be represented by the following formulas (12) and (13):

[0116] f n (X) = H n β n ; (12)

[0117] f ensemble = Voting{f 1 (X), f 2 (X),..., f M (X)}; (13)

[0118] Among them, represents the output matrix of the hidden layer nodes of the first n layers; f n (X) represents the output result of the basic model n; f ensemble represents the voting result of the final output of the model, that is, the training recognition prediction result; n = 1, 2,..., M.

[0119] Through the above construction and training process, the finally trained depression recognition model is obtained. This model determines the membership function and non-membership function based on the relative neighborhood density of the imbalanced training samples, and combines the scoring functions of membership, non-membership, and hesitation degree for training to obtain richer discriminant information; at the same time, during the training process, the model parameters are determined based on the cost-sensitive objective loss function and weighted supervision mechanism of graph intuitionistic fuzzy to ensure the accuracy of subsequent model recognition and prediction.

[0120] The following is an experiment conducted on 5 speech data sets with different noise ratios. During the experiment, each algorithm was run 100 times, and the mean and standard deviation (STD) were recorded as the final results of the class imbalance problem.

[0121] Table 1. Comparison table of running results of different models

[0122]

[0123]

[0124] As can be seen from Table 1 above, the third model IFSCN performs worse than the first model SCN in solving the class-imbalanced depression dataset with outliers. This is because the classical intuitionistic fuzzy method ignores the influence of the relative neighborhood density of imbalanced samples with outliers; the second model performs better than the single-layer first model SCN because the multi-layer random neural network can learn high-level representation information and improve the classification accuracy. Compared with the first model SCN, the second model DSCN, and the third model IFSCN, the deep integrated depression recognition model using the graph intuitionistic fuzzy cost-sensitive objective loss function performs excellently in terms of the average G-mean of all experimental results.

[0125] As Figure 3 shown, an embodiment of the present invention further provides a depression recognition device 30 based on graph intuitionistic fuzzy, including:

[0126] A data acquisition module 31 for acquiring voice data of an object to be processed;

[0127] A feature extraction module 32 for extracting and determining an audio feature sample set corresponding to the voice data;

[0128] A model construction module 33 for constructing a deep model in an incremental manner of adding one node by one node and one hidden layer by one hidden layer, and training the constructed deep model based on the intuitionistic fuzzy numerical pairs corresponding to each training sample in the imbalanced training sample dataset to obtain a trained depression recognition model; the deep model includes multiple hidden layers, and each hidden layer includes multiple nodes; the imbalanced training samples include multiple positive class samples and multiple negative class samples;

[0129] An identification and prediction module 34 for inputting the audio feature sample set into the trained depression recognition model for identification and prediction processing to obtain an identification and prediction result of the object to be processed.

[0130] It should be noted that this device corresponds to the above-mentioned depression recognition method 10 based on graph intuitionistic fuzzy. All implementation manners in the above method embodiments are applicable to this embodiment and can also achieve the same technical effects.

[0131] As Figure 4As shown in the figure, an embodiment of the present invention further provides an electronic device 50, including: a memory 51 for storing one or more computer programs; one or more processors 52 for executing one or more computer programs, and when the computer programs are run by the processor, the depression recognition method 10 based on graph intuitionistic fuzzy as described above is executed. All implementation manners in the above method embodiments are applicable to this embodiment and can achieve the same technical effects. The electronic device 50 is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown in the present invention, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed in the present invention.

[0132] As Figure 5 shown, the electronic device 50 is manifested as a computing device, or a computer system, and may include a CPU 501 (computing unit), which can execute various appropriate actions and processes according to the computer program stored in the ROM 502 (read-only memory) or the computer program loaded from the storage unit 508 into the random access RAM 503 (memory). In the RAM 503, various programs and data required for the operation of the device 500 can also be stored. The CPU 501, ROM 502, and RAM 503 are connected to each other through a bus 504. The I / O interface 505 (input / output interface) is also connected to the bus 504.

[0133] Multiple components in the electronic device 500 are connected to the I / O interface 505, including: an input unit 506, such as a keyboard, a mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a disk, an optical disc, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows the device 500 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0134] The CPU 501 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the CPU 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The CPU 501 executes the various methods and processes described above. For example, in some embodiments, the graph intuitionistic fuzzy-based depression recognition method 10 can be implemented as a computer software program that is tangibly embodied in a computer-readable storage medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the CPU 501, one or more steps of the graph intuitionistic fuzzy-based depression recognition method 10 described above can be executed. Alternatively, in other embodiments, the CPU 501 can be configured to execute the graph intuitionistic fuzzy-based depression recognition method 10 by any other suitable means (e.g., by means of firmware).

[0135] Embodiments of the present invention also provide a computer-readable storage medium that stores instructions which, when run on a computer, cause the computer to execute the graph intuitionistic fuzzy-based depression recognition method 10 described above. All implementation manners in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.

[0136] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. A professional technician can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present invention.

[0137] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0138] In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.

[0139] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0140] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0141] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical discs that can store program codes.

[0142] In addition, it should be noted that in the device and method of the present invention, obviously, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations shall be regarded as equivalent solutions of the present invention. Moreover, the steps of performing the above series of processes can naturally be executed in chronological order according to the described order, but it is not necessary to be executed in chronological order. Some steps can be executed in parallel or independently of each other. For those of ordinary skill in the art, it is understandable that all or any steps or components of the method and device of the present invention can be implemented in any computing device (including processors, storage media, etc.) or a network of computing devices in the form of hardware, firmware, software, or a combination thereof, which can be achieved by those of ordinary skill in the art using their basic programming skills after reading the description of the present invention.

[0143] Therefore, the object of the present invention can also be achieved by running a program or a set of programs on any computing device. The computing device can be a well-known general-purpose device. Thus, the object of the present invention can also be achieved merely by providing a program product containing program code for implementing the method or device. That is to say, such a program product also constitutes the present invention, and a storage medium storing such a program product also constitutes the present invention. Obviously, the storage medium can be any well-known storage medium or any storage medium developed in the future. It should also be noted that in the device and method of the present invention, obviously, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations shall be regarded as equivalent solutions of the present invention. Moreover, the steps of performing the above series of processes can naturally be executed in chronological order according to the described order, but it is not necessary to be executed in chronological order. Some steps can be executed in parallel or independently of each other.

[0144] The above is the preferred embodiment of the present invention. It should be pointed out that for those of ordinary skill in the technical field, without departing from the principle described in the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A method for identifying depression based on graph intuitionistic fuzzy, characterized in that: include: Acquire voice data of the object to be processed; Extracting and determining an audio feature sample set corresponding to the speech data; A deep model is constructed based on an incremental method of adding nodes one by one and hidden layers one by one, and the constructed deep model is trained based on the intuitive fuzzy numerical value corresponding to each training sample in the unbalanced training sample data set to obtain a trained depression recognition model; the deep model includes multiple hidden layers, and each hidden layer includes multiple nodes; the unbalanced training samples include multiple positive samples and multiple negative samples; The audio feature sample set is input into a trained depression recognition model for recognition prediction processing to obtain a recognition prediction result of the object to be processed.

2. The method for identifying depression based on graph intuitionistic fuzzy according to claim 1, characterized in that: Extracting and determining an audio feature sample set corresponding to the speech data includes: Decomposing the voice data to obtain a plurality of voice frames for the voice data; Feature extraction is performed on the plurality of voice frames according to a preset audio feature extraction algorithm to obtain an audio feature sample set corresponding to the voice data.

3. The method for identifying depression based on graph-based intuitionistic fuzzy according to claim 1, characterized in that: The constructed deep model is trained based on the intuitive fuzzy value corresponding to each training sample in the unbalanced training sample data set to obtain a depression recognition model, including: Determining the degree of membership and the degree of non-membership corresponding to each training sample in the unbalanced training sample data set; The sample value and label value corresponding to each training sample, the membership degree and the non-membership degree corresponding to the training sample are input into the constructed deep model as the training array of the current training sample, and the training process is performed on each hidden layer in turn to obtain the depression recognition model.

4. The method for identifying depression based on graph intuitionistic fuzzy according to claim 3 is characterized in that: Determining the degree of membership and the degree of non-membership corresponding to each sample in the unbalanced training sample data set includes: Determine the similarity between the current training sample and the similar training samples of its k nearest neighbors according to the distance between the current training sample and the similar training samples of its k nearest neighbors in the unbalanced training sample data set; According to the similarity between the current training sample and its k nearest neighbor training samples of the same type, the membership degree and the non-membership degree corresponding to each training sample in the unbalanced training sample data set are determined.

5. The method for identifying depression based on graph intuitionistic fuzzy according to claim 3, characterized in that: The sample value and label value corresponding to each training sample, the membership degree and the non-membership degree corresponding to the training sample are input into the constructed deep model as the training array of the current training sample, and the training process is performed on each hidden layer in turn to obtain the depression recognition model, including: Determine the score value corresponding to each training sample according to the membership degree and non-membership degree corresponding to each training sample; The sample value, label value and score value corresponding to each training sample are input into multiple hidden layers of the deep model for weighted processing, and the training residual corresponding to each hidden layer and the output matrix of each node in each hidden layer are obtained; wherein the input weight of each input node in the first hidden layer of the deep model is determined according to the score value corresponding to each training sample; According to the output matrix and residual of each node in each hidden layer, the training prediction result and the trained depression recognition model are obtained.

6. The method for identifying depression based on graph intuitionistic fuzzy according to claim 5, characterized in that: According to the output matrix and residual of each node in each hidden layer, the trained depression recognition model is obtained, including: Based on the objective loss function of graph intuitionistic fuzzy, the model parameters of the deep model are updated to obtain updated model parameters, wherein the model parameters include the input weight of the hidden layer and the output weight of the output layer corresponding to the hidden layer; The trained depression recognition model is determined according to the updated model parameters.

7. The method for identifying depression based on graph intuitionistic fuzzy according to claim 6, characterized in that: Based on the objective loss function of graph intuitionistic fuzzy, the model parameters of the deep model are updated to obtain updated model parameters, including: Determine the input weight of the next hidden layer based on the objective loss function of the graph intuitionistic fuzzy and the output matrix and residual of each node in the current hidden layer; Based on the objective loss function of graph intuition fuzzy, the output matrix of each node in each hidden layer and the scoring function corresponding to the training samples, the output weight of each hidden layer is determined.

8. A depression identification device based on graph intuition fuzzy, characterized in that: include: A data acquisition module, used to acquire the voice data of the object to be processed; A feature extraction module, used to extract and determine an audio feature sample set corresponding to the speech data; A model building module, for building a deep model based on an incremental method of adding nodes one by one hidden layer, and training the built deep model based on the intuitive fuzzy numerical value corresponding to each training sample in the unbalanced training sample data set to obtain a trained depression recognition model; the deep model includes a plurality of hidden layers, and each hidden layer includes a plurality of nodes; the unbalanced training samples include a plurality of positive samples and a plurality of negative samples; The recognition prediction module is used to input the audio feature sample set into the trained depression recognition model for recognition prediction processing to obtain the recognition prediction result of the object to be processed.

9. A computing device, characterized in that include: A memory for storing one or more programs; One or more processors, configured to execute the one or more programs to implement the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: Instructions are stored, and when the instructions are executed on a computer, the computer is caused to execute the method according to any one of claims 1 to 7.