Viscera recognition method based on voice emotion and voice features
Through the internal organ recognition method based on pronunciation emotions and pronunciation characteristics, using emotional common characteristics and multi-scale encoder in European and hyperbolic spatial constraints, the problem of lack of labels for pronunciation and viscera data sets is solved, and the accuracy of internal organ recognition and the application value of traditional Chinese medicine emotional theory is improved.
Patent Information
- Application Number
- CN202510354505.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-03-25
AI Technical Summary
The prior art is difficult to effectively use a large amount of annotated voice emotional data for internal organ recognition, resulting in a lack of labels for voice and internal organ data sets, affecting the accuracy of identification.
The internal organ recognition method based on pronunciation emotions and pronunciation characteristics is adopted. Through the training process of the internal organ classifier, the emotional common characteristics of the internal organ emotions and the internal organ characteristics are extracted. After weighted fusion, the internal organ speech relationship embedded characteristics are obtained, and consistency constraints are carried out in European space and hyperbolic space. Multi-scale speech encoder and dual discriminator are used for feature extraction and constraints.
Make full use of emotional voice data to improve the accuracy and performance of internal organ recognition, verify the application value of traditional Chinese medicine emotional theory, and solve the problems of slow labeling and lack of labeling of data sets.
Smart Images

Figure CN120340535A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of viscera organ recognition, and more specifically to an viscera organ recognition method based on voice emotion and voice features. Background Art
[0002] In recent years, the concept of TCM health preservation and individualized conditioning methods have received extensive attention and application. Among them, sound is one of the important manifestations of human internal organs. By analyzing its corresponding internal organs and their characteristics, it can provide an important reference for formulating individualized conditioning plans.
[0003] Since sounds may correspond to multiple internal organs, the recognition of internal organs based on sounds is a multi-label classification problem. At present, data-driven machine learning is the main method to solve this problem, but it requires a large amount of annotated internal organ speech data, and it is difficult to obtain a large-scale sound dataset with internal organ annotations at this stage.
[0004] However, there are a large number of emotional speech datasets with emotion annotations. According to the theory of emotion in traditional Chinese medicine, the changes of human emotions such as joy, anger, thinking, sadness, fear, etc. are closely related to human internal organs such as heart, liver, spleen, lungs, and kidneys, and they influence each other.
[0005] Therefore, how to make full use of these emotional speech data to alleviate the problem of lack of labels in speech organ data sets is a problem that technical personnel in this field urgently need to solve. Summary of the invention
[0006] In view of this, in order to solve the problems in the background technology, the present invention provides an organ recognition method based on speech emotions and speech features, aiming to explore the intrinsic correlation patterns between emotions in speech signals and organs, construct potential emotion-organ mapping, and thus achieve more accurate organ recognition.
[0007] In order to achieve the above object, the present invention adopts the following technical solution:
[0008] A method for identifying organs based on voice emotion and voice features, comprising:
[0009] Acquire the voice viscera features, output the corresponding viscera classification probability through the viscera classifier, and identify the viscera organs according to the viscera classification probability; wherein,
[0010] The viscera classifier is trained based on the speech emotion feature data and the speech viscera feature data, and the training process includes:
[0011] Extract the common emotional features of speech emotion features and speech viscera features, and obtain the viscera speech relationship embedding features after weighted fusion;
[0012] Obtain the soft labels of the visceral voice relationship embedding features, and preliminarily train the visceral classifier according to the soft labels and the true labels corresponding to the voice emotion features;
[0013] Extract the deep features of the voice emotion features and the voice visceral features, perform consistency constraints based on the Euclidean space and the hyperbolic space, and simultaneously combine the deep features, the visceral voice relationship embedding features, and the consistency constraints to perform secondary training on the preliminarily trained visceral classifier.
[0014] Preferably, the voice emotion feature data and the voice visceral feature data are obtained by using a multi-scale voice encoder to respectively extract features from the voice emotion data and the voice visceral data;
[0015] The multi-scale voice encoder includes three parallel two-dimensional convolutional layers and a three-layer deep convolutional network;
[0016] The two-dimensional convolutional layer is used to extract data features from the time, frequency, and space dimensions;
[0017] The three-layer deep convolutional network is used to further extract multi-scale voice features according to the data features of different dimensions.
[0018] Preferably, a dual discriminator is used to constrain and update the parameters of the multi-scale voice encoder. The dual discriminator includes a domain discriminator and a batch discriminator.
[0019] The constraint method of the domain discriminator is:
[0020] G d (G f ; θ d ) = GRL(softmax(G f (x); θ d ))
[0021]
[0022] In the formula, G d is the domain discriminator, G f represents the multi-scale voice encoder, θ d represents the calculation parameters of the domain discriminator, L d is the cross-entropy loss of the domain discriminator, y d is the true domain category of the domain discrimination;
[0023] The constraint method of the batch discriminator is:
[0024] G b (G f ; θ b ) = softmax(G f (x); θb )
[0025]
[0026] In the formula, G b is the batch discriminator, and G f represents the multi-scale speech encoder, and θ b represents the calculation parameters of the batch discriminator, and L b is the cross-entropy loss of the batch discriminator, and y b is the true domain category of the batch discrimination.
[0027] Preferably, extract the emotional common features of the speech emotion features and the speech zang-fu features, and obtain the zang-fu speech relationship embedding features after weighted fusion; the steps include:
[0028] Calculate the batch attention based on the speech emotion features and the speech zang-fu features;
[0029] Determine the attention weights of the hidden state according to the hidden state of the batch attention;
[0030] Weight the attention weights to the hidden state of the batch attention to obtain the zang-fu speech relationship embedding features.
[0031] Preferably, obtain the soft labels of the fu speech relationship embedding features according to the following formula;
[0032]
[0033] In the formula, S represents the dot product similarity matrix of the normalized speech emotion features and the fu speech relationship embedding features, B s represents the set of speech emotion feature vectors, B r represents the set of fu speech relationship embedding feature vectors, target represents the correct label of the speech emotion features, and Y′ (i) represents the soft label of the fu speech relationship embedding features, and i represents the i-th dimension data of the soft label.
[0034] Preferably, the loss function for preliminary training is:
[0035]
[0036] In the formula, N represents the number of samples of the speech emotion features, M represents the number of samples of the speech zang-fu features, α, β, and λ represent regularization parameters, X s represents the speech emotion features, X r represents the fu speech relationship embedding features, X t represents the speech zang-fu features, Y represents the true class label corresponding to the feature X s and Y′ represents the soft label of the fu speech relationship embedding features.
[0037] Preferably, deep features of speech emotion features and speech zang-fu features are extracted separately;
[0038] The extraction steps include:
[0039] Divide the features into feature blocks and linearly project them into a high-dimensional feature space to obtain high-dimensional feature blocks,
[0040] Perform binary segmentation on the high-dimensional feature blocks from the channel dimension to obtain multi-dimensional feature one and multi-dimensional feature two;
[0041] Enable cross-dimensional linear feature interaction between multi-dimensional feature one and multi-dimensional feature two in the row and column directions, and then obtain deep features according to the following formula;
[0042]
[0043] In the formula, X m1 represents multi-dimensional feature one, X m2 represents multi-dimensional feature two, X m represents the high-dimensional feature block, represents the finally extracted feature.
[0044] Preferably, the process of cross-dimensional linear feature interaction is expressed as follows:
[0045] X (1) *,j = W1X *,j ; X (1) i,* = W2X i,*
[0046] X *,j = W3Cat(X (1) *,j-1 , X (1) *,j , X (1) *,j+1 )
[0047] X i,* = W4Cat(X (1) i-1* , X (1) i,* X (1) i+1,* )
[0048] In the formula, W1, W2, W3, and W4 are weight parameters of the linear layer, X *,j represents the feature vector of the j-th column, X i,j represents the element of the i-th row and the j-th column, Cat represents the tensor dimension concatenation operation, X *,j-1 represents all elements of the (j-1)-th column, X *,j+1Denote all elements in the (j + 1)-th column, X i-1,* Denote all elements in the (i - 1)-th row, X i+1,* Denote all elements in the (i + 1)-th row.
[0049] Preferably, the deep features of the speech emotion feature and the speech zang-fu feature are respectively mapped to the Euclidean space and the hyperbolic space, and consistency constraints are performed according to the cosine distance of the mapped features in the Euclidean space and the hyperbolic space; the formula is expressed as:
[0050]
[0051] Among them, and respectively represent the mapped features of the speech emotion feature in the Euclidean space and the hyperbolic space, and respectively represent the mapped features of the speech zang-fu feature in the Euclidean space and the hyperbolic space; and D es represents the distance between the features and in the Euclidean space, represents the distance between the features and in the hyperbolic space; the calculation formula is as follows:
[0052]
[0053] In the formula, c represents the curvature parameter, and h represents the radius parameter.
[0054] Preferably, the loss function of the secondary training is:
[0055]
[0056] In the formula, represents the deep feature of the speech emotion feature, represents the deep feature of the speech zang-fu feature, Y e and Y p are respectively corresponding true labels, X r represents the zang-fu speech relationship embedding feature, Y′ represents the true label corresponding to X r α, β, and λ are regularization parameters.
[0057] Through the above technical solutions, it can be known that the present invention discloses a zang-fu recognition method based on speech emotion and speech features. Compared with the prior art, it has the following beneficial effects:
[0058] 1) A large amount of labeled speech emotion data can be fully utilized to address practical problems such as slow annotation progress and lack of labels in the initial stage of constructing the speech zang-fu dataset;
[0059] 2) By means of the spatial consistency constraint loss of the emotional viscera, aligning the emotional features with the viscera features can significantly improve the recognition performance of the voice viscera;
[0060] 3) Applying the traditional Chinese medicine emotional theory to the viscera recognition based on voice improves the performance and verifies the correctness and application value of the traditional Chinese medicine emotional theory. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0062] Figure 1 It is a flowchart of the viscera recognition method based on voice emotion and voice features provided by the present invention;
[0063] Figure 2 It is a schematic structural diagram of the multi-scale voice encoder provided by the present invention;
[0064] Figure 3 It is a schematic structural diagram of the voice relationship attention provided by the present invention;
[0065] Figure 4 It is a schematic structural diagram of the voice relationship learning under the double discriminator structure provided by the present invention;
[0066] Figure 5 It is an overall schematic diagram of the cross-domain voice relationship learning method provided by the present invention;
[0067] Figure 6 It is a schematic structural diagram of the stacked dimensional multi-layer perceptron provided by the present invention;
[0068] Figure 7 It is a schematic diagram of the training process of the viscera classifier provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0069] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0070] To make full use of a large amount of labeled voice emotion data, the embodiments of the present invention disclose a viscera recognition method based on voice emotion and voice features; asFigure 1 , including:
[0071] Obtain voice zang-fu organ features, output corresponding zang-fu organ classification probabilities through a zang-fu organ classifier, and identify zang-fu organs according to the zang-fu organ classification probabilities;
[0072] In this embodiment, the zang-fu organ classifier considers voice emotion feature data and is trained in combination with voice zang-fu organ feature data. The training process includes:
[0073] Extract the emotional common features of voice emotion features and voice zang-fu organ features, and obtain zang-fu organ voice relationship embedding features after weighted fusion;
[0074] Obtain the soft labels of the zang-fu organ voice relationship embedding features, and preliminarily train the zang-fu organ classifier according to the soft labels and the true labels corresponding to the voice emotion features;
[0075] Extract the deep features of voice emotion features and voice zang-fu organ features, perform consistency constraints based on Euclidean space and hyperbolic space, and simultaneously combine the deep features, zang-fu organ voice relationship embedding features and consistency constraints to perform secondary training on the preliminarily trained zang-fu organ classifier.
[0076] The present invention innovatively proposes to train a zang-fu organ classifier by considering voice emotion features, thereby improving the recognition ability of the zang-fu organ classifier. Among them, the training process of this application is mainly divided into two parts, and the training process will be described below through embodiments.
[0077] The first step,
[0078] Extract the emotional common features of voice emotion features and voice zang-fu organ features, and obtain zang-fu organ voice relationship embedding features after weighted fusion; extract the soft labels of voice emotion features and zang-fu organ voice relationship embedding features, and preliminarily train the zang-fu organ classifier according to the soft labels.
[0079] S1. In this embodiment, first obtain voice emotion data and voice zang-fu organ data. The present application independently constructs 5 types of voice zang-fu organ data sets with zang-fu organ labels, and the sample distribution is heart (890), liver (861), spleen (3900), lung (2482), kidney (1866).
[0080] For voice emotion data, the publicly released CASIA Chinese Emotion Corpus is selected, which contains 5 kinds of emotion samples (anger, sadness, fear, happiness and surprise) corresponding to zang-fu organs (heart, liver, spleen, lung and kidney).
[0081] Then, perform feature extraction through a multi-scale voice encoder to obtain voice emotion features and voice zang-fu organ features.
[0082] In this embodiment, the multi-scale speech encoder includes three parallel two-dimensional convolutional layers and a three-layer deep convolutional network;
[0083] The two-dimensional convolutional layer is used to extract data features from the time, frequency, and spatial dimensions;
[0084] The three-layer deep convolutional network is used to further extract multi-scale speech features based on the data features of different dimensions.
[0085] In one embodiment, as Figure 2 shown, the multi-scale speech encoder includes three parallel two-dimensional convolutional layers with different convolutional kernel sizes, aiming to extract deep time, frequency, and spatial variation patterns from the primary speech features and simultaneously extract multiple fine-grained information at the same level. Subsequently, the features of different fine-grained levels will be stacked along the channel dimension and fed into a three-layer deep convolutional network to further extract multi-scale speech features. The multi-scale speech encoder will finally output a 512-dimensional deep speech feature representation.
[0086] Preferably, in this application, the spectrogram is used as the primary speech feature, and the 128-dimensional logarithmic Mel (Log-Mel) spectrum and the 40-dimensional Mel-frequency cepstral coefficients (MFCC) are respectively extracted as the primary speech features.
[0087] S2. Extract the emotional common features of the speech emotion features and the viscera speech features, and obtain the viscera speech relationship embedding features after weighted fusion; the steps include:
[0088] Calculate the batch attention based on the speech emotion features and the speech viscera features;
[0089] Determine the attention weights of the hidden state according to the hidden state of the batch attention;
[0090] Weight the attention weights to the hidden state of the batch attention to obtain the viscera speech relationship embedding features.
[0091] This application discovers the potential relationship between emotions and viscera based on speech relationship attention, extracts the emotional common features of speech emotion features and viscera speech features, and generates a speech relationship embedding that mixes the relationship between speech emotion and speech viscera.
[0092] During the model training stage, the speech relationship attention can perform multi-level feature relationship mining on the input batch data to maximize the revelation and integration of the potential emotional relationships contained therein. Refer to Figure 3 , which mainly includes batch attention and hidden state attention;
[0093] Furthermore, using the speech emotion feature data as the source domain training data, it is defined as:
[0094]
[0095] Taking the voice viscera feature data as the target domain training data, it is defined as:
[0096]
[0097] Then the batch attention can be expressed as:
[0098] B = unsqueeze({B s ,B t ) ∈ R 2b×1×d
[0099]
[0100] unsqueeze refers to the expansion of the tensor dimension, aiming to expand the dimension of the original training batch B ∈ R 2b×d to meet the calculation of dot product attention. b is the number of samples, and d is the number of features of the samples.
[0101] In one embodiment, to effectively capture and integrate the emotional relationship information at different levels and scales, an additional hidden state attention mechanism is applied to the hidden state of the batch attention, which is expressed as:
[0102]
[0103] X r = A h H
[0104] In the formula, it refers to the set of all hidden states output by the batch attention calculation, which has met the requirements of dot product attention. n represents the number of hidden states, 2b represents the number of samples, and d represents the feature dimension.
[0105] By performing a global weighted sum on all hidden states to capture the long-range and short-range dependence relationships and cross-scale interaction degrees between positions, it is possible to effectively extract the key emotional common features from the feature representations X i at different scales, and finally linearly weighted fuse to obtain the voice relationship embedding X r .
[0106] In one embodiment, due to the large differences in the data feature distributions between different corpora, it is difficult for the model to capture the optimal emotional relationship features, resulting in a performance decline when dealing with the data imbalance problem. For this reason, a dual discriminator structure is proposed based on the voice relationship attention;
[0107] Referring to Figure 4, the parameters of the multi-scale speech encoder are constrained and updated using a dual discriminator. The dual discriminator includes a domain discriminator and a batch discriminator. First, a domain discriminator structure with a gradient reversal layer is applied, aiming to preliminarily align the feature distributions between two domains, thereby minimizing the impact of feature distribution differences on the calculation of speech relationship attention. The gradient reversal layer multiplies the gradient from the domain discriminator by a negative constant to achieve the "reversal" of the gradient.
[0108] In this embodiment, the constraint method of the domain discriminator is as follows:
[0109] G d (G f ; θ d ) = GRL(softmax(G f (x); θ d ))
[0110]
[0111] In the formula, G d is the domain discriminator, which is a linear binary classifier in this embodiment, aiming to learn a logistic regression G d : R D → [0, 1] to determine whether the given input is from the source domain or the target domain; G f represents the multi-scale speech encoder, and θ d represents the calculation parameters of the domain discriminator. L d is the cross-entropy loss of the domain discriminator, preferably binary cross-entropy loss, used to measure the difference between the domain category predicted by the domain discriminator and the true domain category. y d is the true domain category of the domain discrimination.
[0112] The batch discriminator discriminates the speech relationship embedding output by the speech relationship attention, guiding the model to moderately adjust the differences between samples in different domains while retaining the similarity of samples within the domain according to the following formula, so that it remains within a reasonable range.
[0113] G b (G f ; θ b ) = softmax(G f (x); θ b )
[0114]
[0115] In the formula, G b is the batch discriminator, used to learn a logistic regression to determine whether the given input is from the source domain batch or the speech relationship embedding. G f represents the multi-scale speech encoder, and θb Represents the calculation parameters of the batch discriminator, l b Is the cross-entropy loss of the batch discriminator, y b Is the true domain category of the batch discrimination.
[0116] S3. Obtain the soft labels of the voice-organ relationship embedding features,
[0117] The batch soft label strategy uses the similarity between each source domain feature and the voice-organ relationship embedding as the confidence threshold of the soft label, creating an adaptive soft classification boundary for each voice-organ relationship embedding to encourage the voice-organ relationship embedding to form a tight clustering with the same-category features within the source domain.
[0118] This application extracts the soft labels of the voice emotion features and the voice-organ relationship embedding features respectively according to the following formula;
[0119]
[0120] In the formula, S represents the dot product similarity matrix of the normalized voice emotion features and the voice-organ relationship embedding features, B s Represents the set of voice emotion feature vectors, B r Represents the set of voice-organ relationship embedding feature vectors, target represents the correct label of the voice emotion features, Y′ (i) Represents the soft label of the voice-organ relationship embedding features.
[0121] Y′ (i) Refers to the i-th dimension of the one-hot vector label of the voice-organ relationship embedding. If a source domain sample belongs to the i-th category, the value of the i-th dimension of the one-hot vector label corresponding to its voice-organ relationship embedding is set to S, otherwise its value is Where C represents the number of categories, and target is the correct label of the source domain.
[0122] Furthermore, to solve the problem of inconsistency in the processing flow between the training and testing phases, this application introduces a shared classifier strategy, as Figure 5 shown. The core idea is that in the training phase, the potential voice relationships captured through voice relationship learning are fully utilized and integrated into the learning process of the shared classifier, which can indirectly improve the recognition ability of the shared classifier for the test features that have not undergone voice relationship learning processing.
[0123] In one embodiment, the loss function of the preliminary training is expressed as:
[0124]
[0125] In the formula, N represents the number of samples of the voice emotion features, M represents the number of samples of the voice-organ features, α, β, λ represent regularization parameters, X sDenote the speech emotion feature as X r Denote the viscera-speech relationship embedding feature as X t The speech viscera feature, where Y denotes the feature X s The corresponding true class label, and Y′ denotes the soft label of the viscera-speech relationship embedding feature.
[0126] Step 2:
[0127] To further optimize the above technical solution, the present application extracts the deep features of the speech emotion feature and the viscera-speech feature, performs consistency constraints based on the Euclidean space and the hyperbolic space, and simultaneously combines the deep features, the viscera-speech relationship embedding feature, and the consistency constraint to perform secondary training on the preliminarily trained viscera classifier.
[0128] In one embodiment, the extraction steps of the deep features include:
[0129] Divide the feature into feature blocks and linearly project them into a high-dimensional feature space to obtain high-dimensional feature blocks.
[0130] Bisect the high-dimensional feature blocks from the channel dimension to obtain multi-dimensional feature one and multi-dimensional feature two;
[0131] Perform cross-dimensional linear feature interaction on multi-dimensional feature one and multi-dimensional feature two in the row and column directions, and then obtain deep features through channel fusion.
[0132] The execution process of the above steps can refer to Figure 6 , Figure 6 which is a stacked dimensional multi-layer perceptron structure. Using the dimensional multi-layer perceptron structure to perform deeper encoding on the emotional speech and the viscera speech can extract more accurate features.
[0133] The present application converts the speech signal into a spectrogram as the input. The dimensional multi-layer perceptron first represents the input features divides them into a series of feature blocks, denoted as where F and T respectively represent the frequency and time dimensions of the feature blocks, the number of channels is 1, and the block size is p×p.
[0134] Immediately afterwards, the dimensional multi-layer perceptron linearly projects all the feature blocks into a high-dimensional feature space C to obtain high-dimensional feature blocks Subsequently, bisect the feature blocks in the channel dimension and perform cross-dimensional linear feature interaction in their row and column directions respectively.
[0135] The above process can be expressed as:
[0136]
[0137] X m1 , Xm2 = MLP ft (Transpos(Split c (X m )))
[0138]
[0139] In the formula, MLP c represents a channel perceptron for linearly projecting a two-dimensional feature block into a high-dimensional space. MLP ft represents a cross-dimensional linear interaction perceptron to achieve feature interaction in the time and frequency dimensions for the feature block after spatial transposition.
[0140] In this embodiment, the dimensional multi-layer perceptron structure introduces a cascaded interaction form in feature interaction. Through weighted fusion of linear layers, the block features in each row (or column) not only serve the feature aggregation of the current row (or column), but also contribute to the feature aggregation of adjacent rows (or columns), achieving cross-dimensional normalization, that is
[0141] X (1) *,j = W1X *,j ; X (1) i,* = W2X i,*
[0142] X *,j = W3Cat(X (1) *,j-1 , X (1) *,j , X (1) *,j+1 )
[0143] X i,* = W4Cat(X (1) i-1,* X (1) i,* , X (1) i+1,* )
[0144] In the formula, W1, W2, W3, and W4 are the weight parameters of the linear layer, X *,j represents the feature vector of the j-th column, X i,j represents the element of the i-th row and the j-th column, Cat represents the tensor dimension concatenation operation, X *,j-1 represents all the elements of the (j - 1)-th column, X *,j+1 represents all the elements of the (j + 1)-th column, X i-1,* represents all the elements of the (i - 1)-th row, X i+1,* represents all the elements of the (i + 1)-th row.
[0145] Further, normalize X separately according to the above method m1 and X m2 to achieve cross-dimensional normalization.
[0146] Then, to effectively aggregate spatial information and channel information, the model applies a single convolutional layer between the stacked dimensional multi-layer perceptrons, causing the spatial dimension of the input features to decrease by 2×2 layer by layer while the channel dimension doubles, gradually downsampling the features from to and using residual connections to avoid the problems of vanishing gradients and exploding gradients. Finally, the dimensional multi-layer perceptron inputs the depth aggregation features extracted by multiple blocks into a global average pooling layer for dimensionality reduction, and the formula is expressed as:
[0147]
[0148] In the formula, X m1 represents multi-dimensional feature one, X m2 represents multi-dimensional feature two, X m represents the high-dimensional feature block, represents the finally extracted feature.
[0149] In a preferred embodiment, as Figure 7 shown, by calculating the emotional viscera space consistency constraint loss, the emotional features and viscera features are constrained and aligned in multiple spaces.
[0150] There is a close association between human emotions and the human viscera. Then, in the context of multi-space embedding, emotions and viscera can be respectively mapped into feature vectors to explore their relationships in different geometric spaces.
[0151] Suppose in the Euclidean space, the distance between the emotion vector x e and the viscera vector x p is D es , and in the hyperbolic space, the distance between the emotion vector x e and the viscera vector x p is If the ratio or absolute difference between D es and is within a small threshold range, then it is considered that there is a high spatial consistency between x e and x p .
[0152] In this embodiment, let the mappings of the emotional features and viscera features extracted by the dimensional multi-layer perceptron in the Euclidean space be and and the mappings in the hyperbolic space be and
[0153] Among them, the mapping party is: the mapping function from the Euclidean space to the hyperbolic space is mapping, and the mapping function from the hyperbolic space to the Euclidean space is mapping. The specific expression of the mapping function is:
[0154]
[0155] In the formula, v is the reference point in the space, is a special pair operation, represents the Möbius addition, that is:
[0156]
[0157] In the formula, c is the space curvature constant, usually less than 0, and is uniformly taken as -1 in this article.
[0158] Furthermore, this embodiment performs consistency constraint according to the cosine distance of the mapped features on the Euclidean space and the hyperbolic space; the formula is expressed as:
[0159]
[0160] Among them, and respectively represent the mapped features of the speech emotion features in the Euclidean space and the hyperbolic space, and respectively represent the mapped features of the speech zang-fu features in the Euclidean space and the hyperbolic space; and D es represents the feature and in the Euclidean space,
[0161] represents the feature and in the hyperbolic space; the calculation formula is as follows:
[0162]
[0163] This application takes the cross entropy as the loss function, and the loss function of the secondary training is:
[0164]
[0165] In the formula, represents the deep feature of the speech emotion feature, represents the deep feature of the speech zang-fu feature, Y e and Y p are respectively corresponding true labels, X r represents the zang-fu speech relationship embedding feature, and Y′ represents Xr The corresponding true label, where α, β, and λ are regularization parameters.
[0166] Among them, the loss function of cross-entropy has the following expression:
[0167]
[0168] Among them, y p is the true label of the input sample, while y' i is the predicted label of the classifier.
[0169] When the viscera classifier is trained through the above steps, it can directly identify the viscera classification based on the voice viscera data. Specifically, during the actual application test, only the viscera voice is used as the input, and through the multi-scale voice encoder and the dimensional multi-layer perceptron structure, the multi-label viscera classification probability is obtained by the shared viscera classifier.
[0170] In this application, the above constraints are designed to implicitly guide the model during the learning process to not only pursue good performance in a single space but also ensure that the distances between the emotion vectors and the viscera vectors in multiple spaces are consistent. It is expected that in this way, the association between emotion and viscera can be understood from the perspective of multiple spaces, improving the generalization ability of the model.
[0171] In this specification, each embodiment is described in a progressive manner. Each embodiment focuses on the differences from other embodiments, and the same or similar parts among the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0172] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A viscera recognition method based on speech emotion and speech features, characterized in that Speech viscera features are obtained, and the corresponding viscera classification probabilities are output through a viscera classifier, and the viscera organs are recognized according to the viscera classification probabilities; among them, The viscera classifier is trained based on speech emotion feature data and speech viscera feature data, and the training process includes: Extract the emotional common features of speech emotion features and speech viscera features, and obtain the viscera-speech relationship embedding features after weighted fusion; Obtain the soft labels of the viscera-speech relationship embedding features, and preliminarily train the viscera classifier according to the soft labels and the true labels corresponding to the speech emotion features; Extract the deep features of speech emotion features and speech viscera features, perform consistency constraints based on Euclidean space and hyperbolic space, and at the same time jointly use the deep features, viscera-speech relationship embedding features and consistency constraints to perform secondary training on the preliminarily trained viscera classifier.
2. The viscera recognition method based on voice emotion and voice features according to claim 1, wherein, The speech emotion feature data and speech viscera feature data are obtained by using a multi-scale speech encoder to extract features from speech emotion data and speech viscera data respectively; The multi-scale speech encoder includes three parallel two-dimensional convolutional layers and a three-layer deep convolutional network; The two-dimensional convolutional layer is used to extract data features from the time, frequency and space dimensions; The three-layer deep convolutional network is used to further extract multi-scale speech features according to the data features of different dimensions.
3. The viscera recognition method based on voice emotion and voice features according to claim 2, wherein, Use a dual discriminator to constrain and update the parameters of the multi-scale speech encoder. The dual discriminator includes a domain discriminator and a batch discriminator. The constraint method of the domain discriminator is: G d (G f ; θ d ) = GRL(softmax(G f (x); θ d )) Where G d is the domain discriminator, and G f represents the multi-scale speech encoder. θ d represents the calculation parameters of the domain discriminator. L d is the cross-entropy loss of the domain discriminator, and y d is the true domain category of the domain discrimination; The constraint method of the batch discriminator is: G b (G f ; θ b ) = softmax(G f (x); θ b ) Where G b is the batch discriminator, and G f represents the multi-scale speech encoder, θ b represents the calculation parameters of the batch discriminator, L b is the cross-entropy loss of the batch discriminator, and y b is the true domain category of the batch discrimination.
4. A viscera recognition method based on speech emotion and speech features according to claim 1, characterized in that, Extract the emotional common features of speech emotion features and speech viscera features, and obtain the viscera-speech relationship embedding features after weighted fusion; the steps include: Calculate the batch attention based on the speech emotion features and speech viscera features; Determine the attention weights of the hidden states according to the hidden states of the batch attention; Weight the hidden states of the batch attention with the attention weights to obtain the viscera-speech relationship embedding features.
5. A viscera recognition method based on speech emotion and speech features according to claim 1, characterized in that Obtain the soft labels of the viscera-speech relationship embedding features according to the following formula; where S represents the dot product similarity matrix of the normalized speech emotion features and the abdominal speech relationship embedding features, B s represents the set of speech emotion feature vectors, B r represents the set of abdominal speech relationship embedding feature vectors, target represents the correct label of the speech emotion feature, Y′ (i) represents the soft label of the abdominal speech relationship embedding feature, and i represents the i-th dimension data of the soft label.
6. The viscera recognition method based on voice emotion and voice features according to claim 3, wherein, The loss function of the preliminary training is: where N represents the number of samples of speech emotion features, M represents the number of samples of speech viscera features, α, β, and λ represent regularization parameters, and X s represents speech emotion features, and X r represents the embedding feature of the relationship between the fu-organ and speech, and X t is the speech viscera feature, Y represents the true class label corresponding to feature X s , and Y' represents the soft label of the embedding feature of the relationship between the fu-organ and speech.
7. A viscera recognition method based on voice emotion and voice characteristics according to claim 1, characterized in that Extract the deep features of speech emotion features and speech viscera features respectively; The extraction steps include: Divide the features into feature blocks and linearly project them into a high-dimensional feature space to obtain high-dimensional feature blocks, Bisect the high-dimensional feature blocks from the channel dimension to obtain multi-dimensional feature one and multi-dimensional feature two; Perform cross-dimensional linear feature interaction on multi-dimensional feature one and multi-dimensional feature two in the row and column directions, and then obtain the deep features according to the following formula; where X m1 represents the multi-dimensional feature one, X m2 represents the multi-dimensional feature two, X m represents the high-dimensional feature block, represents the finally extracted feature.
8. A viscera recognition method based on voice emotion and voice features according to claim 7, characterized in that The process of cross-dimensional linear feature interaction is shown as follows: X (1) *,j = W1X *,j ; X (1) i,* = W2X i,* X *,j = W3Cat(X (1) *,j-1 , X (1) *,j , X (1) *,j+1 X i,* = W4Cat(X (1) i-1,* , X (1) i,* , X (1) i+1,* ) Where, W1, W2, W3, and W4 are the weight parameters of the linear layer, and X *,j represents the feature vector of the j-th column, and X i,j represents the element in the i-th row and the j-th column. Cat represents the tensor dimension concatenation operation, and X *,j-1 represents all elements of the (j-1)-th column, and X *,k+1 represents all elements of the (j+1)-th column, and X i-1,* represents all elements of the (i-1)-th row, and X i+1,* represents all elements of the (i+1)-th row.
9. A viscera recognition method based on voice emotion and voice features according to claim 1, characterized in that Map the deep features of speech emotion features and speech viscera features to Euclidean space and hyperbolic space respectively, and perform consistency constraints according to the cosine distance of the mapped features in Euclidean space and hyperbolic space; the formula is expressed as: Among them, and respectively represent the mapping features of speech emotion features in the Euclidean space and the hyperbolic space, and respectively represent the mapping features of speech zang-fu features in the Euclidean space and the hyperbolic space; and D es represents the feature and in the Euclidean space, represents the feature and in the hyperbolic space; the calculation formula is as follows: In the formula, c represents the curvature parameter and h represents the radius parameter.
10. A viscera recognition method based on voice emotion and voice features according to claim 9, characterized in that The loss function of the secondary training is: In the formula, represents the deep feature of the speech emotion feature, represents the deep feature of the speech zang-fu feature, Y e and Y p are respectively the corresponding true labels, X r represents the zang-fu speech relationship embedding feature, Y' represents X r the corresponding true label, and α, β, λ are regularization parameters.
Citation Information
Patent Citations
Traditional Chinese medicine sound smelling diagnosis automatic system supported by smart voice technology
CN112002342A
Hyperbolic space alignment-based multi-modal voice internal organ recognition method
CN117958765A
Voice signal analysis method and device, equipment and medium
CN118969025A
Converting a sequence of speech records of a human subject into a sequence of indicators of a physiological state of the subject
US20240386906A1