A method for recognizing viscera based on voice emotion and voice features

Through the organ recognition method based on speech emotions and speech features, the organ classifier is trained using a multi-scale speech encoder and a dual discriminator, which solves the problems of slow data set annotation and lack of labels in organ recognition, improves the accuracy and performance of organ recognition, and verifies the application value of traditional Chinese medicine emotion theory.

CN120340535BActive Publication Date: 2025-10-17SOUTH CHINA UNIV OF TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510354505.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-10-17
Estimated Expiration
2045-03-25

AI Technical Summary

Technical Problem

Existing technologies make it difficult to effectively utilize annotated speech emotion data for organ recognition. The lack of large-scale annotated speech organ datasets leads to insufficient organ recognition performance.

Method used

An organ recognition method based on speech emotion and speech features is adopted. Features are extracted through a multi-scale speech encoder, combined with speech relation attention and dual discriminator training. The spatial consistency constraint loss of emotional organs is used to construct an organ classifier for organ recognition.

Benefits of technology

By making full use of emotional speech data, the accuracy and performance of organ recognition are improved, the application value of traditional Chinese medicine emotion theory is verified, and the problems of slow data set annotation and lack of labels are solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340535B_ABST
    Figure CN120340535B_ABST
Patent Text Reader

Abstract

The application discloses a viscera recognition method based on voice emotion and voice features, mainly trains a viscera classifier based on voice emotion feature data and voice viscera feature data; wherein, the training steps include extracting emotional common features of voice emotion features and voice viscera features, obtaining viscera voice relationship embedding features after weighted fusion; extracting soft labels of voice emotion features and viscera voice relationship embedding features, and preliminarily training the viscera classifier according to the soft labels; extracting deep features of voice emotion features and voice viscera features, performing consistency constraint based on Euclidean space and hyperbolic space, and simultaneously combining the deep features, viscera voice relationship embedding features and consistency constraint to perform secondary training on the preliminarily trained viscera classifier. The viscera recognition method in the application can fully utilize a large amount of voice emotion data with labels, and can improve the recognition performance of viscera.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of viscera recognition, and more particularly to a viscera recognition method based on voice emotion and voice features. BACKGROUND

[0002] In recent years, the concept of traditional Chinese medicine health care and individualized conditioning methods have received widespread attention and application. Among them, voice, as one of the important manifestations of human viscera, can provide important reference for formulating individualized conditioning programs by analyzing the corresponding viscera and viscera characteristics.

[0003] Since a voice may correspond to multiple viscera, the viscera recognition based on voice is a multi-label classification problem. At present, data-driven machine learning is the main method to solve this problem, but it requires a large amount of labeled voice data of viscera, and it is difficult to obtain a large-scale voice dataset labeled with viscera at this stage.

[0004] However, there are currently a large number of emotion voice datasets labeled with emotions. The theory of traditional Chinese medicine emotion believes that the emotional changes of the human body such as joy, anger, thought, sadness, and fear are closely related to the viscera of the human body such as the heart, liver, spleen, lung, and kidney, and influence each other.

[0005] Therefore, how to make full use of these emotional voice data to alleviate the problem of lack of voice viscera dataset labels is a problem that needs to be solved by those skilled in the art. SUMMARY

[0006] Therefore, in order to solve the problems in the background art, the present application provides a viscera recognition method based on voice emotion and voice features, which aims to mine the internal correlation pattern between emotions and viscera in voice signals, and construct a potential emotion-viscera mapping, so as to realize more accurate viscera recognition.

[0007] In order to achieve the above purpose, the present application adopts the following technical solutions:

[0008] A viscera recognition method based on voice emotion and voice features, comprising,

[0009] Obtaining voice viscera features, outputting corresponding viscera classification probabilities through a viscera classifier, and identifying viscera organs according to the viscera classification probabilities; wherein,

[0010] The viscera classifier is trained based on voice emotion feature data and voice viscera feature data, and the training process includes:

[0011] Extracting emotional common features of voice emotion features and voice viscera features, and obtaining viscera voice relationship embedding features after weighted fusion;

[0012] The soft label of the visceral voice relationship embedding feature is obtained, and the visceral classifier is preliminarily trained according to the soft label and a real label corresponding to the voice emotion feature;

[0013] Deep features of the voice emotion feature and the voice visceral feature are extracted, and consistency constraints are performed based on Euclidean space and hyperbolic space, and the preliminarily trained visceral classifier is secondarily trained based on the deep features, the visceral voice relationship embedding feature and the consistency constraints.

[0014] Preferably, the voice emotion feature data and the voice visceral feature data are obtained by using a multi-scale voice encoder to extract features of voice emotion data and voice visceral data, respectively.

[0015] The multi-scale voice encoder includes three parallel two-dimensional convolution layers and a three-layer deep convolution network.

[0016] The two-dimensional convolution layer is configured to extract data features from time, frequency and space dimensions.

[0017] The three-layer deep convolution network is configured to further extract multi-scale voice features based on the data features of different dimensions.

[0018] Preferably, the parameters of the multi-scale voice encoder are updated by using a double discriminator, and the double discriminator includes a domain discriminator and a batch discriminator.

[0019] The constraint mode of the domain discriminator is as follows:

[0020] G d (G f ;θ d )=GRL(softmax(G f (x);θ d ))

[0021]

[0022] In the formula, G d is the domain discriminator, G f represents the multi-scale voice encoder, θ d represents the calculation parameters of the domain discriminator, L d is the cross-entropy loss of the domain discriminator, and y d is the real field category of the domain discriminator.

[0023] The constraint mode of the batch discriminator is as follows:

[0024] G b (G f ;θ b )=softmax(G f (x);θb

[0025]

[0026] where G b is the batch discriminator, G f represents a multi-scale speech encoder, θ b represents the calculation parameters of the batch discriminator, L b is the cross-entropy loss of the batch discriminator, y b is the true field category of the batch discrimination.

[0027] Preferably, the extracted emotional common features of the speech emotion features and the speech zang-visceral features are weighted and fused to obtain the zang-visceral speech relationship embedding features; the step comprises:

[0028] The batch attention is calculated based on the speech emotion features and the speech zang-visceral features;

[0029] According to the hidden state of the batch attention, the attention weight of the hidden state is determined;

[0030] The attention weight is weighted to the hidden state of the batch attention to obtain the zang-visceral speech relationship embedding features.

[0031] Preferably, the soft label of the zang-visceral speech relationship embedding features is obtained according to the following formula:

[0032]

[0033] where S represents a dot product similarity matrix of the normalized speech emotion features and the zang-visceral speech relationship embedding features, B s represents a set of speech emotion feature vectors, B r represents a set of zang-visceral speech relationship embedding feature vectors, target represents a correct label of the speech emotion features, Y′ (i) represents a soft label of the zang-visceral speech relationship embedding features, and i represents the i-th dimension data of the soft label.

[0034] Preferably, the loss function of the preliminary training is:

[0035]

[0036] where N represents the sample quantity of the speech emotion features, M represents the sample quantity of the speech zang-visceral features, α, β, λ represent regularization parameters, X s represents the speech emotion features, X r represents the zang-visceral speech relationship embedding features, X t represents the speech zang-visceral features, Y represents the corresponding true category label of the features X s , and Y′ represents a soft label of the zang-visceral speech relationship embedding features.​

[0037] Preferably, the deep-level features of the speech emotion features and the speech Zangfu features are extracted respectively;

[0038] The extraction step comprises:

[0039] The features are divided into feature blocks and linearly projected into a high-dimensional feature space to obtain high-dimensional feature blocks,

[0040] The high-dimensional feature blocks are bisected from the channel dimension to obtain multi-dimensional feature one and multi-dimensional feature two;

[0041] The multi-dimensional feature one and the multi-dimensional feature two are subjected to cross-dimensional linear feature interaction in the row and column directions, and then the deep-level features are obtained according to the following formula:

[0042]

[0043] In the formula, X m1 represents the multi-dimensional feature one, X m2 represents the multi-dimensional feature two, X m represents the high-dimensional feature block, represents the final extracted features.

[0044] Preferably, the process of cross-dimensional linear feature interaction is represented as follows:

[0045] X (1) *,j = W1X *,j ; X (1) i,* = W2X i,*

[0046] X *,j = W3Cat(X (1) *,j-1 , X (1) *,j , X (1) *,j+1 )

[0047] X i,* = W4Cat(X (1) i-1* , X (1) i,* X (1) i+1,* )

[0048] In the formula, W1, W2, W3 and W4 are weight parameters of linear layers, X *,j represents the feature vector of the jth column, X i,j represents the element of the ith row and the jth column, Cat represents a tensor dimension splicing operation, X *,j-1 represents all elements of the j-1th column, X*,j+1 represents all elements of the j+1th column, X i-1,* represents all elements of the i-1th row, X i+1,* represents all elements of the i+1th row.

[0049] Preferably, the deep features of the voice emotion features and the voice zangfu features are respectively mapped to the Euclidean space and the hyperbolic space, and consistency constraints are performed according to the cosine distances of the mapped features in the Euclidean space and the hyperbolic space; the formula is represented as:

[0050]

[0051] wherein, and respectively represent the mapped features of the voice emotion features in the Euclidean space and the hyperbolic space, and respectively represent the mapped features of the voice zangfu features in the Euclidean space and the hyperbolic space; and D es represents the feature and in the Euclidean space, represents the feature and in the hyperbolic space; the calculation formula is as follows:

[0052]

[0053] In the formula, c represents a curvature parameter, and h represents a radius parameter.

[0054] Preferably, the loss function of the secondary training is:

[0055]

[0056] In the formula, represents the deep feature of the voice emotion feature, represents the deep feature of the voice zangfu feature, Y e and Y p respectively are corresponding true labels, X r represents the zangfu voice relationship embedding feature, Y' represents X r corresponding true labels, and α, β and λ are regularization parameters.

[0057] According to the technical solution, the application discloses a zangfu recognition method based on voice emotion and voice features, which has the following beneficial effects compared with the prior art:

[0058] 1) can make full use of a large number of labeled speech emotion data to deal with the slow labeling progress and label scarcity of the speech Zang organ data set construction in the early stage;

[0059] 2) With the help of the spatial consistency constraint loss of emotion Zang organs, the emotion features and Zang organ features are aligned, which can significantly improve the recognition performance of speech Zang organs;

[0060] 3) The TCM emotion theory is used for speech-based Zang organ recognition, which improves the performance and verifies the correctness and application value of the TCM emotion theory. BRIEF DESCRIPTION OF DRAWINGS

[0061] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of the provided drawings.

[0062] Figure 1 The Zang organ recognition method based on speech emotion and speech features provided by the present application is shown in the flow chart.

[0063] Figure 2 The multi-scale speech encoder structure schematic diagram provided by the present application is shown in the schematic diagram.

[0064] Figure 3 The speech relationship attention structure schematic diagram provided by the present application is shown in the schematic diagram.

[0065] Figure 4 The speech relationship learning structure schematic diagram under the double discriminator structure provided by the present application is shown in the schematic diagram.

[0066] Figure 5 The overall schematic diagram of the cross-domain speech relationship learning method provided by the present application is shown in the schematic diagram.

[0067] Figure 6 The stacked dimension multi-layer perceptron structure schematic diagram provided by the present application is shown in the schematic diagram.

[0068] Figure 7 The training process schematic diagram of the Zang organ classifier provided by the present application is shown in the schematic diagram. DETAILED DESCRIPTION

[0069] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0070] To make full use of a large amount of annotated speech emotion data, an organ recognition method based on speech emotion and speech feature is disclosed in the embodiment of the application. Figure 1 , comprising:

[0071] The speech organ feature is acquired, the corresponding organ classification probability is output by the organ classifier, and the organ is recognized according to the organ classification probability;

[0072] In the embodiment, the organ classifier considers the speech emotion feature data and is trained in combination with the speech organ feature data, and the training process comprises:

[0073] The emotional common features of the speech emotion feature and the speech organ feature are extracted, and the organ speech relationship embedding feature is obtained after weighted fusion;

[0074] The soft label of the organ speech relationship embedding feature is acquired, and the organ classifier is preliminarily trained according to the soft label and the real label corresponding to the speech emotion feature;

[0075] Deep features of the speech emotion feature and the speech organ feature are extracted, consistency constraint is performed based on the Euclidean space and the hyperbolic space, and the organ classifier preliminarily trained is secondarily trained based on the deep features, the organ speech relationship embedding feature and the consistency constraint.

[0076] The application innovatively proposes to train the organ classifier by considering the speech emotion feature, so as to improve the organ recognition ability of the organ classifier. The training process of the application mainly comprises two parts, which will be described below by means of an embodiment.

[0077] The first step,

[0078] The emotional common features of the speech emotion feature and the speech organ feature are extracted, and the organ speech relationship embedding feature is obtained after weighted fusion; the soft label of the speech emotion feature and the organ speech relationship embedding feature is extracted, and the organ classifier is preliminarily trained according to the soft label.

[0079] S1, in the embodiment, speech emotion data and speech organ data are acquired first, and the application independently constructs a speech organ data set with five types of organ labels, and the sample distribution is heart (890), liver (861), spleen (3900), lung (2482) and kidney (1866).

[0080] The speech emotion data is selected from the CASIA Chinese emotion corpus publicly released, which contains five types of emotion samples (anger, sadness, fear, joy and surprise) corresponding to the organs (heart, liver, spleen, lung and kidney).

[0081] Then feature extraction is performed through a multi-scale speech encoder to obtain speech emotion features and speech viscera features.

[0082] In this embodiment, the multi-scale speech encoder includes three parallel two-dimensional convolution layers and a three-layer deep convolution network.

[0083] The two-dimensional convolution layer is used to extract data features from the time, frequency and spatial dimensions.

[0084] The three-layer deep convolution network is used to further extract multi-scale speech features according to the data features of different dimensions.

[0085] In one embodiment, as shown in Figure 2 The multi-scale speech encoder includes three parallel two-dimensional convolution layers with different convolution kernel sizes, which are designed to extract deep time, frequency and spatial change patterns from primary speech features and simultaneously extract multiple fine-grained information at the same level. Subsequently, the different fine-grained features are stacked along the channel dimension and fed into a three-layer deep convolution network to further extract multi-scale speech features. The multi-scale speech encoder finally outputs a 512-dimensional deep speech feature representation.

[0086] Preferably, the present application takes spectrograms as primary speech features, and extracts 128-dimensional Log-Mel spectrum and 40-dimensional Mel Frequency Cepstral Coefficient (MFCC) as primary speech features, respectively.

[0087] S2, extract the emotional common features of the speech emotion features and the viscera speech features, and obtain the viscera speech relationship embedding features after weighted fusion; the step includes:

[0088] Calculate batch attention based on the speech emotion features and the speech viscera features;

[0089] Determine the attention weight of the hidden state according to the hidden state of the batch attention;

[0090] Weight the attention weight to the hidden state of the batch attention to obtain the viscera speech relationship embedding features.

[0091] The present application discovers the potential relationship between emotions and viscera based on speech relationship attention, extracts the emotional common features of the speech emotion features and the viscera speech features, and generates the speech relationship embedding of the relationship between mixed speech emotions and speech viscera.

[0092] During the model training phase, the speech relationship attention can perform multi-level feature relationship mining on the input batch data to maximize the revelation and integration of the potential emotional relationship contained therein. Referring to Figure 3 , mainly includes batch attention and hidden state attention.

[0093] Further, the voice emotion feature data is defined as the source domain training data, as follows:

[0094]

[0095] The voice visceral feature data is defined as the target domain training data, as follows:

[0096]

[0097] The batch attention can be expressed as:

[0098] B = unsqueeze({B s ,B t}) ∈ R 2b×1×d

[0099]

[0100] unsqueeze refers to tensor dimension expansion, which is to expand the original training batch B ∈ R 2b×d to meet the calculation of dot product attention. b is the number of samples, and d is the number of features of the sample.

[0101] In an embodiment, to effectively capture and integrate the emotional relationship information at different levels and scales, an additional hidden state attention mechanism is applied to the hidden state of the batch attention, expressed as:

[0102]

[0103] X r = A h H

[0104] In the formula, refers to the set of all hidden states calculated by the batch attention, which meets the requirements of dot product attention, n represents the number of hidden states, 2b represents the number of samples, and d represents the feature dimension.

[0105] By globally weighting and summing all hidden states, the long and short range dependencies and cross-scale interactions between different positions can be captured, and key emotional common features can be effectively extracted from different scale feature representations X i , and finally linearly weighted and fused to obtain voice relationship embedding X r .

[0106] In an embodiment, due to the large difference in data feature distribution between different corpora, it is difficult for the model to capture the optimal emotional relationship features, resulting in performance degradation when dealing with data imbalance. Therefore, a double discriminator structure is proposed based on voice relationship attention;

[0107] Referring to Figure 4 , the parameters of the multi-scale speech encoder are updated by using a double discriminator, wherein the double discriminator comprises a domain discriminator and a batch discriminator. First, a domain discriminator structure with a gradient flipping layer is applied, aiming to preliminarily align the feature distribution between two domains, so as to reduce the influence of feature distribution difference on the calculation of speech relationship attention as much as possible. The gradient flipping layer will multiply the gradient from the domain discriminator by a negative constant, thereby realizing the "flipping" of the gradient.

[0108] In the embodiment, the constraint mode of the domain discriminator is as follows:

[0109] G d (G f ; θ d ) = GRL(softmax(G f (x); θ d ))

[0110]

[0111] In the formula, G d is a domain discriminator, which is a linear binary classifier in the embodiment, and the purpose is to learn a logistic regression G d : R D → [0, 1] to judge whether the given input is from the source domain or the target domain; G f represents a multi-scale speech encoder, θ d represents the calculation parameters of the domain discriminator, L d is the cross-entropy loss of the domain discriminator, which is preferably binary cross-entropy loss, used to measure the difference between the domain class predicted by the domain discriminator and the real domain class, y d is the real domain class of the domain discrimination.

[0112] The batch discriminator discriminates the speech relationship embedding output by the speech relationship attention, guides the model to retain the similarity of the samples in the domain, and moderately adjusts the difference between the samples in different domains according to the following formula, so that it is kept within a reasonable range.

[0113] G b (G f ; θ b ) = softmax(G f (x); θ b )

[0114]

[0115] In the formula, G b is a batch discriminator, which is used to learn a logistic regression to judge whether the given input is from the source domain batch or the speech relationship embedding, Gf denotes a multi-scale speech encoder, θ b denotes the computation parameter of the batch discriminator, l b is the cross-entropy loss of the batch discriminator, y b is the real domain class of the batch discrimination.

[0116] S3, obtaining a soft label of the zang-fu speech relationship embedding feature,

[0117] The batch soft label strategy takes the similarity between each source domain feature and the speech relationship embedding as the confidence threshold of the soft label, creates an adaptive soft classification boundary for each speech relationship embedding, and encourages the speech relationship embedding to form a close cluster with the same class features in the source domain.

[0118] The application extracts the soft label of the speech emotion feature and the zang-fu speech relationship embedding feature according to the following formula respectively;

[0119]

[0120] In the formula, S represents the dot product similarity matrix of the normalized speech emotion feature and the zang-fu speech relationship embedding feature, B s denotes a set of speech emotion feature vectors, B r denotes a set of zang-fu speech relationship embedding feature vectors, target denotes the correct label of the speech emotion feature, Y′ (i) denotes the soft label of the zang-fu speech relationship embedding feature.

[0121] Y′ (i) denotes the i-th dimension of the one-hot vector label of the speech relationship embedding, if a source domain sample belongs to the i-th class, the value of the i-th dimension of the one-hot vector label corresponding to the speech relationship embedding is set to S, otherwise the value is where C represents the number of classes, and target is the correct label of the source domain.

[0122] Further, in order to solve the inconsistency problem of the processing flow in the training and testing stages, the application introduces a shared classifier strategy, as shown in Figure 5 The core idea is to fully utilize the potential speech relationship captured through speech relationship learning in the training stage, and integrate it into the learning process of the shared classifier, which can indirectly improve the recognition ability of the shared classifier for test features that have not been processed through speech relationship learning.

[0123] In an embodiment, the loss function of the preliminary training is represented as:

[0124]

[0125] In the formula, N represents the sample number of the voice emotion feature, M represents the sample number of the voice Zangfu feature, a, b, and l represent regularization parameters, X s represents the voice emotion feature, X r represents the Zangfu voice relationship embedding feature, X t the voice Zangfu feature, Y represents the feature X s the corresponding real class label, Y' represents the soft label of the Zangfu voice relationship embedding feature.

[0126] In the second step,

[0127] To further optimize the above technical solution, the voice emotion feature and the deep-level feature of the voice Zangfu feature are extracted, and consistency constraints are performed based on the Euclidean space and the hyperbolic space; and the deep-level feature, the Zangfu voice relationship embedding feature, and the consistency constraints are combined to perform secondary training on the Zangfu classifier preliminarily trained.

[0128] In an embodiment, the extraction step of the deep-level feature includes:

[0129] The features are divided into feature blocks, and linearly projected into a high-dimensional feature space to obtain high-dimensional feature blocks,

[0130] The high-dimensional feature blocks are divided into two parts in the channel dimension to obtain multi-dimensional feature one and multi-dimensional feature two;

[0131] The multi-dimensional feature one and the multi-dimensional feature two are subjected to cross-dimension linear feature interaction in the row and column directions, and then deep-level features are obtained through channel fusion.

[0132] The execution process of the above steps can refer to Figure 6 , Figure 6 The above steps are executed in a stacked dimension multi-layer perception structure, and the dimension multi-layer perception structure is used to encode the emotional voice and the Zangfu voice in a deeper level, so that more accurate features can be extracted.

[0133] In the present application, the voice signal is converted into a spectrogram as input, and the dimension multi-layer perception first linearly projects the input feature representation into a series of feature blocks, denoted as where F and T represent the frequency and time dimensions of the feature blocks respectively, the channel number is 1, and the block size is p x p.

[0134] Then, the dimension multi-layer perception linearly projects all the feature blocks into a high-dimensional feature space C to obtain high-dimensional feature blocks Subsequently, the feature blocks are divided into two parts in the channel dimension, and linear feature interaction is performed in the row and column directions respectively.

[0135] The above process can be represented as:

[0136]

[0137] X m1 , X m2 = MLP ft (Transpos(Split c (X m )))

[0138]

[0139] where MLP c represents a channel-aware machine, which is used to linearly project the two-dimensional feature block into a high-dimensional space. MLP ft represents a cross-dimensional linear interaction-aware machine to achieve feature interaction in the time and frequency dimensions of the transposed feature block.

[0140] In this embodiment, the dimension multi-layer perception structure introduces a cascaded interaction form in feature interaction, and the weighted fusion of linear layers makes the block features of each row (or column) not only serve the feature aggregation of the current row (or column), but also contribute to the feature aggregation of the adjacent row (or column), realizing cross-dimensional normalization, i.e.

[0141] X (1) *,j = W1X *,j ; X (1) i,* = W2X i,*

[0142] X *,j = W3Cat(X (1) *,j-1 , X (1) *,j , X (1) *,j+1 )

[0143] X i,* = W4Cat(X (1) i-1,* X (1) i,* , X (1) i+1,* )

[0144] where W1, W2, W3 and W4 are weight parameters of linear layers, X *,j represents the feature vector of the jth column, X i,j represents the element of the ith row and the jth column, Cat represents a tensor dimension concatenation operation, X *,j-1 represents all elements of the j-1th column, X *,j+1 represents all elements of the j+1th column, X i-1,*X i+1,* represents all elements of the i+1th row.

[0145] Further, X m1 , X m2 are respectively calculated by the above method.

[0146] Then, for effective aggregation of spatial information and channel information, the model applies a single convolution layer between stacked dimensional multi-layer perceptrons, so that the spatial dimension of the input feature is reduced by 2x2 layer by layer while the channel dimension is doubled, gradually reducing the feature from to and using residual connection to avoid gradient vanishing and gradient explosion problems. Finally, the dimensional multi-layer perceptron inputs the deep aggregated features extracted by multiple blocks into a global average pooling layer for dimension reduction, which is expressed by the formula:

[0147]

[0148] In the formula, X m1 represents the multi-dimensional feature one, X m2 represents the multi-dimensional feature two, and X m represents the high-dimensional feature block. represents the final extracted feature.

[0149] In a preferred embodiment, as shown in Figure 7 , by calculating the emotional visceral spatial consistency constraint loss, the emotional features and visceral features are constrained and aligned in multiple spaces.

[0150] There is a close relationship between human emotions and human viscera, so in the context of multi-space embedding, emotions and viscera can be mapped into feature vectors respectively to explore their relationship in different geometric spaces.

[0151] Suppose in Euclidean space, the distance between the emotional vector x e and the visceral vector x p is D es , and in hyperbolic space, the distance between the emotional vector x e and the visceral vector x p is If the ratio or absolute difference between D es and is within a small threshold range, it is considered that there is a high spatial consistency between x e and x p .

[0152] In this embodiment, suppose the mapping of the emotional features and visceral features extracted by the dimensional multi-layer perceptron in Euclidean space is and The mapping in hyperbolic space is and

[0153] wherein the mapping function from the Euclidean space to the hyperbolic space is the mapping function from the hyperbolic space to the Euclidean space is the mapping function is specifically expressed as

[0154]

[0155] wherein v is a reference point in the space, is a special pair operation, denotes the Möbius addition, that is:

[0156]

[0157] wherein c is a space curvature constant, usually less than 0, and in this document, the value is uniformly taken as -1.

[0158] Further, the embodiment performs consistency constraint according to the cosine distance of the mapped features in the Euclidean space and the hyperbolic space; the formula is expressed as:

[0159]

[0160] wherein, and denote the mapped features of the speech emotion features in the Euclidean space and the hyperbolic space respectively, and denote the mapped features of the speech Zangfu features in the Euclidean space and the hyperbolic space respectively; and D es denotes the feature and in the Euclidean space,

[0161] denotes the distance between the features and in the hyperbolic space; the calculation formula is as follows:

[0162]

[0163] The present application takes the cross entropy as the loss function, and the loss function of the secondary training is as follows:

[0164]

[0165] wherein, denotes the deep feature of the speech emotion feature, denotes the deep feature of the speech Zangfu feature, Y e and Y p are respectively corresponding true label, X r represents the zang-fu voice relationship embedding feature, Y' represents X r corresponding true label, a, b, and l are regularization parameters.

[0166] wherein the loss function of cross-entropy is expressed as:

[0167]

[0168] wherein y p is the true label of the input sample, and y' i is the predicted label of the classifier.

[0169] After the zang-fu classifier is trained through the above steps, the zang-fu classification can be directly recognized based on the voice zang-fu data. Specifically, in actual application testing, only the zang-fu voice is taken as the input, and the multi-label zang-fu classification probability is obtained by the shared zang-fu classifier through the multi-scale voice encoder and the multi-dimensional multilayer perception structure.

[0170] In the present application, the above constraint aims to implicitly guide the model in the learning process to not only pursue good performance in a single space, but also to ensure that the distance between the emotion vector and the zang-fu vector in multiple spaces remains consistent. It is expected that in this way, the association between emotion and zang-fu is understood from multiple spatial perspectives, and the generalization ability of the model is improved.

[0171] The embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the related parts can be referred to the method part.

[0172] The above description of the disclosed embodiments enables a person skilled in the art to implement or use the present application. Various modifications to the embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for identifying organs based on speech emotion and speech features, characterized in that: Acquire the voice viscera features, output the corresponding viscera classification probability through the viscera classifier, and identify the viscera organs according to the viscera classification probability; wherein, The viscera classifier is trained based on the speech emotion feature data and the speech viscera feature data, and the training process includes: Extract the common emotional features of speech emotion features and speech viscera features, and obtain the viscera-speech relationship embedding features after weighted fusion; Obtaining soft labels of the embedded features of the organ speech relationship, and performing preliminary training on the organ classifier based on the soft labels and the true labels corresponding to the speech emotion features; The deep-level features of speech emotion features and speech viscera features are extracted, and consistency constraints are performed based on Euclidean space and hyperbolic space. At the same time, the deep-level features, viscera speech relationship embedding features and consistency constraints are combined to perform secondary training on the initially trained viscera classifier.

2. The method for identifying organs based on speech emotion and speech features according to claim 1, characterized in that: The speech emotion feature data and the speech viscera feature data are obtained by extracting the features of the speech emotion data and the speech viscera data respectively using a multi-scale speech encoder; The multi-scale speech encoder includes three parallel two-dimensional convolutional layers and a three-layer deep convolutional network; The two-dimensional convolutional layer is used to extract data features from time, frequency and space dimensions; The three-layer deep convolutional network is used to further extract multi-scale speech features based on data features of different dimensions.

3. The method for identifying organs based on speech emotion and speech features according to claim 2, characterized in that: The parameters of the multi-scale speech encoder are constrained and updated using a dual discriminator, which includes a domain discriminator and a batch discriminator. The constraints of the domain discriminator are: G d (G f ;θ d )=GRL(softmax(G f (x);θ d )) Where G d is the domain discriminator, G f represents the multi-scale speech encoder, θ d represents the calculation parameters of the domain discriminator, L d is the cross entropy loss of the domain discriminator, y d The real domain category for domain discrimination; The constraints of the batch discriminator are: G b (G f ;θ b )=softmax(G f (x);θ b ) Where G b is the batch discriminator, G f represents the multi-scale speech encoder, θ b Represents the calculation parameters of the batch discriminator, L b is the cross entropy loss of the batch discriminator, y b is the true domain category for batch discrimination.

4. The method for identifying organs based on speech emotion and speech features according to claim 1, characterized in that: Extract the common emotional features of speech emotion features and speech viscera features, and obtain the viscera speech relationship embedding features after weighted fusion; the steps include: Calculate batch attention based on speech emotion features and speech organ features; Determine the attention weight of the hidden state based on the hidden state of the batch attention; The attention weight is weighted to the hidden state of the batch attention to obtain the viscera and voice relationship embedding feature.

5. The method for identifying organs based on speech emotion and speech features according to claim 1, characterized in that: Obtain the soft label of the viscera and corpus phonetic relationship embedding feature according to the following formula; Where S represents the dot product similarity matrix of the normalized speech emotion feature and the embedded feature of the viscera speech relationship, B s Represents the speech emotion feature vector set, B r represents the set of embedded feature vectors of the internal organs’ speech relationship, target represents the correct label of the speech emotion feature, and Y′ (i) The soft label represents the embedded feature of the organ-speech relationship, and i represents the i-th dimension data of the soft label.

6. The method for identifying organs based on voice emotion and voice features according to claim 3, characterized in that: The loss function for the initial training is: Where N is the number of samples of speech emotion features, M is the number of samples of speech viscera features, α, β, and λ are regularization parameters, and X is the number of samples of speech emotion features. s represents the emotional characteristics of speech, X r represents the embedded features of organ phonetic relations, X t Voice organ features, Y represents feature X s The corresponding true category label, Y′ represents the soft label of the viscera phonetic relationship embedding feature.

7. The method for identifying organs based on speech emotion and speech features according to claim 1, characterized in that: Extracting deep features of speech emotion features and speech viscera features respectively; The extraction steps include: Divide the features into feature blocks and linearly project them into high-dimensional feature space to obtain high-dimensional feature blocks. The high-dimensional feature is split into two parts from the channel dimension to obtain multi-dimensional feature 1 and multi-dimensional feature 2; Make multidimensional feature 1 and multidimensional feature 2 interact with each other in the row and column directions, and then obtain the deep features according to the following formula; Where, X m1 Represents multidimensional feature one, X m2 Represents multidimensional feature 2, X m represents a high-dimensional feature block, Represents the final extracted features.

8. The method for identifying organs based on speech emotion and speech features according to claim 7, characterized in that: The process of cross-dimensional linear feature interaction is expressed as follows: X (1) *,j =W1X *,j ;X (1) i,* =W2X i,* X *,j =W3Cat(X (1) *,j-1 ,X (1) *,j ,X (1) *,j+1 X i,* =W4Cat(X (1) i-1,* ,X (1) i,* ,X (1) i+1,* ) Where W1, W2, W3 and W4 are the weight parameters of the linear layer, X *,j represents the eigenvector of the jth column, X i,j represents the element in row i and column j, Cat represents the tensor dimension concatenation operation, X *,j-1 represents all elements in the j-1th column, X *,k+1 represents all elements in the j+1th column, X i-1,* represents all elements in row i-1, X i+1,* Represents all elements in row i+1.

9. The method for identifying organs based on speech emotion and speech features according to claim 1, characterized in that: The deep features of speech emotion features and speech viscera features are mapped to Euclidean space and hyperbolic space respectively, and consistency constraints are imposed based on the cosine distance of the mapped features in Euclidean space and hyperbolic space; the formula is expressed as: in, and Respectively represent the mapping characteristics of speech emotion features in Euclidean space and hyperbolic space, and Respectively represent the mapping characteristics of speech viscera features in Euclidean space and hyperbolic space; and D es Representation characteristics and The distance in Euclidean space, Representation characteristics and Distance in hyperbolic space; calculated as follows: Where c represents the curvature parameter and h represents the radius parameter.

10. The method for identifying organs based on speech emotion and speech features according to claim 9, characterized in that: The loss function of the secondary training is: Where, Deep features that represent emotional characteristics of speech, Indicates the deep-level characteristics of the voice organs, Y e and Y p They are The corresponding true label, X r represents the embedded features of the organ phonetic relationship, and Y′ represents X r The corresponding true label, α, β, λ are regularization parameters.

Citation Information

Patent Citations

  • Traditional Chinese medicine sound smelling diagnosis automatic system supported by smart voice technology

    CN112002342A

  • Hyperbolic space alignment-based multi-modal voice internal organ recognition method

    CN117958765A