Classification method and system for data enhancement and hybrid expert mechanism feature selection
Through the feature selection method of data enhancement and hybrid expert mechanism, the recognition accuracy and robustness of speech analysis in the detection of cognitive function changes are improved, the problems of data scarcity and insufficient feature fusion are solved, and efficient cognitive abnormality screening is achieved.
Patent Information
- Application Number
- CN202510705540.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-05-29
AI Technical Summary
Existing technologies suffer from data scarcity, low feature selection efficiency, and limited feature fusion effects, resulting in insufficient recognition accuracy of speech analysis in detecting changes in cognitive function.
The method of data enhancement and hybrid expert mechanism multi-level feature selection is adopted to generate enhanced speech through cross-sample paired text-to-speech synthesis. Combined with multi-level feature extraction and co-attention mechanism, multimodal features are dynamically selected and integrated to improve the diversity of speech data and the generalization ability of the model.
It significantly improves the recognition accuracy and robustness of the speech classification system, especially in small samples and variable semantic scenarios, and achieves high-precision recognition of cognitive abnormalities, which is suitable for low-cost, non-invasive large-scale screening.
Smart Images

Figure CN120612925A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a technology in the field of information processing based on speech analysis, specifically a classification method and system based on text-to-speech data enhancement and hybrid expert mechanism multi-level feature selection. Background Art
[0002] With the advancement of artificial intelligence and speech processing technologies, methods for detecting cognitive changes based on speech analysis have become an important area of research in feature recognition. Studies have found that the early stages of cognitive changes may manifest as disorganized speech, decreased speech rate, frequent pauses, unclear pronunciation, and repetitive words, and these characteristics are highly quantifiable. Compared with traditional detection methods, speech analysis offers advantages such as ease of data collection, low cost, and high non-invasiveness, making it particularly suitable for large-scale screening in community or remote settings. Summary of the Invention
[0003] In response to the shortcomings of existing technologies in terms of data scarcity, low feature selection efficiency and limited feature fusion effect, the present invention proposes a classification method and system for data enhancement and hybrid expert mechanism feature selection. It uses a speech data enhancement module based on speech-to-text (TTS) technology to improve data diversity and model generalization ability; uses multi-level acoustic and text feature extraction to more comprehensively characterize language changes; uses a hybrid expert mechanism (MoE) mechanism to achieve dynamic selection of multimodal features and improve feature utilization efficiency; uses a co-attention mechanism to optimize the fusion method between different modal features and enhance the interactive expression ability between features, significantly improving the recognition accuracy and system robustness in a multimodal environment, and overcoming the above-mentioned shortcomings.
[0004] The present invention is achieved through the following technical solutions:
[0005] The present invention relates to a classification method for data enhancement and hybrid expert mechanism multi-level feature selection, comprising:
[0006] Step 1: Collect the subject's natural speech input as the original audio signal, which provides the basic data source for subsequent speech recognition and feature extraction, and helps capture the patient's natural language.
[0007] Step 2: Use a pre-trained ASR model (such as Whisper) to transcribe the original audio to obtain the text transcript corresponding to the speech sample;
[0008] Step 3: Re-input the transcribed text into a TTS model (such as FishSpeech). By pairing samples across different samples (i.e., the speaker is from sample A, the text content is from sample B, A≠B, and A and B belong to the same category), an enhanced speech is synthesized that retains the speaker's timbre but has different semantic content. The synthesized speech and the original speech together constitute the extended speech dataset.
[0009] Step 4: For the original speech and TTS enhanced speech samples, extract the corresponding modal features through different encoders, including:
[0010] a. Spectrogram features (Spec): First, the audio is converted into a corresponding spectrogram. Then, the spectrogram is processed using the ResNet18 deep neural network model to extract spectrogram features and capture the audio's time-frequency information. Spectrogram features provide high-order frequency domain information of the audio signal and are a common feature in speech signal analysis.
[0011] b. Mel-Frequency Cepstral Coefficient (MFCC) features: MFCC features are processed using a bidirectional long short-term memory (BiLSTM) network. These features reflect the human auditory system's perception of sound and have strong biological relevance. Using BiLSTM effectively captures the temporal dependencies within MFCC features.
[0012] c. Text features: ASR transcription is performed using the Whisper model to convert the speech content into text. The BERT pre-trained model is then used to perform semantic understanding and feature extraction on the transcribed text, providing semantic-level feature support for subsequent classification.
[0013] d. Wav2Vec2 features: The Transformer-based Wav2Vec2 model extracts deep time-domain features from the original audio signal. The Wav2Vec2 model is capable of capturing long-term dependencies in speech and handling more complex speech patterns.
[0014] Step 5: Input the three modal features extracted by ResNet18, BiLSTM, and BERT into the corresponding MoE submodules containing several expert networks and gating networks, and dynamically select and output the most discriminative feature representation;
[0015] The MoE submodule includes: k expert networks , a gating network corresponding to the features of each set of inputs And normalized weights, ensuring that each weight is positive and the sum is 1 Softmax function, where: each expert network is based on the feature Get the output , the gating network generates the weight vector , the MoE output of each type of feature is expressed as the weighted sum of the outputs of each expert, specifically: Spectral graph feature: MFCC features: ;Text features: , ; is the weight matrix, is the bias term.
[0016] Step 6: Input the features selected by MoE and the Wav2Vec2 features that have not been through MoE into the co-attention mechanism module, and perform weighted adjustment on Wav2Vec2 by calculating the correlation between features, so as to improve the interaction effect between different modalities. Specifically: First, the spectrogram features output by the MoE module are , MFCC features and text features Splice and get the intermediate fusion feature representation , and through the attention transformation function Generate fusion weight vector ; Then the fusion weight and deep acoustic features extracted by the Wav2Vec2 model Perform element-by-element multiplication and perform aggregation operations on the time series dimension to obtain the final weighted Wav2Vec2 feature vector , and its calculation formula is ,in: , for the time step features of Wav2Vec2 After weight adjustment, the final output is the low-dimensional representation after fusion.
[0017] This process realizes the explicit collaborative modeling between the original time domain features and other modal features, thereby enhancing the comprehensive perception ability of Wav2Vec2 on spectrum, rhythm and semantics; and 、 、 They are concatenated together to form the final fused feature representation, which is then input into the downstream classification module.
[0018] Step 7: All optimized multimodal features are spliced and fused to generate the final fused representation, which specifically includes: early fusion and late fusion. The early fusion strategy fuses all modal features before entering the co-attention mechanism and unifies the modeling interaction weights; while the late fusion strategy processes each type of feature separately and integrates the text features in the final classification stage. The present invention verifies the performance of the two strategies through experimental comparison. The results show that late fusion performs better in maintaining feature independence, improving model stability and classification accuracy. Therefore, this system ultimately adopts the late fusion strategy for implementation;
[0019] Step 8: Input the fused features into a deep neural network classifier (such as a multi-layer perceptron MLP) and output the prediction result for the current sample, that is, judge whether the speaker is a cognitively abnormal (or cognitively normal) individual.
[0020] Technical Effects
[0021] The present invention constructs a TTS enhancement strategy for cross-sample speech-text pairing, and combines multimodal feature extraction with the MoE module to achieve dynamic selection of spectrograms, MFCCs, and text features. At the same time, it introduces the co-attention mechanism Wav2Vec2 to fuse information with other features to build a fully optimized speech feature classification system architecture. Compared with the existing technology, the present invention significantly enhances the adaptability of the speech classification system to a small number of samples, variable semantics, and pronunciation differences; improves the discriminability of feature selection and the robustness of the system; achieves classification performance significantly better than existing methods on public datasets (with an accuracy rate of over 85%); and achieves accurate recognition of changes in implicit key features in speech, making it suitable for low-cost, non-invasive, large-scale speech screening scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 Flowchart of the present invention;
[0023] Figure 2 Schematic diagram of the system structure of the present invention;
[0024] Figure 3 This is a schematic diagram of the present invention in the pre-treatment stage;
[0025] Figure 4 Schematic diagram of the generation path of different types of features of the present invention;
[0026] Figure 5 Schematic diagram of the feature selection structure of the hybrid expert mechanism of the present invention;
[0027] Figure 6 Schematic diagram of the co-attention mechanism and classification structure of the present invention;
[0028] Figure 7 A schematic diagram of the deployment of the present invention in an actual application scenario;
[0029] Figure 8 This is a performance comparison chart of the proposed method and the existing technology under the accuracy index;
[0030] In the figure, the horizontal axis is the name of each comparison method, and the vertical axis is the classification accuracy (%). DETAILED DESCRIPTION
[0031] like Figure 2 As shown, this embodiment relates to a method for realizing the above Figure 1 The data enhancement and hybrid expert mechanism multi-level feature selection classification system of the method shown includes: an automatic speech recognition (ASR) module, a text-to-speech synthesis (TTS) module, a multimodal feature extraction module, a hybrid expert mechanism (MoE) module, a co-attention mechanism module, a feature fusion module and a classification module, wherein: the automatic speech recognition module transcribes the original audio through a pre-trained ASR model to obtain text transcription content corresponding to the speech sample; the text-to-speech synthesis unit re-inputs the transcribed text into the TTS model, and synthesizes enhanced speech that retains the speaker's timbre but has different semantic content through cross-sample pairing. The synthesized speech and the original speech together constitute the extended speech dataset; the multimodal feature extraction extracts spectrogram features, Mel-frequency cepstral coefficient (MFCC) features, deep temporal audio embedding features and semantic text features from the original audio and TTS enhanced speech samples through different encoders; the mixture of experts (MoE) module inputs the spectrogram features, Mel-frequency cepstral coefficient (MFCC) features and deep temporal audio embedding features into the corresponding MoE sub-modules and obtains feature representations; the co-attention mechanism module weights and adjusts the semantic text features by calculating the correlation between features, thereby improving the interaction effect between different modalities; the feature fusion module splices and fuses all optimized multimodal features to generate the final fused representation; the classification module inputs the fused features into the deep neural network classifier and outputs the prediction result for the current sample, that is, judging whether the speaker is cognitively abnormal or cognitively normal.
[0032] like Figure 3 As shown, the automatic speech recognition (ASR) module includes: an audio preprocessing unit, an acoustic modeling unit and a language decoding unit, wherein: the audio preprocessing unit performs endpoint detection, normalization and denoising on the input speech signal; the acoustic modeling unit inputs the extracted features into the ASR model (such as Whisper) for inference; and the language decoding unit outputs the final text transcription result based on the language model.
[0033] like Figure 3 As shown, the TTS voice data enhancement module includes: an original voice collection and recognition unit, a voice-text cross-sample pairing device, a text-to-speech synthesis unit and an enhanced voice set construction unit, wherein: the original voice collection and recognition unit collects the original voice set And input into the automatic speech recognition module, and transcribed into the corresponding text set through models such as Whisper ; Speech-text cross-sample pairer from Select a voice and compared with text from different speech samples of the same category (where i≠j) is paired; the text-to-speech synthesis unit will pair the result ( ) Input TTS model (such as FishSpeech) to generate synthetic speech samples ; Enhanced speech set construction unit for all and different Perform traversal synthesis operations on the combination of The synthesized speech and the original speech are combined to form the expanded speech training dataset.
[0034] Each speech in the original speech set is generated by a picture-text description task.
[0035] The synthesized speech is kept While carrying the speaker's timbre, rhythm and emotional characteristics, semantic content.
[0036] This TTS enhancement module significantly improves the model's robustness and recognition capabilities when processing complex speech scenarios (such as unclear pronunciation, short sentences or semantic confusion), effectively enhances the system's generalization ability for different speaker features and semantic combinations, and provides richer and more balanced data support for intelligent feature classification tasks.
[0037] like Figure 4 As shown, the multimodal feature extraction module includes a ResNet18 spectrogram feature extraction unit, a BiLSTM MFCC modeling unit, a Whisper transcription combined with BERT text encoding unit, and a Wav2Vec2 deep embedding extraction unit. These modules are responsible for extracting structural features of different modalities from the enhanced speech and transcribed text, respectively, laying the foundation for subsequent feature fusion.
[0038] like Figure 5 As shown, the hybrid expert mechanism module dynamically selects multi-level audio and text features to optimize the discriminability and robustness of the features. The module includes: a feature acquisition unit, an MoE expert network, a gating network and a weighting unit, and a weighted output unit, wherein: the feature acquisition unit collects 1) spectrogram features (extracted by ResNet18); 2) MFCC features (extracted by BiLSTM) and 3) text features (extracted by BERT); the MoE expert network independently processes the three input features through three sub-networks to obtain multiple expert outputs; the gating network and weighting unit use the softmax function to calculate weights and dynamically adjust the weight vector to optimize feature selection; the weighted output unit multiplies the expert network output by the corresponding weight for each feature to obtain the weighted feature.
[0039] For example, spectrogram features, MFCC features, and text features are processed by the expert network and the gating network, respectively, to generate weighted final feature representations. These weighted feature representations will enter the downstream co-attention mechanism to further optimize the fusion expression of cross-modal information.
[0040] This embodiment effectively reduces the interference of redundant features and improves the discriminability of features through the hybrid expert mechanism module. Compared with traditional feature selection methods, the MoE module can be more flexible and precise in the selection of multimodal features, thereby significantly improving the accuracy and robustness of the final classification. In particular, the application of this module improves classification accuracy and optimizes the generalization ability of the model in cognitive classification tasks.
[0041] like Figure 6 As shown, the co-attention mechanism module fuses the weighted features of multiple feature sources and classifies the cognitive situation according to these fused features. The module includes: a feature acquisition unit, a co-attention mechanism unit, a classification unit and an output unit, wherein: the feature acquisition unit collects the spectrum map features, MFCC features, text features and Wav2Vec2 features that are screened and weighted by the MoE module; the co-attention mechanism unit performs weighted fusion on the input features (spectrogram features, MFCC features, text features and Wav2Vec2 features) to obtain a comprehensive feature matrix, and further multiplies the input features element by element with the Wav2Vec2 features to obtain weighted Wav2Vec2 features, and then splices all the weighted features to obtain a comprehensive feature representation; the classification unit uses a deep neural network (such as MLP) to predict and classify the comprehensive feature representation; the output unit outputs according to the predicted classification results.
[0042] like Figure 7 The following figure shows the application architecture of this embodiment based on the aforementioned classification system, including a user device, a classification system located in a server / cloud processing center, and an output terminal. The user device collects the subject's voice input and uploads the data via a mobile terminal or local application. The classification system sequentially performs audio signal preprocessing, feature extraction, expert mechanism selection, and fusion classification operations, returning the classification results to the output terminal for display and report generation by the physician system. This architecture is suitable for a variety of deployment scenarios, including hospitals and mobile health devices, and is both practical and scalable.
[0043] Experiments were conducted on the ADReSSo dataset, using Whisper as the ASR model, FishSpeech as the TTS module, and BERT-base as the text feature extractor. Training parameters were batch size 32, learning rate 1e-4, 200 training rounds, and five runs per experiment, with the average taken. This method demonstrated higher classification accuracy and robustness than existing methods in cognitive classification tasks, particularly when applied to diverse speech data and smaller sample sets. A comparison of the effectiveness of this method and existing methods in predicting cognitive abnormality (CA) and cognitive normality (CN) is shown in Table 1.
[0044] Table 1
[0045] As shown in Table 1, by comparing the performance of different methods in terms of accuracy, precision, recall rate, and F1 score, we can comprehensively evaluate the effectiveness and advantages of the methods in cognitive impairment screening.
[0046] Accuracy: As shown in Table 1, our method achieved a high classification accuracy of 85.71% for both CA and CN. Compared to other methods, this result demonstrates that our method is more accurate in classification, especially when using multimodal feature input and data augmentation, which allows it to better adapt to the diversity of different samples and reduce classification errors.
[0047] Accuracy: Our method achieved an accuracy of 93.10% for CA, significantly higher than other compared methods, particularly single-modality methods such as Wav2Vec2+TB (72.50% accuracy) and Whisper-TLmedium (77.78% accuracy). This demonstrates that our method is more effective in distinguishing CA individuals and reducing false positives.
[0048] Recall: This method achieved a recall rate of 94.29%, exceeding most other methods. For example, Whisper-TLmedium achieved a recall rate of 77.78%, and Wav2Vec2+TB achieved a recall rate of 72.73%. This high recall rate indicates that this method can effectively identify a large proportion of CA patients, reducing underreporting. This is particularly significant given that early-stage CA symptoms may be mild.
[0049] F1 score: This method also significantly outperforms other methods in terms of F1 score, reaching 84.38%, higher than other methods such as Whisper-TLmedium (F1 score of 77.78%) and Wav2Vec2+TB (F1 score of 76.32%). This shows that this method achieves a good balance between ensuring high precision and maintaining a high recall rate.
[0050] like Figure 8 The following is a bar chart comparing the classification accuracy of our method with existing techniques. This experiment was conducted on the publicly available ADReSSo cognitive impairment recognition dataset, evaluating the performance of different methods on the same test set to ensure a fair and consistent comparison. The bar chart further supports the practicality and robustness of our system in cognitive impairment screening tasks.
[0051] In summary, compared with existing technologies, this method performs outstandingly across multiple key performance indicators. These results demonstrate that, by combining text-to-speech (TTS) data augmentation with dynamic feature selection using a mixture of experts (MoE) mechanism, this method not only improves the recognition of cognitive functions but also addresses the challenges of complex speech features and diverse data in practical applications, demonstrating its high application value and practicality.
[0052] The above-mentioned specific implementation can be partially adjusted in different ways by those skilled in the art without departing from the principles and purpose of the present invention. The scope of protection of the present invention shall be based on the claims and shall not be limited by the above-mentioned specific implementation. All implementation schemes within its scope shall be subject to the constraints of the present invention.
Claims
1. A classification method based on data enhancement and hybrid expert mechanism multi-level feature selection, characterized in that: include: Step 1: Collect the subject's natural speech input as the original audio signal to provide the basic data source for subsequent speech recognition and feature extraction; Step 2: Use the pre-trained ASR model to transcribe the original audio to obtain the text transcription content corresponding to the speech sample; Step 3: Re-input the transcribed text into the TTS model and synthesize enhanced speech that retains the speaker's timbre but has different semantic content through cross-sample pairing. The synthesized speech and the original speech together form an extended speech dataset. Step 4: For the original speech and TTS enhanced speech samples, extract the corresponding modal features through different encoders; Step 5: Input the three modal features extracted by ResNet18, BiLSTM, and BERT into the corresponding MoE submodules containing several expert networks and gating networks, and dynamically select and output the most discriminative feature representation; Step 6: The features selected by MoE and the Wav2Vec2 features that have not been processed by MoE are input into the co-attention mechanism module. Wav2Vec2 is weighted and adjusted by calculating the correlation between features to improve the interaction effect between different modalities. Step 7: All optimized multimodal features are concatenated and fused to generate the final fusion representation; Step 8: Input the fused features into the deep neural network classifier and output the prediction result for the current sample, that is, determine whether the speaker is a cognitively abnormal or cognitively normal individual.
2. The classification method of data enhancement and hybrid expert mechanism multi-level feature selection according to claim 1 is characterized in that: The corresponding modal features specifically include: a. Spectrogram features (Spec): First, the audio is converted into a corresponding spectrogram. Then, the spectrogram is processed using the ResNet18 deep neural network model to extract spectrogram features and capture the time-frequency information of the audio. Spectrogram features provide high-order frequency domain information of the audio signal and are a common feature in speech signal analysis. b. Mel-Frequency Cepstral Coefficient (MFCC) features: MFCC features are processed using a bidirectional long short-term memory (BiLSTM) network. These features reflect the sound perception characteristics of the human auditory system and have strong biological relevance. By using BiLSTM, the temporal dependencies in MFCC features can be effectively captured. c. Text features: The Whisper model is used for ASR transcription to convert speech content into text. Then, the BERT pre-trained model is used to perform semantic understanding and feature extraction on the transcribed text, providing semantic-level feature support for subsequent classification. d. Wav2Vec2 features: The deep time domain features of the original audio signal are extracted through the Wav2Vec2 model based on the Transformer architecture. The Wav2Vec2 model can capture long-term dependencies in speech and can handle more complex speech patterns.
3. The classification method of data enhancement and hybrid expert mechanism multi-level feature selection according to claim 1 is characterized in that: The MoE submodule includes: k expert networks , a gating network corresponding to the features of each set of inputs And normalized weights, ensuring that each weight is positive and the sum is 1 Softmax function, where: each expert network is based on the feature Get the output , the gating network generates the weight vector , the MoE output of each type of feature is expressed as the weighted sum of the outputs of each expert, specifically: Spectral graph feature: MFCC features: ;Text features: , ; is the weight matrix, is the bias term.
4. The classification method of data enhancement and hybrid expert mechanism multi-level feature selection according to claim 1 is characterized in that: The step 6 is as follows: first, the spectrum features output by the MoE module are , MFCC features and text features Splice and get the intermediate fusion feature representation , and through the attention transformation function Generate fusion weight vector ; Then the fusion weight and deep acoustic features extracted by the Wav2Vec2 model Perform element-by-element multiplication and perform aggregation operations on the time series dimension to obtain the final weighted Wav2Vec2 feature vector , and its calculation formula is ,in: , for the time step features of Wav2Vec2 After weight adjustment, the final output is the low-dimensional representation after fusion.
5. The classification method of data enhancement and hybrid expert mechanism multi-level feature selection according to claim 1 is characterized in that: The step 7 includes: early fusion and late fusion. The early fusion strategy fuses all modal features before entering the co-attention mechanism and unifies the modeling interaction weights; while the late fusion strategy processes each type of feature separately and integrates the text features in the final classification stage. The present invention verifies the performance of the two strategies through experimental comparison. The results show that late fusion performs better in maintaining feature independence, improving model stability and classification accuracy. Therefore, this system ultimately adopts the late fusion strategy for implementation.
6. A classification system that implements the data enhancement and hybrid expert mechanism multi-level feature selection method of any one of claims 1 to 5, characterized in that: include: Automatic speech recognition (ASR) module, text-to-speech synthesis (TTS) module, multimodal feature extraction module, mixture of experts (MoE) module, co-attention mechanism module, feature fusion module and classification module, among which: the automatic speech recognition module transcribes the original audio through the pre-trained ASR model to obtain the text transcription content corresponding to the speech sample; the text-to-speech synthesis unit re-inputs the transcribed text into the TTS model, and synthesizes the enhanced speech that retains the speaker's timbre but has different semantic content through cross-sample pairing. The synthesized speech and the original speech together constitute the extended speech dataset; the multimodal feature extraction is based on the original For audio and TTS enhanced speech samples, spectrogram features, Mel-frequency cepstral coefficient (MFCC) features, deep temporal audio embedding features, and semantic text features are extracted through different encoders. The Mixture of Experts (MoE) module inputs the spectrogram features, Mel-frequency cepstral coefficient (MFCC) features, and deep temporal audio embedding features into the corresponding MoE submodules to obtain feature representations. The co-attention mechanism module weights and adjusts the semantic text features by calculating the correlation between features, thereby improving the interaction effect between different modalities. The feature fusion module splices and fuses all optimized multimodal features to generate the final fused representation. The classification module inputs the fused features into the deep neural network classifier and outputs the prediction result for the current sample, that is, judging whether the speaker is a cognitively abnormal or cognitively normal individual.
7. The classification system of data enhancement and hybrid expert mechanism multi-level feature selection according to claim 6 is characterized in that: The TTS voice data enhancement module includes: an original voice collection and recognition unit, a voice-text cross-sample pairing device, a text-to-speech synthesis unit and an enhanced voice set construction unit, wherein: the original voice collection and recognition unit collects the original voice set And input into the automatic speech recognition module, and transcribe it into the corresponding text set through Whisper ; Speech-text cross-sample pairer from Select a voice and compared with text from different speech samples of the same category (where i≠j) is paired; the text-to-speech synthesis unit will pair the result ( ) Input TTS model to generate synthetic speech samples ; Enhanced speech set construction unit for all and different Perform traversal synthesis operations on the combination of The synthesized speech and the original speech are combined to form the expanded speech training dataset.
8. The classification system of data enhancement and hybrid expert mechanism multi-level feature selection according to claim 6 is characterized in that: The hybrid expert mechanism module includes: a feature acquisition unit, an MoE expert network, a gating network and a weighting unit, and a weighted output unit, wherein: the feature acquisition unit collects spectrum features, MFCC features, and text features; the MoE expert network independently processes the three input features through three sub-networks to obtain multiple expert outputs; the gating network and the weighting unit use the softmax function to calculate the weights and dynamically adjust the weight vector to optimize feature selection; the weighted output unit multiplies the expert network output by the corresponding weight for each feature to obtain the weighted feature.
9. The classification system of data enhancement and hybrid expert mechanism multi-level feature selection according to claim 6, characterized in that: The co-attention mechanism module includes: a feature acquisition unit, a co-attention mechanism unit, a classification unit and an output unit, wherein: the feature acquisition unit collects the spectrum graph features, MFCC features, text features and Wav2Vec2 features that are screened and weighted by the MoE module; the co-attention mechanism unit performs weighted fusion on the input spectrum graph features, MFCC features, text features and Wav2Vec2 features to obtain a comprehensive feature matrix, and then further multiplies them element by element with the Wav2Vec2 features to obtain weighted Wav2Vec2 features, and then splices all the weighted features to obtain a comprehensive feature representation; the classification unit uses a deep neural network to predict and classify the comprehensive feature representation; the output unit outputs according to the predicted classification results.
10. An application architecture based on the system according to any one of claims 6 to 9, characterized in that: include: The user device side, the classification system located in the server / cloud processing center, and the output side, wherein: the user device side collects the subject's voice input and uploads the data through a mobile terminal or local application; The classification system performs audio signal preprocessing, feature extraction, expert mechanism selection and fusion classification operations in sequence, and returns the classification results to the output end.
Citation Information
Patent Citations
Multi-modal data fusion method and device for speech recognition, equipment and medium
CN115762484A
Training method, using method, device and equipment of multi-modal pre-training model
CN116756574A
Alzheimer's disease diagnosis system based on audio and text modal fusion
CN117577306A
Abnormal session problem identification method and device, electronic equipment and readable medium
CN117612520A
Hybrid expert visual question-answering method and system based on strong visual semantics
CN118070816A
Cited By
Underwater sound target recognition system and method based on multi-modal depth feature fusion
CN121789648A