Autistic child speech recognition optimization method and system based on feature fusion and multi-task learning, and application

Through the feature fusion and multi-task learning method of dual feature encoder, the problem of insufficient generalization ability in dealing with autistic children with autism is solved, and more accurate speech recognition and language development evaluation are achieved.

CN120183443APending Publication Date: 2025-06-20EAST CHINA NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510335467.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The existing automatic speech recognition (ASR) model is insufficient in generalization when dealing with the speech of children with autism, making it difficult to capture the unique speech characteristics of children with autism, resulting in a high recognition error rate.

Method used

The feature fusion and multi-task learning method of dual feature encoder are used to extract the common phonological features and the unique phonological features of autistic children through the general feature encoder and ASD feature encoder, and feature fusion is performed through the attention mechanism. At the same time, language development assessment tasks were introduced as auxiliary training objectives, and supervision signals were constructed using Mullen evaluation data to guide the model to learn pronunciation characteristics related to language ability.

Benefits of technology

Through feature fusion and multi-task learning, the ASR model's transcription ability of the autistic children is improved, the recognition error rate is reduced, and the generalization ability of the model is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005321589990000033
    Figure BDA0005321589990000033
  • Figure BDA0005321589990000044
    Figure BDA0005321589990000044
  • Figure BDA0005321589990000051
    Figure BDA0005321589990000051
Patent Text Reader

Abstract

The invention discloses a feature fusion and multi-task learning-based voice recognition optimization method for children with autism, and the method optimizes the voice recognition result of the children with autism through the feature fusion and multi-task learning of a double-feature encoder. Comprising the following steps: step 1, performing feature extraction on input voice signals of autistic children by using a double-feature encoder; step 2, splicing the features obtained in the step 1 on a feature dimension; generating an attention weight through a full connection layer, and carrying out weighted adjustment on the features to realize feature fusion; and step 3, taking the fusion features in the step 2 as input, performing task optimization through multi-task learning, and obtaining an optimized speech recognition result and a language development condition. The invention further discloses an optimization system for implementing the optimization method, and the optimization system has a wide application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of speech recognition, and relates to an optimized method, system and application for speech recognition of autistic children based on feature fusion and multi-task learning. Background Art

[0002] Existing automatic speech recognition (ASR) models are mainly designed for the speech of the general population [1-3] , and the speech transcription effect for autistic children (ASD) is significantly limited. The speech of autistic children has particularities, such as unclear pronunciation, abnormal intonation, repetitive language, etc. [4-6] , while general ASR models only adopt a single feature encoder and are difficult to capture both general speech features and speech features unique to ASD at the same time. In addition, the training objective of existing models is single (only optimizing the mapping from speech to text), lacking effective utilization of the implicit language development information in speech, resulting in insufficient generalization ability of the model in the ASD speech scenario. During training, traditional automatic speech recognition (ASR) systems are usually trained based on the speech of adults or typical normal children, lacking targeted modeling of the speech characteristics of ASD children. Therefore, when processing the speech of ASD children, a relatively high recognition error rate often occurs.

[0003] Wav2Vec 2.0 [7] is a powerful self-supervised learning speech feature extraction model. By pre-training on a large amount of unlabeled speech data, it can directly extract highly expressive speech features from the original waveform. Its end-to-end modeling method makes it an ideal general speech feature encoder. Existing research on dual feature encoders for speech recognition based on Wav2Vec 2.0 aims to improve the accuracy of speech recognition for children with autism spectrum disorder (ASD). However, the speech of autistic children often shows different characteristics from those of typical children in terms of pronunciation, intonation, and rhythm, such as unclear pronunciation, abnormal intonation, or too fast speech rate. At the same time, the design of Wav2Vec 2.0 does not specifically consider the speech characteristics of ASD children. Therefore, when facing ASD speech, it cannot effectively capture their unique pronunciation and intonation abnormalities. Summary of the Invention

[0004] In order to solve the deficiencies of the existing technology, the purpose of the present invention is to provide an optimized method, system and application for speech recognition of autistic children based on feature fusion and multi-task learning. The present invention aims at the speech recognition of autistic children (ASD), and through the feature fusion and multi-task learning of a dual feature encoder, improves the transcription ability of the automatic speech recognition (ASR) model for the speech of autistic children.

[0005] In the present invention, the optimized method for speech recognition of autistic children includes the following steps:

[0006] Step 1: Use a dual - feature encoder to extract features from the input speech signals of children with autism respectively;

[0007] Step 2: Concatenate the features obtained in Step 1 in the feature dimension; generate attention weights through a fully - connected layer, and perform weighted adjustment on the features to achieve feature fusion;

[0008] Step 3: Use the fused features in Step 2 as input, and optimize the task through multi - task learning to obtain the optimized speech recognition results and language development status.

[0009] Specifically, the dual - feature encoder used in the present invention consists of two encoders: one is a general - feature encoder, which uses the pre - trained ASR model of Wav2Vec 2.0 and aims to extract general features applicable to all speech, and processes the original audio of the input speech signals of children with autism; the other is an ASD - feature encoder, which is specifically designed for the speech of children with autism and aims to capture specific abnormal features in ASD speech, such as unclear pronunciation, abnormal intonation, etc., and processes the Mel spectrogram of the input speech signals of children with autism. By designing a feature - fusion strategy, these two types of features are effectively fused based on the attention mechanism to generate the final feature representation, thereby improving the recognition performance of ASD speech. At the same time, a multi - task learning method is adopted, and a language development assessment task is introduced as an auxiliary training target. The supervision signal is constructed using the language development scores of children with autism (such as Mullen assessment data) to guide the model to more accurately learn the speech features related to language ability, thereby indirectly improving the adaptability of the ASR model to the speech of children with autism.

[0010] The goal of the ASR task is to transcribe the speech of children with ASD into text. This task trains the model to learn the mapping relationship between speech and text, and uses the traditional cross - entropy loss function to minimize the error of the model during speech transcription. With a large amount of speech - to - text labeled data, only using the traditional cross - entropy loss can effectively learn the mapping between speech and text and provide accurate speech recognition output. However, in the speech recognition of children with ASD, due to the difficulty of obtaining a large amount of labeled data of children with ASD speech, if only the traditional cross - entropy loss is used, it is difficult for the model to learn the unique language features of children with ASD. In order to enable the model to better focus on the language development features of children with ASD, the present invention introduces a language development assessment task to achieve the learning of ASD - specific features through predicting and learning the language development level of children with ASD. The language development assessment task is based on the Mullen assessment scale [8]The provided language development score requires the model to not only identify the text content in the speech but also evaluate the child's language ability by analyzing the speech features. This task is trained using the mean squared error (MSE) loss function between the language development score from the Mullen assessment scale and the score predicted by the model, with the goal of minimizing the difference between the predicted language development score of the model and the true score. The introduction of this task enables the model to extract information related to the language development level from the speech features, thereby improving the model's transcription ability in the ASR scenario.

[0011] The innovations of the present invention mainly include:

[0012] Dual-feature encoder design: Extract general speech features and speech features unique to children with autism spectrum disorder (ASD) through a general feature encoder and an ASD feature encoder respectively, and generate a more robust feature representation through a fusion strategy.

[0013] Multi-task learning: Introduce the language development assessment task, use the language development score provided by the Mullen assessment scale as a supervision signal, further improve the model's transcription ability for the speech of children with autism spectrum disorder, avoid overfitting of the model in a single task, and improve the generalization ability.

[0014] In the present invention, when training and optimizing the ASR task in multi-task learning, the ASR training data of each child used is manually annotated by medical experts to obtain the data annotated by medical experts; in addition, the audio is segmented and aligned according to the time of each sentence, and the same Mullen assessment score is used for the same child, that is, the Mullen assessment score corresponding to each sentence of the same child is the same.

[0015] The method mentioned in the present invention mainly includes the following parts:

[0016] ① General feature encoder: Based on wav2vec2.0

[0017] Load the pre-trained wav2vec2.0 model. Since the number of children with special needs is limited, to avoid overfitting, the present invention chooses to freeze all the parameters of wav2vec 2.0 and use it only as a general feature extractor. Input the audio data, and obtain the general feature sequence through the general feature encoder (wav2vec 2.0) T is the number of time steps, and D represents the feature dimension.

[0018] ② ASD feature extractor: Based on CNN

[0019] Initialize a convolutional neural network (CNN) to extract ASD-specific features. Input the mel spectrogram of the audio, and obtain the ASD feature sequence through the ASD feature extractor The time steps and dimensions of the ASD feature sequence need to be aligned with the time step T and dimension D of the general feature sequence; if the time steps are inconsistent, truncate the ASD feature sequence to align it with the time step of the general feature sequence F obtained by the general feature extractor. generic of the general feature sequence.

[0020] ③ Feature fusion module: Design a feature fusion based on the attention mechanism

[0021] Concatenate the general feature sequence and the ASD feature sequence along the feature dimension to obtain a joint feature matrix:

[0022]

[0023] Generate attention scores through a learnable fully connected layer (attention network):

[0024] S = ReLU(W a ·F contact + b a )

[0025] where ReLU is the activation function, which performs a non-linear transformation on the audio features, is the weight matrix of the learnable fully connected layer, is the bias term, H is the hidden layer dimension, that is, the size of the middle layer of the attention network, and the output score matrix

[0026] Perform a linear transformation on the score matrix S to map it to scalar attention weights:

[0027] A = W v ·S + b v

[0028] where, represents the linear transformation matrix that maps the hidden layer to scalar weights, is the bias term; the output represents the attention weight at each time step.

[0029] Perform softmax normalization on the attention weights to ensure that the weights sum to 1, and finally obtain the normalized attention weight vector

[0030]

[0031] where, α t ∈[0,1] represents the attention weight at the t-th time step.

[0032] According to the normalized attention weights, dynamically adjust the contributions of the general features and the ASD features, and finally obtain the fusion features containing general speech information and ASD-specific information

[0033] F fused = α⊙F generic +(1 - α)⊙F ASD

[0034] ⊙ represents the Hadamard product at each time step.

[0035] ④ Multi - task learning module

[0036] Optimize the ASR task and / or the language development assessment task;

[0037] ASR decoder: The Transformer decoder is used for speech transcription. Input F fused into the decoder to generate text predictions.

[0038] Language development assessment network: Initialize a regression network (fully - connected layer) for predicting the Mullen language development score. Perform temporal dimension pooling (such as average pooling) on F fused to obtain a global feature vector, and predict the Mellen language development score through the fully - connected regression network.

[0039] During the multi - task learning process, the present invention uses the language development task to provide additional supervision signals to guide the model to extract deep features related to language ability from speech. By jointly performing the ASR task and the language development task, the model can more comprehensively model the correlation between speech and language ability and avoid overfitting to a single task.

[0040] ⑤ Loss calculation

[0041] ASR loss: Calculate the cross - entropy loss L between the transcribed text and the true label ASR .

[0042] Language development loss: Calculate the mean squared error L between the predicted score and the true Mullen score language .

[0043] Joint loss: Weighted sum (formula is as follows), where β is a hyperparameter.

[0044] L = βL ASR +(1 - β)L language

[0045] The cross - entropy loss between the transcribed text and the true label is expressed as follows:

[0046]

[0047] where M represents the number of words in the dictionary; y i represents the true label in one - hot encoding; Represents the probability of the character predicted by the model, that is, the output value of softmax;

[0048] The mean square error between the predicted score and the true Mullen score

[0049]

[0050] Where N represents the total number of training samples; Represents the true Mullen developmental score of the i-th sample; Represents the language development score of the i-th sample predicted by the model.

[0051] During the backpropagation and parameter update process, fix wav2vec 2.0 and only update the parameters of the ASD feature encoder, feature fusion module, and multi-task module.

[0052] The speech recognition optimization method in the present invention utilizes the wav2vec2.0 model, which is pre-trained on large-scale unlabeled speech data and can extract general speech features with high expressiveness. It is especially suitable for the speech scenarios of ASD children with scarce data and has the ability of self-supervised training; in the design of the optimization method, end-to-end modeling is used to directly extract features from the original audio, avoiding the limitations of manually designed features.

[0053] In a specific implementation process of the present invention, in the Chinese speech recognition task, the character error rate (CER) is used as the evaluation index, and the calculation method is as follows:

[0054]

[0055] Where S c Represents the number of characters that are inconsistent with the corresponding positions of the true text in the prediction result; D c Represents the number of characters that exist in the true text but do not exist in the prediction result; I c Represents the number of extra characters that appear in the prediction result but do not exist in the true text; L represents the total number of characters in the true text.

[0056] The present invention also provides an optimization system for implementing the above optimization method. The optimization system includes: a feature extraction and fusion module, a multi-task learning module;

[0057] The feature extraction and fusion module includes a general feature encoder, an ASD feature encoder, and a feature fuser, which are used to extract general feature sequences and ASD feature sequences, and perform weighted fusion through an attention mechanism to generate a fusion feature representation suitable for speech recognition and language development evaluation;

[0058] The multi-task learning module includes a Transformer decoder and a regression network, and is used to perform automatic speech recognition tasks and language development assessment tasks.

[0059] In a specific implementation, the optimization system may further include: a voice acquisition module and an optimization module;

[0060] The voice acquisition module is used to receive the voice input of autistic children through an audio acquisition device for subsequent feature extraction;

[0061] The optimization module adjusts and updates the trainable parameters of the feature extraction and fusion module and the multi-task learning module based on backpropagation, optimizes the model weights, and improves the system performance.

[0062] The present invention also provides the above optimization method, or the application of the above optimization system in improving the speech recognition accuracy of children, predicting speech development scores, etc.

[0063] The beneficial effects of the present invention include: The present invention discloses an optimization method for speech transcription of autistic children. Through the feature fusion of a dual feature encoder and multi-task learning, it aims to improve the transcription ability of the ASR model for the speech of autistic children. The present invention designs a general feature encoder and an ASD feature encoder to extract general speech features and speech features unique to autistic children respectively, aiming to enhance the adaptability to heterogeneous speech (such as strengthening the contribution of ASD features to pronunciation-blurred segments, effectively coping with the diversity of the speech of autistic children, such as repetitive language and non-standard pronunciation). The present invention jointly optimizes the ASR task and the language development assessment task through a multi-task learning method, aiming to achieve the collaborative optimization of multi-task learning, and at the same time, enhance the interpretability of the model. The present invention directly processes the original children's speech, reducing the dependence on labeled data. The present invention has scalability and can be migrated to other special population speech recognition scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0065] Figure 1 is a flowchart of the method of the present invention.

[0066] Figure 2 is a structural diagram of the speech recognition optimization of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0067] The present invention will be further described in detail below in conjunction with the following specific embodiments and accompanying drawings. The processes, conditions, experimental methods, etc. for implementing the present invention are all common knowledge and well-known common sense in the art except for the specifically mentioned content below, and the present invention has no particularly restricted content.

[0068] The following are the explanations of the terms involved in the present invention:

[0069] Autism Spectrum Disorder(ASD): Autism Spectrum Disorder.

[0070] Automatic Speech Recognition(ASR): Automatic Speech Recognition

[0071] In existing speech recognition technologies (such as wav2vec 2.0), it is difficult for a single feature encoder to simultaneously capture general speech features and the speech features unique to children with autism. In addition, the training objectives of existing models are single, mainly optimizing the mapping from speech to text, and the implicit language development information in speech is not effectively utilized, resulting in insufficient generalization ability of the model when processing the speech of children with autism.

[0072] To solve this problem, the present invention proposes an optimization method for speech transcription of children with autism. This method improves the transcription ability of the ASR model for the speech of children with autism through the feature fusion of a dual feature encoder and a multi-task learning method. By fusing general speech features and the speech features unique to children with autism, the speech features of children with ASD are enriched, thus obtaining a more robust feature representation. At the same time, the multi-task learning method is adopted, introducing the language development assessment task as an auxiliary training objective, and constructing a supervision signal using the language development scores of children with autism (such as Mullen assessment data) to guide the model to more accurately learn the speech features related to language ability, thereby indirectly improving the adaptability of the ASR model to the speech of children with autism.

[0073] I. Data Preparation

[0074] Collect the speech data of 50 children with autism spectrum disorder, about 30 minutes of audio per person, sampling rate 16kHZ, and the format is WAV. Each audio is annotated with the corresponding transcribed text, and the language development score is provided through the Mullen Early Assessment Scale. The data is divided into a training set and a test set. The training set includes the data of 40 children (about 20 hours), and the test set includes the data of 10 children (about 5 hours).

[0075] The general feature encoder used in the present invention selects the pre-trained wav2vec2.0-base model (https: / / huggingface.co / facebook / wav2vec2-base), fixes all parameters, and only serves as a static feature extractor.

[0076] wav2vec 2.0 is a deep learning model for end-to-end speech recognition tasks, proposed by the Facebook AI Research Institute. It is the successor version of the wav2vec model and is also based on the contrastive learning framework, but the model has been optimized. The entire model mainly includes the following parts:

[0077] ① Obtain the latent representation of the audio: The original audio is input into a multi-layer (7 layers in the experiment) convolutional neural network to extract the latent representation, obtaining latent representation vectors for multiple time steps;

[0078] ② Obtain a high-quality representation of the audio containing context information: wav2vec 2.0 introduces a Masked Predictive Learning task, similar to the masked language model task in the BERT model. In the pre-training stage, the model needs to predict the content of the randomly masked audio segments.

[0079] ③ Quantization: Convert the latent representation obtained in step ① into a discrete representation, which is achieved through the Gumbel-Softmax technique. During the quantization process, a diversity loss is used.

[0080] ④ Training loss function: The weighted sum of the contrastive loss and the diversity loss.

[0081] - Contrastive loss: The similarity between the context representation ② and the quantized representation ③. The goal is for ② to have a high similarity with the quantized positive sample.

[0082] - Diversity loss: The goal is to make each code (or vector) in each codebook be evenly used. Ideally: the frequency of each code in the codebook being selected should be close to the uniform distribution 1 / N (where N is the number of codes).

[0083] After the input audio is processed by wav2vec2.0, a general feature sequence is output (T is the number of time steps).

[0084] The ASD feature encoder uses a convolutional neural network, and the input is the Mel spectrogram of the audio (80-dimensional Mel filter bank, window length of 25 ms, frame shift of 10 ms). Convolutional neural network design: The first layer is Conv2D (number of filters = 64, kernel size = 3×3, activation function = ReLU); the second layer is MaxPooling2D (pooling size = 2×2); the third layer is Conv2D (number of filters = 128, kernel size = 3×3, activation function = ReLU); the fourth layer is global average pooling + fully connected layer (output dimension = 768); finally, the ASD feature sequence is output

[0085] The feature fusion module first ensures that the time steps T of F ASD are consistent with those of F generic .

[0086] Concatenate the general feature sequence and the ASD feature sequence:

[0087]

[0088] Generate attention scores through a fully connected layer:

[0089] S = ReLU(W a ·F contact + b a ),

[0090] Map to scalar weights:

[0091] A = W v ·S + b v ,

[0092] Softmax normalization:

[0093]

[0094] Feature fusion:

[0095] F fused = α ⊙ F generic + (1 - α) ⊙ F ASD

[0096] The multi-task learning module includes the ASR task and the language development assessment task. In the ASR task, the decoder uses a Transformer decoder (4 layers, 4-head attention, hidden layer dimension of 768). The loss function uses cross-entropy loss L ASR : The language development assessment task uses a regression network (global average pooling + fully connected layer, with an output dimension of 1) to predict the Mullen language development score. Loss function: mean squared error loss Combined loss: L = 0.7L ASR + 0.3L language .

[0097] In the present invention, the attention mechanism can be used to dynamically adjust the weights of different features and adaptively fuse general speech features and speech features of children with ASD according to the input content. This dynamic nature enables the model to flexibly capture key information.

[0098] In addition to the attention mechanism, in the actual implementation process, methods such as weighted average, concatenation + fully connected layer, and gating mechanism can also be considered for feature fusion; however, the above methods have disadvantages such as static weights that cannot be adjusted according to the input (weighted average), may introduce redundant information and have high computational complexity (concatenation + fully connected layer), and lower flexibility than the attention mechanism (gating mechanism), and the overall performance is lower than that of the attention mechanism.

[0099] References

[0100] [1] Radford A, Kim J W, Xu T, et al. Robust speech recognition via large-scale weak supervision[C] / / International conference on machine learning. PMLR, 2023: 28492 - 28518.

[0101] [2] Gao Z, Li Z, Wang J, et al. Funasr: A fundamental end-to-end speech recognition toolkit[J]. arXiv preprint arXiv:2305.11013, 2023.

[0102] [3] Zhang B, Wu D, Peng Z, et al. Wenet 2.0: More productive end-to-end speech recognition toolkit[J]. arXiv preprint arXiv:2203.15455, 2022.

[0103] [4] Zhao Jinzhu, Tang Lina, He Tianyi, et al. Language development characteristics of children with autism spectrum disorder[J]. Chinese Journal of Child Health Care, 2021, 29(09): 969 - 972.

[0104] [5] Hu Xinyu, Xu Lian, Liu Min, et al. Comparative study on language abilities of preschool children with autism spectrum disorder and developmental delay [J]. Chinese Scientific Journal of Hearing and Speech Rehabilitation, 2023, 21(04): 348 - 351.

[0105] [6] Pan Xiuyu, Li Honghua, Wang Bing, et al. Analysis of language development status of children with autism spectrum disorder [J]. Journal of Educational Biology, 2021, 9(04): 262 - 265 + 295.

[0106] [7] Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael Auli: wav2vec2.0: A Framework for Self - Supervised Learning of Speech Representations. NeurIPS 2020

[0107] [8] E.M. Mullen, et al., Mullen scales of early learning, AGS Circle Pines, MN, 1995.

[0108] The protection scope of the present invention is not limited to the above embodiments. Without departing from the spirit and scope of the inventive concept, changes and advantages that can be conceived by those skilled in the art are included in the present invention, and the scope of protection is defined by the appended claims.

Claims

1. A method for optimizing speech recognition for children with autism based on feature fusion and multi-task learning, characterized in that: The method optimizes the speech recognition results of autistic children through feature fusion of dual feature encoders and multi-task learning; and comprises the following steps: Step 1: Use a dual feature encoder to extract features from the input speech signal of the autistic child; Step 2: Concatenate the features obtained in step 1 in the feature dimension; generate attention weights through the fully connected layer, and perform weighted adjustment on the features to achieve feature fusion; Step 3: Use the fused features in step 2 as input, perform task optimization through multi-task learning, and obtain optimized speech recognition results and language development.

2. The method according to claim 1, characterized in that In step 1, the dual feature encoder includes a general feature encoder and an autism spectrum disorder ASD feature encoder; The universal feature encoder processes the original audio of the input autistic child speech signal to obtain a universal feature sequence; the autism spectrum disorder ASD feature encoder processes the Mel spectrum map of the input autistic child speech signal to obtain an autism spectrum disorder feature sequence.

3. The method according to claim 2, characterized in that The universal feature encoder uses a pre-trained wav2vec2.0 model as a static feature extractor to output a universal feature sequence; The ASD feature encoder uses a convolutional neural network, and outputs an ASD feature sequence after processing including a convolution layer, a pooling layer, and a connection layer; and / or, The convolutional neural network includes a first convolutional layer, a maximum pooling layer, a second convolutional layer, an average pooling layer and a fully connected layer.

4. The method according to claim 1, characterized in that In step 2, before feature fusion, it is necessary to ensure that the time step of the ASD feature sequence is consistent with the time step of the general feature sequence; when the time steps are inconsistent, the ASD feature sequence is truncated to align the time steps; and / or, Step 2 further includes: Step 2.

1. Concatenate the universal feature sequence and the ASD feature sequence: in, represents the speech feature matrix extracted by the general feature encoder; represents the speech feature matrix unique to ASD children extracted by the ASD feature encoder; T is the number of time steps; D represents the feature dimension; Step 2.

2. Generate attention scores through a learnable fully connected layer: S=ReLU(W a ·F contact +b a ); Among them, ReLU is the activation function; is the learnable fully connected layer weight matrix; is the bias term; H is the hidden layer dimension; Step 2.

3. Perform a linear transformation on the attention score matrix S and map it to a scalar attention weight: A=W v ·S+b v ; in, A linear transformation matrix that maps hidden layers to scalar weights; is the bias term, outputting the attention weight for each time step Step 2.

4. Normalize the scalar attention weights in step 2.3: α t ∈[0,1] represents the attention weight at the tth time step; Step 2.

5. Perform feature fusion based on the normalized attention weights: F fused =α⊙F generic +(1-α)⊙F ASD ; where ⊙ represents the Hadamard product at each time step.

5. The method according to claim 1, characterized in that In step 3, the multi-task learning is used to optimize the ASR task and / or the language development assessment task; The ASR task generates speech recognition text prediction results by inputting fusion features into the Transformer decoder; The language development assessment task obtains a global feature vector by pooling the fused features in the time dimension, and uses a fully connected regression network to predict the Mullen language development score.

6. The method according to claim 5, characterized in that Optimize the ASR task and language development assessment task through a joint loss function: L=βL ASR +(1-β)L language , Among them, L ASR represents the cross entropy loss between the transcribed text and the true label, L language represents the mean square error between the predicted score and the true Mullen score, and β is a hyperparameter; The cross entropy loss between the transcribed text and the true label is expressed as follows: Where M represents the number of words in the dictionary; y i represents the true label; Represents the probability of the word predicted by the model; The mean square error between the predicted score and the true Mullen score is expressed as follows: Where N represents the total number of training samples; represents the true Mullen developmental score of the i-th sample; Represents the language development score of the i-th sample predicted by the model.

7. The method according to claim 1, characterized in that During the back-propagation and parameter update process, wav2vec2.0 is fixed, and only the parameters of the ASD feature encoder, feature fusion module, and multi-task module are updated; and / or, The word error rate CER is used as the evaluation indicator in the Chinese speech recognition task, which is expressed as follows: Among them, S c Indicates the number of inconsistencies between the predicted result and the corresponding position characters in the real text; D c Indicates the number of characters that exist in the real text but not in the predicted result; I c represents the number of extra characters that appear in the predicted results but do not exist in the real text; L represents the total number of characters in the real text.

8. An optimization system for implementing the optimization method according to any one of claims 1 to 7, characterized in that: The optimization system includes: a feature extraction and fusion module and a multi-task learning module; The feature extraction and fusion module includes a universal feature encoder, an ASD feature encoder, and a feature fuser, which are used to extract universal feature sequences and ASD feature sequences, and perform weighted fusion through an attention mechanism to generate a fused feature representation suitable for speech recognition and language development assessment; The multi-task learning module includes a Transformer decoder and a regression network, which is used to perform automatic speech recognition tasks and language development assessment tasks.

9. The optimization system according to claim 8, characterized in that: The optimization system also includes: a voice collection module and an optimization module; The voice collection module is used to receive the voice input of the autistic child through the audio collection device for subsequent feature extraction; The optimization module adjusts and updates the trainable parameters of the feature extraction fusion module and the multi-task learning module based on back propagation, optimizes the model weights, and improves system performance.

10. Use of the optimization method according to any one of claims 1 to 7, or the optimization system according to claim 8 or 9, in improving the accuracy of children's speech recognition and predicting speech development scores.