Model training method, speech recognition method and related equipment

By encoding the features and training the loss values ​​of audio and annotated text, the problem of low vocabulary recognition accuracy of traditional speech recognition models in professional fields is solved, and the model training efficiency and recognition accuracy are improved.

CN120708601APending Publication Date: 2025-09-26MASHANG CONSUMER FINANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510312975.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Traditional speech recognition models have low vocabulary recognition accuracy in professional fields, resulting in insufficient accuracy in speech recognition.

Method used

By encoding the features of the words in the first audio and the annotated text, predicting the probability value of the output word, and training the model based on the loss value, and using the vocabulary training model in the preset field, the training efficiency and recognition accuracy of the model are improved.

Benefits of technology

It improves the model training efficiency and speech recognition accuracy, especially in vocabulary recognition in professional fields, and enhances the recognition efficiency and accuracy of audio text.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708601A_ABST
    Figure CN120708601A_ABST
Patent Text Reader

Abstract

The invention relates to artificial intelligence, and provides a model training method, a speech recognition method and related equipment. The model training method comprises the following steps: carrying out feature coding on a first audio to obtain a first coding feature, and carrying out feature coding on a first vocabulary in a labeled text of the first audio to obtain a second coding feature; predicting a probability value of outputting the first vocabulary according to the first coding feature and the second coding feature; calculating a first loss value according to the first coding feature and the probability value; a model is trained based on the first loss value. The method can improve the training effect of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and specifically to a model training method, a speech recognition method, and related equipment. Background Art

[0002] With the development of artificial intelligence, automatic speech recognition technology is constantly improving. Speech recognition models can directly convert audio into text, significantly improving work efficiency. However, when it comes to audio recognition in some professional fields, traditional speech recognition models struggle to accurately identify vocabulary in relevant fields, resulting in low speech recognition accuracy. Summary of the Invention

[0003] The present application provides a model training method, a speech recognition method and related equipment to solve the technical problem that the vocabulary in the relevant professional field cannot be accurately recognized, resulting in low accuracy of speech recognition.

[0004] A first aspect of an embodiment of the present application provides a model training method, which includes: feature encoding a first audio to obtain a first encoding feature, and feature encoding a first word in an annotated text of the first audio to obtain a second encoding feature; predicting and outputting a probability value of the first word based on the first encoding feature and the second encoding feature; calculating a first loss value based on the first encoding feature and the probability value; and training a model based on the first loss value.

[0005] A second aspect of an embodiment of the present application provides a speech recognition method, which includes: recognizing a second audio, obtaining a third word in the text of the second audio and time frame information of the third word in the second audio; extracting an audio segment from the second audio based on the time frame information; and determining a fifth word in the text of the second audio from the third word and a preset fourth word based on the audio segment.

[0006] A third aspect of an embodiment of the present application provides a model training device, which includes: an encoding unit for feature encoding a first audio to obtain a first encoding feature, and feature encoding a first word in an annotated text of the first audio to obtain a second encoding feature; a prediction unit for predicting and outputting a probability value of the first word based on the first encoding feature and the second encoding feature; a calculation unit for calculating a first loss value based on the first encoding feature and the probability value; and a training unit for training a model based on the first loss value.

[0007] A fourth aspect of an embodiment of the present application provides a speech recognition device, comprising: a recognition module for recognizing a second audio, obtaining a third word in the text of the second audio and time frame information of the third word in the second audio; an extraction module for extracting an audio segment from the second audio based on the time frame information; and a determination module for determining a fifth word in the text of the second audio from the third word and a preset fourth word based on the audio segment.

[0008] A fifth aspect of an embodiment of the present application provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method provided in the first or second aspect above when executing the computer program.

[0009] A sixth aspect of an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the method provided in the first or second aspect are implemented.

[0010] A seventh aspect of the embodiments of the present application provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the steps in the method provided in the first or second aspect above.

[0011] In the model training method of this embodiment, by feature encoding the first word in the annotated text of the first audio, since it is not necessary to feature encode the entire annotated text, the efficiency of obtaining the second encoding feature can be improved. At the same time, the dimension of the second encoding feature obtained by feature encoding the first word is smaller than the encoding feature obtained by feature encoding the entire annotated text. Therefore, when predicting and outputting the probability value of the first word based on the second encoding feature, the amount of calculation can be reduced, thereby improving the efficiency of determining the probability value, and further improving the efficiency of model training. In addition, using the first audio corresponding to the annotated text including the first word to participate in the training of the model can avoid invalid training of the model due to the fact that the first word is not included in the annotated text, thereby improving the training effect of the model.

[0012] In the speech recognition method of this embodiment, by recognizing the second audio, the third word in the text of the second audio and the time frame information of the third word in the second audio can be obtained. Based on the time frame information, the audio segment containing the third word can be accurately extracted from the second audio. Based on the audio segment, the fifth word in the text of the second audio can be determined from the third word and a preset fourth word, thereby improving the recognition accuracy of the fifth word in the text of the second audio. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments of the present application. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0014] Figure 1 This is a schematic diagram of an application scenario of a model training method and a speech recognition method provided in an embodiment of the present application; Figure 2 This is a flow chart of a model training method provided in an embodiment of the present application; Figure 3 This is a flowchart of another model training method provided in an embodiment of the present application; Figure 4 This is a flow chart of a speech recognition method provided by an embodiment of the present application; Figure 5 It is a structural diagram of the model provided in the embodiment of the present application; Figure 6 This is a functional module diagram of a model training device provided in an embodiment of the present application; Figure 7 This is a functional module diagram of a speech recognition device provided by an embodiment of the present application; Figure 8 It is a structural diagram of an electronic device for implementing a model training method and a speech recognition method provided in an embodiment of the present application. DETAILED DESCRIPTION

[0015] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0016] In the following, the terms "first" and "second" are used for descriptive purposes only and are not to be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or design schemes. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a concrete way.

[0017] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application belongs. The terms used in the specification of this application are for the purpose of describing specific embodiments only and are not intended to limit this application. It should be understood that, unless otherwise specified in this application, " / " means or. For example, A / B can mean A or B. "And / or" in this application is merely a way to describe the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. "At least one" means one or more. "Multiple" means two or more than two. For example, at least one of a, b or c can mean: a, b, c, a and b, a and c, b and c, a, b and c.

[0018] Some terminology explanations: End-to-end speech recognition: refers to a speech recognition method that directly converts audio sequences into text sequences. This method is different from the multi-module recognition method used by traditional speech recognition. Traditional speech recognition is generally divided into an acoustic module and a language module. The acoustic module is responsible for converting audio sequences into phoneme sequences, and the language module is responsible for converting these phoneme sequences into text sequences.

[0019] Transformer model: A time series model with a self-attention mechanism. It can effectively encode time series information in the encoding network (i.e., the attention layer and the feedforward network layer). It has good processing capabilities and high speed for time series information, and is therefore widely used in fields such as natural language processing, computer vision, machine translation, and speech recognition.

[0020] Conformer model: A network structure used for end-to-end speech recognition tasks. It combines the convolutional network and the Transformer model. A convolutional module is introduced after the multi-headed self-attention (MHSA) layer of the Transformer model. The MHSA layer captures global context information, and the convolutional module extracts local features, thereby better achieving unified modeling of global and local features.

[0021] The Connectionist Temporal Classification (CTC) loss function is specifically designed for sequence learning problems. Traditional sequence labeling algorithms require perfect alignment of input and output symbols at every moment. CTC, however, expands the label set and adds empty elements. After the sequence is labeled with the expanded label set, all predicted sequences that can be converted to true sequences through a mapping function are correct predictions, eliminating the need for data alignment.

[0022] Hot words: refers to words that frequently appear in specific scenarios. In the embodiment of the present application, hot words may include words related to the address field, such as "XX District, XX City".

[0023] In end-to-end speech recognition technology, deep networks often have stronger generalization capabilities. However, related speech recognition models are extremely dependent on training data. Especially when it comes to audio recognition in specialized fields, traditional speech recognition models struggle to accurately identify vocabulary in those fields, resulting in low speech recognition accuracy.

[0024] Based on the above problems, in order to improve the accuracy of speech recognition, the embodiment of the present application provides a model training method for improving the training effect of the model. In addition, the embodiment of the present application also provides a speech recognition method that can improve the recognition efficiency and accuracy of words in audio text.

[0025] See also Figure 1 , Figure 1 Schematic diagram of an application scenario of a model training method and a speech recognition method provided in an embodiment of the present application. The scenario may include various electronic devices 100 and a server 200.

[0026] The electronic device 100 can be a mobile phone, a tablet computer, a desktop computer, a laptop computer, a handheld computer, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a cellular phone, a personal digital assistant (PDA), an augmented reality (AR) device, a virtual reality (VR) device, an artificial intelligence (AI) device, a wearable device, an in-vehicle device, a smart home device and / or a smart city device. The embodiments of the present application do not impose any special restrictions on the specific type of the electronic device 100.

[0027] Server 200 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, i.e., Content Delivery Network (CDN), as well as big data and artificial intelligence platforms, but is not limited to these.

[0028] It should be noted that the method in the embodiment of the present application can be performed independently by the electronic device 100 or the server 200, or can be performed jointly by the server 200 and the electronic device 100. When performed independently by the electronic device 100 or the server 200, the model training and application process can be implemented independently by the electronic device 100 or the server 200. For example, a trained model can be obtained by fine-tuning the model on the electronic device 100. Accordingly, after training, the electronic device 100 can use the trained model to recognize audio. The above process can also be performed independently by the server 200. When performed jointly by the server 200 and the electronic device 100, the server 200 can train the model and then deploy the trained model to the electronic device 100, and the electronic device 100 can implement the speech recognition process, or part of the model training or application process can be implemented by the electronic device 100, and part of the process can be implemented by the server 200, and the two can cooperate to implement the model training or application process. In actual application, specific configuration can be made according to the situation, and no specific limitation is made here.

[0029] It should be noted that when the model training method and speech recognition method provided in the embodiment of the present application are executed separately by the server 200 or the electronic device 100, the above-mentioned application scenario may also only include any single device in the server 200 or the electronic device 100, or the server 200 and the electronic device 100 may also be considered to be the same device. In actual application, when the model training method and speech recognition method provided in the embodiment of the present application are jointly executed by the server 200 and the electronic device 100, the server 200 and the electronic device 100 may also be the same device, that is, the server 200 and the electronic device 100 may be different functional modules of the same device, or virtual devices virtualized from the same physical device.

[0030] In one possible implementation, the user may provide audio through the electronic device 100, and the server 200 may use the speech recognition method of the embodiment of the present application to determine the domain-specific vocabulary contained in the text of the audio and return it to the electronic device 100 for presentation.

[0031] In the embodiment of the present application, the electronic device 100 and the server 200 can be directly or indirectly connected to each other through one or more networks. The network can be a wired network or a wireless network. For example, the wireless network can be a mobile cellular network or a Wireless Fidelity (Wi-Fi) network. Of course, it can also be other possible networks, and the embodiment of the present application does not limit this. It should be noted that Figure 1 The examples shown are just for illustration. In fact, the number of terminal devices and servers is not limited and is not specifically limited in the embodiments of this application.

[0032] The model training method and speech recognition method provided by the embodiments of the present application are described below in combination with the application scenarios described above and with reference to the accompanying drawings. It should be noted that the above application scenarios are only shown to facilitate understanding of the spirit and principles of the embodiments of the present application, and the embodiments of the present application are not limited in this respect.

[0033] like Figure 2 FIG. 1 is a flow chart of a model training method provided by an embodiment of the present application. The model training method is applied to electronic devices, for example, Figure 1 The electronic device 100. According to different requirements, the order of the steps in the flowchart can be changed, and some steps can be omitted.

[0034] S201 , feature encoding is performed on the first audio to obtain a first encoding feature, and feature encoding is performed on the first word in the annotated text of the first audio to obtain a second encoding feature.

[0035] In at least one embodiment of the present application, the first audio may include voice signals in a variety of application scenarios, such as customer service call quality inspection scenarios, voice text error correction scenarios, intelligent voice question and answer scenarios, etc.

[0036] In some embodiments, the electronic device can filter the corresponding voice signal as the first audio from the voice signal in any application scenario based on the vocabulary corresponding to the preset field. The vocabulary corresponding to the preset field can be set according to the purpose of the model. For example, if the model is used to identify address words, such as "XX District, XX City", the vocabulary corresponding to the preset field can include any words related to the address; if the model is used to identify words related to the insurance field, such as "premium", the vocabulary corresponding to the preset field can include any words related to the insurance field. The embodiment of the present application determines the first audio from the voice signal in any application scenario through the vocabulary corresponding to the preset field, which can ensure that the annotated text of the first audio used to train the model includes the vocabulary corresponding to the preset field, thereby improving the effectiveness of the first audio and enabling targeted training of the model.

[0037] Specifically, the electronic device captures audio from any user for vocabulary corresponding to a preset domain as user audio, and uses a sound analysis instrument to obtain the pitch frequency of each frame of speech from the user audio. The electronic device then performs feature encoding on the pitch frequency of each frame of speech to obtain a first voiceprint feature for the vocabulary corresponding to the preset domain. The electronic device then obtains second voiceprint features for multiple voice signals, where the dimension of the second voiceprint feature is greater than or equal to the dimension of the first voiceprint feature.

[0038] Furthermore, the electronic device extracts at least one third voiceprint feature from the second voiceprint feature, and the dimension of the third voiceprint feature is equal to the dimension of the first voiceprint feature. In one example, the electronic device can determine a segmentation window based on the dimension of the first voiceprint feature, and the electronic device uses the segmentation window and a preset spacing to segment the second voiceprint feature to obtain at least one third voiceprint feature. The preset spacing can be set and adjusted according to actual needs. For example, if the number of frames of the user audio is 10 frames, the first voiceprint feature can be a 10x256 vector; if the number of frames of the voice signal is 12 frames, the second voiceprint feature can be a 12x256 vector; therefore, the dimension of the segmentation window determined by the electronic device is: a 10x256 vector. If the preset spacing is 1, the electronic device can obtain 3 third voiceprint features, namely: the 1st row vector to the 10th row vector in the second voiceprint feature, the 2nd row vector to the 11th row vector in the second voiceprint feature, and the 3rd row vector to the 12th row vector in the second voiceprint feature. In another example, the electronic device can segment the second voiceprint feature based on the dimension of the first voiceprint feature to obtain a third voiceprint feature. Continuing with the above example, two third voiceprint features can be obtained: the first to tenth row vectors in the second voiceprint feature, and the ninth to eighteenth row vectors in the second voiceprint feature.

[0039] Furthermore, the electronic device calculates the similarity between each third voiceprint feature and the first voiceprint feature using the Cosine formula. If the similarity between the third voiceprint feature and the first voiceprint feature in any voice signal is greater than a similarity threshold, the voice signal is determined to be the first audio signal. In other embodiments, the electronic device may also calculate the similarity between each third voiceprint feature and the first voiceprint feature using the Euclidean distance formula, but practical applications are not limited to this.

[0040] In at least one embodiment of the present application, the annotated text may be text corresponding to the first audio. For example, if the first audio is a voice signal related to Mandarin, the annotated text may be a call text described in Chinese, etc. The electronic device may identify the first vocabulary from the annotated text based on the vocabulary corresponding to the preset field. Specifically, the electronic device calculates the first similarity between each vocabulary in the annotated text and the vocabulary corresponding to the preset field, and determines the vocabulary in the annotated text whose first similarity is greater than a first threshold as the first vocabulary, wherein the first threshold can be set and adjusted according to actual needs. The embodiment of the present application can quickly identify the first vocabulary from the annotated text through the vocabulary corresponding to the preset field.

[0041] In at least one embodiment of the present application, in the process of using the first audio to train the model, in order to enable the model to better capture the intrinsic correlation between the first audio and the first vocabulary in the annotated text of the first audio, while reducing the consumption of computing resources, the electronic device can use the same encoding network to perform feature encoding on the first audio and the first vocabulary respectively to obtain vector representations of the first audio and the first vocabulary on the same dimension.

[0042] In at least one embodiment of the present application, the electronic device may perform feature encoding on the first audio through the encoding network in the model. The encoding network may include m first encoding layers and n second encoding layers, where m and n are positive integers greater than or equal to 1, and the mth first encoding layer is connected to the first second encoding layer. The first encoding layer is used to capture local information of the first audio, and the second encoding layer is used to capture global information of the first audio. The attention dimensions of the first encoding layer and the second encoding layer are the same to meet the encoding requirements for the first audio. Each first encoding layer may include a convolutional layer, a first pooling layer, and a first fully connected layer, and each second encoding layer may include an attention layer, a second pooling layer, and a second fully connected layer.

[0043] In at least one embodiment of the present application, after the first audio is processed by the convolution layer in the first first coding layer, it is processed by the first pooling layer in the first first coding layer, and then processed by the first fully connected layer in the first first coding layer to obtain the coding features of the first first coding layer, and the coding features of the first first coding layer are passed to the second first coding layer until the coding features of the mth first coding layer are obtained. The coding features of the mth first coding layer are passed to the first second coding layer for processing until the coding features of the nth second coding layer are obtained, and the coding features of the nth second coding layer are used as the first coding features. The first coding feature can be a vector of NxMx256, where N can represent the number of first audios input to the coding network, M can represent the number of frames of the first audio, and 256 can represent the output dimension of the coding network.

[0044] In at least one embodiment of the present application, in order to improve the training effect and training efficiency of the model, the electronic device can perform feature encoding on the first audio through the encoding network in the model after the encoding network in the model is trained. The method of training the encoding network in the electronic device can refer to Figure 3 In the embodiment of the present application, since the encoding network has been pre-trained, it is ensured that the encoding network can accurately encode the features of the first audio, thereby avoiding the impact on the training of the model caused by the encoding network's inability to accurately represent the first audio, thereby improving the training effect of the model.

[0045] In at least one embodiment of the present application, the electronic device may perform feature encoding on the first word through a coding network to obtain a second coding feature. The second coding feature is generated in a similar manner to the first coding feature, and this application will not repeat the description.

[0046] S202: predict and output a probability value of the first word based on the first coding feature and the second coding feature.

[0047] In at least one embodiment of the present application, the probability value of the first word may include: a probability value of the first word output by the model for each frame of audio in the first audio. For example, if the first audio has 10 frames, then the probability value of the first word output by the model for the first audio is 10.

[0048] In at least one embodiment of the present application, the electronic device predicts and outputs the probability value of the first word based on the first coding feature and the second coding feature, including: the electronic device performs a linear transformation on the first coding feature to obtain the key vector and value vector corresponding to each attention head, and performs a linear transformation on the second coding feature to obtain the query vector corresponding to each attention head; the electronic device performs an attention operation on the query vector, the key vector, and the value vector to obtain the first vector output by each attention head; the electronic device determines the second vector of the model for the first audio based on the first vector output by each attention head; the electronic device determines the probability value of the first word based on the second vector. The embodiment of the present application analyzes the first coding feature and the second coding feature through multi-head attention, which allows the model to focus on the context information of the first word from multiple dimensions, thereby ensuring the accuracy of the probability value of the first word.

[0049] In some embodiments, the model may further include a recognition network, which may include a multi-head attention layer, a fusion layer, a third fully connected layer, and a decoding layer. The network parameters of the multi-head attention layer may include, but are not limited to, a first weight matrix, a second weight matrix, a third weight matrix, a first bias vector, a second bias vector, and a third bias vector corresponding to each attention head. The fusion layer may include weight coefficients corresponding to each attention head, and the network parameters of the third fully connected layer may include, but are not limited to, a fourth weight matrix and a fourth bias vector.

[0050] In some embodiments, when determining the key vector K, value vector V, and query vector Q corresponding to each attention head, the electronic device may perform a linear transformation on the first coding feature based on the first weight matrix and the first bias vector corresponding to each attention head to obtain the key vector K corresponding to each attention head. The electronic device may perform a linear transformation on the first coding feature based on the second weight matrix and the second bias vector corresponding to each attention head to obtain the value vector V corresponding to each attention head. The electronic device may perform a linear transformation on the second coding feature based on the third weight matrix and the third bias vector corresponding to each attention head to obtain the query vector Q corresponding to each attention head. The embodiment of the present application determines the key vector and the value vector based on the first coding feature, and determines the query vector based on the second coding feature, thereby realizing cross-level coupling between the acoustic features of the first audio and the first vocabulary.

[0051] In some embodiments, in the process of determining the second vector, the electronic device can fuse the first vectors output by all attention heads in the model through the weight coefficient corresponding to each attention head in the fusion layer to obtain an intermediate vector. The electronic device transforms the intermediate vector based on the fourth weight matrix and the fourth bias vector of the third fully connected layer to obtain a second vector, which is used to indicate the probability values ​​corresponding to multiple words output by the model for the first audio. The embodiment of the present application combines the first vectors output by all attention heads in the model to integrate information from multiple dimensions, which is conducive to accurately determining the second vector.

[0052] In some embodiments, during the process of determining the probability value of the first word, if the first word is included in the multiple words output by the model for the m-th frame of the first audio, the electronic device selects a corresponding element from the second vector as the probability value of the first word corresponding to the m-th frame based on the first word. If the first word is not included in the multiple words output by the model for the m-th frame of the first audio, the electronic device determines the probability value of the first word corresponding to the m-th frame as a second threshold value, where the second threshold value can be set and adjusted according to actual needs, for example, the second threshold value can be set to 0.

[0053] S203: Calculate a first loss value according to the first coding feature and the probability value.

[0054] In at least one embodiment of the present application, the first loss value can be used to measure the difference between the model's predicted value (e.g., the second word output by the model for the first audio) and the true value (e.g., the first word in the annotated text of the first audio). For example, the first loss value can be used to quantify the loss of a portion of the network layer in the model, for example, the first loss value can be used to quantify the loss of the recognition network in the model.

[0055] In at least one embodiment of the present application, the electronic device calculates a first loss value based on the first coding feature and the probability value, including: the electronic device obtains a third vector based on the first coding feature and the probability value; normalizes the third vector to obtain a fourth vector; decodes the fourth vector to obtain a second vocabulary output by the model for the first audio; and calculates the first loss value based on the first vocabulary and the second vocabulary.

[0056] In some embodiments, the electronic device can obtain a vector based on the probability values ​​of the first vocabulary corresponding to multiple frames in the first audio. For example, if the first encoding feature can be an NxMx256 vector, the vector corresponding to the probability values ​​of the first vocabulary can be an NxMx1 vector. The electronic device can calculate the dot product of the first encoding feature and the vector corresponding to the probability values ​​of the first vocabulary to obtain a third vector. The third vector is used to represent the weighted aggregate feature of the first encoding feature. Continuing with the above example, after calculation, the third vector can be represented as an NxM vector.

[0057] In some embodiments, in order to obtain the importance of each frame in the first audio, the electronic device may perform row normalization on the third vector to obtain a fourth vector. For example, if the third vector is , after normalizing each row vector in the third vector, the fourth vector can be obtained as .

[0058] In some embodiments, during the process of determining the second vocabulary, the electronic device may decode the fourth vector through a decoding layer in the recognition network to obtain the second vocabulary output by the model for each first audio.

[0059] In some embodiments, during calculation of the first loss value, the electronic device determines a first number of first audio tracks, where the first number indicates the number of first audio tracks corresponding to a second vocabulary different from the first vocabulary. The electronic device determines a second number N of first audio tracks in the input model and determines a first loss value based on the first number and the second number. The first loss value can be determined based on a ratio of the first number to the second number.

[0060] S204: Train the model based on the first loss value.

[0061] In at least one embodiment of the present application, during the process of training a model based on a first loss value, the electronic device may set a convergence condition for the model, for example, the first loss value must meet a preset range. If the first loss value is not within the preset range, the network parameters of the recognition network in the model (e.g., the network parameters of the multi-head attention layer, the network parameters of the fusion layer, and the network parameters of the third fully connected layer) are adjusted, and the model is iteratively trained until the first loss value is within the preset range, completing model training. The preset range can be customizable.

[0062] In other embodiments, during the process of training the model based on the first loss value, if the change trend of the first loss value becomes gentle and tends to be stable, the electronic device can determine that the model has converged.

[0063] In some other embodiments, the electronic device may further set a maximum number of iterations. When the model training reaches the maximum number of iterations, it is determined that the model has converged.

[0064] In the model training method of this embodiment, by feature encoding the first word in the annotated text of the first audio, since it is not necessary to feature encode the entire annotated text, the efficiency of obtaining the second encoding feature can be improved. At the same time, the dimension of the second encoding feature obtained by feature encoding the first word is smaller than the encoding feature obtained by feature encoding the entire annotated text. Therefore, when predicting and outputting the probability value of the first word based on the second encoding feature, the amount of calculation can be reduced, thereby improving the efficiency of determining the probability value, and further improving the efficiency of model training. In addition, using the first audio corresponding to the annotated text including the first word to participate in the training of the model can avoid invalid training of the model due to the fact that the first word is not included in the annotated text, thereby improving the training effect of the model.

[0065] like Figure 3 FIG. 1 is a flow chart of another model training method provided by an embodiment of the present application. The model training method is applied to electronic devices, for example, Figure 1 The electronic device 100. According to different requirements, the order of the steps in the flowchart can be changed, and some steps can be omitted.

[0066] S301: Perform feature encoding on the first audio to obtain a first encoding feature.

[0067] In at least one embodiment of the present application, the first audio may include speech signals in a variety of application scenarios. The electronic device may perform feature encoding on the first audio using a coding network in the model to obtain first coding features. The coding network may be a deep learning network, a recurrent neural network, or the like. The coding network may also be an encoding module of a Conformer model, and this application does not impose any specific limitations thereon.

[0068] In at least one embodiment of the present application, the electronic device may perform feature encoding on the first audio in detail. Figure 2 The detailed description of step S201 is omitted here.

[0069] S302: Determine a second loss value according to the first audio and the first coding feature.

[0070] In at least one embodiment of the present application, the second loss value may be used to indicate the loss condition of the encoding network in the model.

[0071] In at least one embodiment of the present application, the electronic device determines a second loss value based on the first audio and the first coding feature, including: the electronic device maps the first coding feature to obtain a fifth vector; determines the posterior probability of the fifth vector to the first audio, and calculates the second loss value based on the posterior probability of the fifth vector to the first audio.

[0072] The fifth vector may be obtained by mapping the first encoding feature via a first linear layer, and the first linear layer may be pre-set and adjustable. For example, the first encoding feature of NxMx256 may be mapped to a fifth vector of NxMx5000 via the first linear layer, where 5000 may represent the total number of preset characters in a preset dictionary, and the preset dictionary may be pre-set and adjustable according to actual needs.

[0073] The calculation formula for the second loss value is: ,in, Can represent the second loss value, It can represent the first audio, can represent the fifth vector, It can represent the posterior probability corresponding to the fifth vector. The posterior probability of the fifth vector to the first audio can be obtained by processing the fifth vector through an activation function. The activation function can be a softmax function.

[0074] The embodiment of the present application can obtain a fifth vector by mapping the first coding feature. The fifth vector can be used to determine the posterior probability of the fifth vector to the first audio, thereby quickly determining the second loss value.

[0075] S303: Determine a third loss value according to the annotated text and the decoding vector of the first encoding feature.

[0076] In at least one embodiment of the present application, during the model training process, the electronic device can determine the loss of the model in conjunction with a decoding network. The decoding network can be a decoding module of a Transformer model, but actual applications are not limited thereto.

[0077] In at least one embodiment of the present application, the electronic device may decode the first coding feature through a decoding network to obtain a decoding vector. For example, the decoding network may include k decoding layers, k is a positive integer greater than or equal to 1, and each decoding layer in the decoding network may include: a first attention layer with a mask operation, a second attention layer, and a feedforward neural network layer. The input of the first decoding layer in the decoding network is the first coding feature, and the input of the second decoding layer to the kth decoding layer in the decoding network includes the first coding feature and the output of the previous decoding layer. The electronic device uses the k decoding layers in the decoding network to decode the first coding feature in sequence to obtain a decoding vector. The decoding vector can be a vector of NxKx256, where N can represent the number of first audios input to the encoding network, K can represent the length of the text corresponding to the first audio, and 256 can represent the output dimension.

[0078] In at least one embodiment of the present application, the electronic device determines a third loss value based on the decoded vector and the annotated text. In one example, if the encoding network is able to accurately encode the first audio feature, the third loss value can be used to indicate the loss of the decoding network. In another example, if the decoding network is able to accurately decode the first feature code, the third loss value can also be used to indicate partial loss of the encoding network.

[0079] In at least one embodiment of the present application, the electronic device performs mapping processing on the decoded vector to obtain a sixth vector, determines a posterior probability of the sixth vector for the annotated text, and calculates a third loss value based on the posterior probability of the sixth vector for the annotated text.

[0080] In some embodiments, the sixth vector may be obtained by mapping the decoded vector through a second linear layer, and the second linear layer may be pre-set or adjusted. For example, the decoded vector of NxKx256 may be mapped to the sixth vector of NxMx5000 through the second linear layer.

[0081] In some embodiments, the third loss value is calculated as follows: ,in, It can represent the third loss value, Z can represent the total number of preset characters in the preset dictionary, represents the posterior probability that the decoding network predicts the decoding vector as the i-th preset character, It can represent the probability corresponding to the i-th preset character when the i-th preset character is a real character in the annotation text (for example, ), , It can be set and adjusted according to actual needs. When the i-th preset character is not a real character in the annotation text (for example, ), , where Z may represent the total number of preset characters in the preset dictionary.

[0082] In the embodiment of the present application, a sixth vector can be obtained by mapping the decoded vector, and the posterior probability of the sixth vector to the annotated text can be obtained through the sixth vector, so that the third loss value can be determined.

[0083] S304: Train the model based on the second loss value and the third loss value.

[0084] In at least one embodiment of the present application, to accurately determine the loss of the coding network in the model, the electronic device can calculate the total loss value of the coding network in the model based on the second loss value and the third loss value. The total loss value can be determined as a weighted sum of the second loss value and the third loss value, or as the product of the second loss value and the third loss value, but is not limited to this in practical applications.

[0085] In at least one embodiment of the present application, the electronic device adjusts the coding network in the model based on the total loss value of the coding network until the total loss value reaches a first preset condition. The first preset condition may include, but is not limited to: the total loss value no longer decreases, and the total loss value is less than or equal to a third threshold value. The third threshold value can be set and adjusted as needed.

[0086] In another embodiment, the electronic device may further adjust the encoding network in the model according to the total loss value of the encoding network until the number of adjustments or the learning rate reaches a second preset condition. The second preset condition may include, but is not limited to: the number of adjustments reaching a preset number or the learning rate meeting a preset requirement.

[0087] In another embodiment, the electronic device may further adjust the encoding network in the model based on the second loss value until the first loss value satisfies the first preset condition and / or the number of adjustments or the learning rate reaches the second preset condition.

[0088] In multiple embodiments of the present application, by combining the second loss value of the encoding network and the third loss value of the decoding network, the loss of the encoding network in the model can be accurately quantified, thereby accurately training the model.

[0089] like Figure 4 FIG. 1 is a flow chart of a speech recognition method provided by an embodiment of the present application. The speech recognition method is applied to electronic devices, for example, Figure 1 The electronic device 100. According to different requirements, the order of the steps in the flowchart can be changed, and some steps can be omitted.

[0090] S401 , recognizing a second audio, obtaining a third word in the text of the second audio and time frame information of the third word in the second audio.

[0091] In at least one embodiment of the present application, the second audio may be a voice signal in a variety of application scenarios, such as a customer service call quality inspection scenario, a voice text error correction scenario, an intelligent voice question and answer scenario, and the like.

[0092] In at least one embodiment of the present application, the electronic device can recognize the second audio through a model. The model can include an encoding network and a recognition network. Figure 5 As shown, Figure 5 is a schematic diagram of the model structure. Figure 5 As shown, the model includes an encoding network, which can include m first encoding layers and n second encoding layers. Each first encoding layer can include a convolutional layer, a first pooling layer, and a first fully connected layer, and each second encoding layer can include an attention layer, a second pooling layer, and a second fully connected layer. The model also includes a recognition network, which includes a multi-head attention layer, a fusion layer, a third fully connected layer, and a decoding layer.

[0093] In at least one embodiment of the present application, the electronic device performs feature encoding on the second audio using the encoding network in the model to obtain a third feature code, and then uses the recognition network in the model to predict the third feature code to obtain a third word in the text of the second audio and time frame information of the third word in the second audio. The third word can be an address word or a word related to other professional fields. The time frame information can be used to indicate the time frame in which the third word occurs in the second audio.

[0094] S402: Extract an audio segment from the second audio based on the time frame information.

[0095] In at least one embodiment of the present application, the electronic device may determine the audio corresponding to the time frame information in the second audio as an audio segment. For example, if the second audio includes 100 frames of audio and the time frame information is frames 5 to 8, the electronic device may extract the audio between frames 5 and 8 from the second audio as the audio segment.

[0096] In another embodiment, the electronic device determines the target frame based on the time frame information and the preset number of frames, and the electronic device uses the audio corresponding to the target frame in the second audio as an audio segment, wherein the preset number of frames can be set and adjusted according to actual needs. For example, the second audio includes 100 frames of audio, and the time frame information is the 5th to 8th frames. If the preset number of frames is 2 frames, the target frame can be the 3rd to 10th frames, and the electronic device extracts the audio in the 3rd to 10th frames from the second audio as an audio segment. This embodiment can avoid the situation where the third word is not included in the audio segment or the included content is incomplete due to deviations in the time frame information by presetting the number of frames, thereby ensuring the accuracy and completeness of the audio segment extraction.

[0097] S403 : Based on the audio segment, determine a fifth word in the text of the second audio from the third word and the preset fourth word.

[0098] In at least one embodiment of the present application, the preset fourth vocabulary may include vocabulary related to a specific professional field. For example, the preset fourth vocabulary may include address words, such as "XX District, XX City." Another example is that the preset fourth vocabulary may also include vocabulary related to the insurance field, such as "premium."

[0099] In at least one embodiment of the present application, the electronic device determines the fifth word in the text of the second audio from the third word and the preset fourth word based on the audio clip, including: the electronic device extracts audio features from the audio clip, and obtains features corresponding to each fourth word; the electronic device calculates the similarity between the audio features and the features corresponding to each fourth word; if the maximum similarity among the similarities is greater than a preset threshold, the fourth word corresponding to the maximum similarity is determined as the fifth word; if the maximum similarity is less than or equal to the preset threshold, the third word is determined as the fifth word.

[0100] In some embodiments, during the process of extracting audio features from an audio clip, the electronic device may perform feature encoding on the audio clip using a coding network to obtain a fourth feature code. The electronic device may pre-train a neural network model, which may include a multi-layer time-delay neural network. The electronic device inputs the fourth feature code corresponding to the audio clip into the neural network model to obtain the audio features.

[0101] In some embodiments, each fourth word may correspond to a feature, and the feature corresponding to each fourth word may be obtained by encoding the third audio corresponding to the fourth word through a neural network model, wherein the neural network model may be formed by stacking multiple time-delay neural networks. For example, each user sends a third audio for the fourth word "XX City XX District", and 20 third audios may be obtained. The neural network model may perform feature encoding on the 20 third audios respectively, and obtain 20 256 vectors, 20 Taking the mean of the 256 vectors, we get 1 256 features.

[0102] In some embodiments, the electronic device may calculate the similarity between the audio feature and the feature corresponding to each fourth word using formulas such as the Cosine formula and the Euclidean distance formula. In other embodiments, the electronic device may calculate the difference between each element in the audio feature and the corresponding element in the feature corresponding to the fourth word. For example, the difference between the element in the first row and tenth column of the audio feature and the element in the first row and tenth column of the feature corresponding to the fourth word may be calculated. Further, the electronic device determines the similarity between the audio feature and the feature corresponding to each fourth word based on all the calculated differences. In one example, the electronic device performs a weighted sum calculation on all the differences to obtain the similarity. In another example, the electronic device calculates the sum of all the differences to obtain the similarity.

[0103] In the embodiment of the present application, feature extraction is performed on an audio clip using a neural network model to accurately obtain audio features. By comparing the similarity between the audio features and the features corresponding to each fourth word, when the maximum similarity among multiple similarities is greater than a preset threshold, the fourth word with the maximum similarity is determined as the fifth word in the text of the second audio, thereby improving the accuracy of determining the fifth word. When the maximum similarity is less than or equal to the preset threshold, the third word is determined as the fifth word in the text of the second audio, thereby avoiding incorrect replacement of the fifth word and thus improving the accuracy of the fifth word.

[0104] In at least one embodiment of the present application, the electronic device obtains a corresponding response result from a preset question-and-answer library based on the fifth word in the text of the second audio. In another embodiment, the electronic device analyzes the fifth word in the text of the second audio using a large language model to obtain a response result for the second audio.

[0105] In the speech recognition method of this embodiment, by recognizing the second audio, the third word in the text of the second audio and the time frame information of the third word in the second audio can be obtained. Based on the time frame information, the audio segment containing the third word can be accurately extracted from the second audio. Based on the audio segment, the fifth word in the text of the second audio can be determined from the third word and a preset fourth word, thereby improving the recognition accuracy of the fifth word in the text of the second audio.

[0106] like Figure 6 , is a functional module diagram of a model training device provided by an embodiment of the present application. The model training device 61 includes an encoding unit 610, a prediction unit 611, a calculation unit 612, a training unit 613 and a determination unit 614. The module / unit referred to in this application refers to a unit that can be processed by a processor (e.g. Figure 8 The processor 1101 shown in FIG. 1 is obtained and is capable of performing a series of computer-readable instruction segments that are stored in a memory (eg, Figure 81102).

[0107] In one embodiment, the encoding unit 610 is used to perform feature encoding on the first audio to obtain a first encoding feature, and to perform feature encoding on the first word in the annotated text of the first audio to obtain a second encoding feature; the prediction unit 611 is used to predict and output the probability value of the first word based on the first encoding feature and the second encoding feature; the calculation unit 612 is used to calculate the first loss value based on the first encoding feature and the probability value; the training unit 613 is used to train the model based on the first loss value.

[0108] In one embodiment, the prediction unit 611 is specifically used to: perform a linear transformation on the first encoding feature to obtain a key vector and a value vector corresponding to each attention head, and perform a linear transformation on the second encoding feature to obtain a query vector corresponding to each attention head; perform an attention operation on the query vector, the key vector, and the value vector to obtain a first vector output by each attention head; based on the first vector output by each attention head, determine the second vector of the model for the first audio; and based on the second vector, determine the probability value of the first vocabulary.

[0109] In one embodiment, the calculation unit 612 is specifically used to: obtain a third vector based on the first coding feature and the probability value; normalize the third vector to obtain a fourth vector; decode the fourth vector to obtain a second vocabulary output by the model for the first audio; and calculate a first loss value based on the first vocabulary and the second vocabulary.

[0110] In one embodiment, the determination unit 614 is used to determine the second loss value based on the first audio and the first encoding feature; the determination unit 614 is also used to determine the third loss value based on the annotated text and the decoding vector of the first encoding feature; the training unit 613 is also used to train the model based on the second loss value and the third loss value.

[0111] In one embodiment, the determination unit 614 is specifically configured to: perform mapping processing on the first coding feature to obtain a fifth vector; determine a posterior probability of the fifth vector for the first audio; and calculate a second loss value based on the posterior probability.

[0112] In multiple embodiments of the present application, by feature encoding the first word in the annotated text of the first audio, since it is not necessary to feature encode the entire annotated text, the efficiency of obtaining the second encoding feature can be improved. At the same time, the dimension of the second encoding feature obtained by feature encoding the first word is smaller than the encoding feature obtained by feature encoding the entire annotated text. Therefore, when predicting and outputting the probability value of the first word based on the second encoding feature, the amount of calculation can be reduced, thereby improving the efficiency of determining the probability value, and further improving the efficiency of model training. In addition, by using the first audio corresponding to the annotated text including the first word to participate in the training of the model, it is possible to avoid invalid training of the model due to the fact that the first word is not included in the annotated text, thereby improving the training effect of the model.

[0113] like Figure 7 , is a functional module diagram of a speech recognition device provided by an embodiment of the present application. The speech recognition device 71 includes a recognition module 710, an extraction module 711 and a determination module 712. The module / unit referred to in this application refers to a unit that can be processed by a processor (e.g. Figure 8 The processor 1101 shown in FIG. 1 is obtained and is capable of performing a series of computer-readable instruction segments that are stored in a memory (eg, Figure 8 1102).

[0114] In one embodiment, the recognition module 710 is used to recognize the second audio, obtain the third word in the text of the second audio and the time frame information of the third word in the second audio; the extraction module 711 is used to extract the audio segment from the second audio based on the time frame information; the determination module 712 is used to determine the fifth word in the text of the second audio from the third word and the preset fourth word based on the audio segment.

[0115] In one embodiment, the determination module 712 is specifically used to: extract audio features from the audio clip and obtain features corresponding to each fourth word; calculate the similarity between the audio features and the features corresponding to each fourth word; if the maximum similarity among the similarities is greater than a preset threshold, determine the fourth word corresponding to the maximum similarity as the fifth word; if the maximum similarity is less than or equal to the preset threshold, determine the third word as the fifth word.

[0116] In various embodiments of the present application, by recognizing the second audio, a third word in the text of the second audio and information about the time frame of the third word in the second audio can be obtained. Based on the time frame information, an audio segment containing the third word can be accurately extracted from the second audio. Using the audio segment, a fifth word in the text of the second audio can be determined from the third word and a preset fourth word, thereby improving the recognition accuracy of the fifth word in the text of the second audio.

[0117] Figure 8 Schematic diagram of the structure of an electronic device for implementing a model training method and a speech recognition method provided in an embodiment of the present application. Figure 8 The electronic device 100 is used to perform Figure 2 、 Figure 4 The method shown.

[0118] The electronic device 100 includes at least one processor 1101 , a memory 1102 , and at least one network interface 1103 .

[0119] The processor 1101 is, for example, a general-purpose central processing unit (CPU), a network processor (NP), a graphics processing unit (GPU), a neural-network processing unit (NPU), a data processing unit (DPU), a microprocessor, or one or more integrated circuits for implementing the solution of the present application. For example, the processor 1101 includes an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The PLD is, for example, a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0120] The memory 1102 may be, for example, a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, a random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, an optical disc storage (including a compact disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but is not limited thereto. Optionally, the memory 1102 exists independently and is connected to the processor 1101 via the internal connection 1104. Alternatively, the memory 1102 and the processor 1101 may be integrated together.

[0121] The network interface 1103 uses any transceiver-like device for communicating with other devices or communication networks. For example, the network interface 1103 includes at least one of a wired network interface and a wireless network interface. For example, the wired network interface is an Ethernet interface. For example, the Ethernet interface is an optical interface, an electrical interface, or a combination thereof. For example, the wireless network interface is a wireless local area network (WLAN) interface, a cellular network interface, or a combination thereof.

[0122] In some embodiments, the processor 1101 includes one or more CPUs, such as Figure 8 CPU0 and CPU1 are shown in the figure.

[0123] In some embodiments, the electronic device 100 optionally includes multiple processors, such as Figure 8 1 and 1105. Each of these processors is, for example, a single-CPU or a multi-CPU. A processor herein optionally refers to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).

[0124] In some embodiments, electronic device 100 further includes internal connections 1104. Processor 1101, memory 1102, and at least one network interface 1103 are connected via internal connections 1104. Internal connections 1104 include pathways for transmitting information between these components. Internal connections 1104 may optionally be a single board or bus. Internal connections 1104 may optionally be divided into an address bus, a data bus, a control bus, and the like.

[0125] In some embodiments, the electronic device 100 further includes an input / output interface 1106 , which is connected to the internal connection 1104 .

[0126] Optionally, the processor 1101 implements the method in the above embodiment by reading the program code 910 stored in the memory 1102, or the processor 1101 implements the method in the above embodiment by internally stored program code. In the case where the processor 1101 implements the method in the above embodiment by reading the program code 910 stored in the memory 1102, the memory 1102 stores the program code that implements the method provided in the embodiment of the present application.

[0127] For more details on how the processor 1101 implements the above functions, please refer to the descriptions in the previous method embodiments, which will not be repeated here.

[0128] This embodiment also provides a computer storage medium, which stores computer instructions. When the computer instructions are executed on an electronic device, the electronic device executes the above-mentioned related method steps to implement the model training method and speech recognition method in the above-mentioned embodiment.

[0129] This embodiment also provides a computer program product. When the computer program product is run on an electronic device, the electronic device executes the above-mentioned related steps to implement the model training method and speech recognition method in the above-mentioned embodiment.

[0130] In addition, an embodiment of the present application also provides a device, which can specifically be a chip, component or module, and the device may include a connected processor and memory; wherein the memory is used to store computer-executable instructions, and when the device is running, the processor can execute the computer-executable instructions stored in the memory to enable the chip to execute the model training method and speech recognition method in the above-mentioned method embodiments.

[0131] Among them, the electronic device, computer storage medium, computer program product or chip provided in this embodiment is used to execute the corresponding method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method provided above, and will not be repeated here.

[0132] Through the description of the above implementation methods, technical personnel in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0133] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0134] The units described as separate components may or may not be physically separate, and the components shown as units may be one physical unit or multiple physical units, that is, they may be located in one place or distributed in multiple places. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0135] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0136] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a device (which can be a single-chip microcomputer, chip, etc.) or a processor to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program code.

[0137] The above are only specific embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any changes or replacements within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A model training method, characterized in that: The method comprises: Performing feature encoding on the first audio to obtain a first encoding feature, and performing feature encoding on a first word in the annotated text of the first audio to obtain a second encoding feature; Predicting and outputting a probability value of the first vocabulary word based on the first coding feature and the second coding feature; Calculating a first loss value according to the first coding feature and the probability value; A model is trained based on the first loss value.

2. The model training method according to claim 1, characterized in that The predicting and outputting the probability value of the first word according to the first coding feature and the second coding feature includes: Performing a linear transformation on the first encoded feature to obtain a key vector and a value vector corresponding to each attention head, and performing a linear transformation on the second encoded feature to obtain a query vector corresponding to each attention head; Performing an attention operation on the query vector, the key vector, and the value vector to obtain a first vector output by each attention head; Determining a second vector of the model for the first audio based on the first vector output by each attention head; Based on the second vector, a probability value of the first vocabulary word is determined.

3. The model training method according to claim 1, characterized in that The calculating a first loss value according to the first coding feature and the probability value includes: Obtaining a third vector based on the first coding feature and the probability value; Normalizing the third vector to obtain a fourth vector; Decoding the fourth vector to obtain a second vocabulary output by the model for the first audio; The first loss value is calculated according to the first vocabulary and the second vocabulary.

4. The model training method according to claim 1, characterized in that The method further comprises: determining a second loss value based on the first audio and the first coding feature; determining a third loss value based on the annotated text and the decoded vector of the first encoding feature; The model is trained based on the second loss value and the third loss value.

5. The model training method according to claim 4, characterized in that The determining a second loss value according to the first audio and the first coding feature includes: Performing mapping processing on the first coding feature to obtain a fifth vector; Determine a posterior probability of the fifth vector for the first audio, and calculate the second loss value according to the posterior probability.

6. A speech recognition method, characterized in that: The method comprises: Recognize the second audio, obtain a third word in the text of the second audio and time frame information of the third word in the second audio; extracting an audio segment from the second audio based on the time frame information; Based on the audio segment, a fifth word in the text of the second audio is determined from the third word and a preset fourth word.

7. The speech recognition method according to claim 6, characterized in that The step of determining, based on the audio segment, a fifth word in the text of the second audio from the third word and a preset fourth word includes: Extracting audio features from the audio clip, and obtaining features corresponding to each fourth word; Calculating the similarity between the audio feature and the feature corresponding to each fourth word; If the maximum similarity among the similarities is greater than a preset threshold, determining the fourth word corresponding to the maximum similarity as the fifth word; If the maximum similarity is less than or equal to the preset threshold, the third word is determined as the fifth word.

8. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the model training method according to any one of claims 1 to 5 or the speech recognition method according to any one of claims 6 to 7 is implemented.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the model training method according to any one of claims 1 to 5 or the speech recognition method according to any one of claims 6 to 7.

10. A computer program product, characterized in that The computer program product includes a computer program, which, when executed by a processor, implements the model training method according to any one of claims 1 to 5 or the speech recognition method according to any one of claims 6 to 7.